Method and system for haplotype level transposon insertion identification based on physical typing
By aligning sequencing data to two haplotype genomes and allocating sequences based on alignment scores and mismatches, different transposon insertion identification methods were employed. This solved the problem of insufficient accuracy in haplotype-level identification in existing methods, achieving more precise transposon insertion identification and genetic information resolution.
Patent Information
- Application Number
- CN202511034782.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing methods are insufficient for haplotype-level transposon insertion identification, resulting in inadequate accuracy and failure to fully utilize the advantages of HiFi sequences.
By aligning sequencing data to two haplotype genomes, sequence allocation was performed based on alignment scores and the number of mismatches. Different transposon insertion identification methods were used to identify haplotype-level transposon insertions for single-end long reads and paired-end short reads. The alignment files were then separated based on the physical typing results to perform haplotype-level transposon insertion identification.
It improves the accuracy and resolution of transposon insertion identification, and can more accurately reflect the insertion of transposons in different haplotypes, providing more refined genetic information.
Smart Images

Figure CN120526846B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of genome sequencing, in particular to a haplotype level transposon insertion identification method and system based on physical typing. BACKGROUND
[0002] Transposon is a kind of exogenous DNA sequence that can spontaneously transpose, which widely exists in host genomes. Since the insertion site of transposon after transposition has no directionality, it will cause harmful effects such as gene transcription silencing, loss of protein function, etc., which seriously affect the stability and normal physiological function of host genomes. Therefore, accurate identification of transposon insertion site is of great significance for in-depth understanding of biological processes such as genome structure and function, gene expression regulation, and biological evolution.
[0003] At present, transposon insertion identification mainly relies on genome sequencing technology and bioinformatics analysis methods. The developed transposon insertion identification methods are mainly for second-generation sequencing data. A large number of genome sequences are obtained by sequencing, and transposon insertion signals such as broken alignment sequences and abnormal insertion fragment sizes after alignment are used to filter these insertion signals to obtain high-confidence transposon insertion events. With the development of sequencing technology, haplotype genomes of a large number of species have been completely assembled. Physical typing of sequencing sequences, i.e. distinguishing the sequence belonging to haplotype, respectively identifying transposon insertion in haplotype genomes, can distinguish transposon insertion events at haplotype level, which helps to improve the research of transposon insertion affecting allele expression difference, three-dimensional genome, etc. High Fidelity (HiFi) genome sequencing has the advantages of long read length and high accuracy, which can play an important role in transposon insertion identification.
[0004] However, the existing transposon insertion identification methods have many shortcomings, for example, the existing methods do not fully utilize transposon insertion signals, resulting in unstable accuracy of transposon insertion identification; the existing methods cannot achieve haplotype level transposon insertion identification; and there is no transposon insertion identification method based on HiFi sequence. In view of the above defects of the existing methods, a haplotype level transposon insertion precise identification method suitable for HiFi and second-generation sequencing is developed. SUMMARY
[0005] The present application provides a haplotype level transposon insertion identification method and system based on physical typing, which solves the problem of insufficient accuracy caused by the difficulty of existing methods to achieve haplotype level identification.
[0006] To solve the above technical problems, the present application provides a haplotype level transposon insertion identification method based on physical typing, comprising the following steps:
[0007] Step S1: aligning the sequencing data to two sets of haplotype genomes respectively, obtaining alignment scores and mismatch numbers of the sequencing data aligned to the two sets of haplotype genomes;
[0008] Step S2: for single-end long read sequencing data, assigning the sequencing sequence to the haplotype with a larger alignment score; for paired-end short read sequencing data, adding the alignment scores and the mismatch numbers of the paired-end short reads as the alignment score and the mismatch number of each pair of sequences, and assigning the sequencing sequence to the haplotype with a larger alignment score; if the alignment scores of the sequencing sequence aligned to the two sets of haplotype genomes are the same, assigning the sequencing sequence to the haplotype with a smaller mismatch number; if the alignment scores and the mismatch numbers of the sequencing sequence aligned to the two sets of haplotype genomes are the same, assigning the sequencing sequence to both sets of haplotypes;
[0009] Step S3: based on the physical typing results, separating the alignment files of the two sets of haplotype genomes, and based on the alignment files, identifying the transposon insertion at the haplotype level.
[0010] Preferably, different transposon insertion identification methods are used for the single-end long read and the paired-end short read in step S3.
[0011] Preferably, the transposon insertion identification method for the single-end long read comprises the following steps:
[0012] Step S301: performing haplotype-level variation identification on the alignment file, taking structural variations with an insertion length exceeding a set threshold as large fragment insertion type structural variations, and extracting insertion sequences of the large fragment insertion type structural variations from the alignment file;
[0013] Step S302: performing self nucleic acid sequence alignment analysis on the insertion sequence, detecting the repeat structure inside the insertion sequence, and verifying the end sequence integrity of the insertion sequence;
[0014] Step S303: aligning the full-length transcriptome data to the insertion sequence, identifying potential coding regions, analyzing the domains of the insertion sequence according to the potential coding regions, and inferring the functional domains of the insertion sequence through the domains;
[0015] Step S304: aligning the extracted insertion sequence with a transposon database, determining the transposon type of the insertion sequence according to the database alignment results, domain and functional domain information, updating the annotation information in the database if the insertion sequence is a known transposon, and defining a new family and submitting it to the database if it is an unknown transposon.
[0016] Preferably, the transposon insertion identification method for the paired-end short read comprises the following steps:
[0017] Step S311: performing haplotype-level variant calling on the alignment file, screening abnormal alignment sequences in the whole genome range of the alignment file through local depth anomaly and abnormal alignment signal;
[0018] Step S312: aligning the abnormal alignment sequences with target transposon sequences, and taking completely matched sites as candidate insertion sites;
[0019] Step S313: experimentally verifying the local depth anomaly and abnormal alignment signal of the candidate insertion site by the three-primer method, and confirming the authenticity of the insertion.
[0020] Preferably, in step S3, the alignment files of the two sets of haplotype genomes are separated based on the physical typing results, including the following steps: according to the sequence number list in the typing result of the sequencing data, extracting the sequence alignment records respectively belonging to the two sets of haplotype genomes from the alignment file to form two sets of typing result files.
[0021] The application also provides a haplotype-level transposon insertion identification system based on physical typing, which is realized based on the above-mentioned haplotype-level transposon insertion identification method based on physical typing, and includes a data input module, a haplotype alignment module, a physical typing distribution module, a haplotype alignment file generation module, and a haplotype-level transposon insertion identification module.
[0022] The data input module receives raw sequencing data, removes low-quality sequences, and distinguishes between single-end long reads and double-end short reads.
[0023] The haplotype alignment module aligns the sequencing data to two sets of haplotype genomes respectively, calculates the alignment score of each sequence, and outputs two sets of raw alignment files of haplotype genomes.
[0024] The physical typing distribution module distributes each sequencing sequence to a haplotype genome according to the alignment score, and outputs the typing result of the sequencing sequence.
[0025] The haplotype alignment file generation module extracts the sequences distributed to each haplotype from the raw alignment file according to the typing result, and generates haplotype-specific alignment files.
[0026] The haplotype-level transposon insertion identification module detects transposon insertion in the haplotype-specific alignment file, and determines the insertion type in combination with structural variation and functional verification.
[0027] Preferably, the physical typing distribution module adopts different distribution methods for single-end long reads and double-end short reads.
[0028] The single-end long reads are distributed to the haplotype with a higher alignment score.
[0029] The paired-end short reads are assigned to haplotypes with higher total alignment scores by adding alignment scores of the paired-end short reads.
[0030] Preferably, the haplotype-level transposon insertion identification module employs different transposon insertion identification methods for the single-end long reads and the paired-end short reads.
[0031] Preferably, the method for transposon insertion identification of the haplotype-level transposon insertion identification module on the single-end sequencing sequences comprises the following steps:
[0032] Step 1: performing haplotype-level variation identification on the alignment file, taking structural variations with insertion lengths exceeding a set threshold as large-insert-type structural variations, and extracting insertion sequences of the large-insert-type structural variations from the alignment file;
[0033] Step 2: performing self-nucleic-acid-sequence alignment analysis on the insertion sequences, detecting repeat structures within the insertion sequences, and verifying end sequence integrity of the insertion sequences;
[0034] Step 3: aligning full-length transcriptome data to the insertion sequences, identifying potential coding regions, analyzing domains of the insertion sequences according to the potential coding regions, and inferring functional domains of the insertion sequences through the domains;
[0035] Step 4: aligning the extracted insertion sequences to a transposon database, determining a transposon type of the insertion sequences according to database alignment results, domain and functional domain information, updating annotation information in the database if the insertion sequences are known transposons, and defining a new family and submitting to the database if the insertion sequences are unknown transposons.
[0036] Preferably, the method for transposon insertion identification of the haplotype-level transposon insertion identification module on the paired-end short reads comprises the following steps:
[0037] Step 1: performing haplotype-level variation identification on the alignment file, screening abnormal alignment sequences in a whole genome range of the alignment file through local depth abnormality and abnormal alignment signals;
[0038] Step 2: aligning the abnormal alignment sequences to target transposon sequences, and taking completely matched sites as candidate insertion sites;
[0039] Step 3: experimentally verifying local depth abnormality and abnormal alignment signals of the candidate insertion sites through a three-primer method, and confirming authenticity of the insertion.
[0040] The present application has at least the following advantages:
[0041] 1. By respectively aligning sequencing data to two sets of haplotype genomes and assigning sequencing sequences according to alignment scores, sequencing sequences can be more accurately positioned to specific haplotypes, reducing errors caused by sequence similarity and other factors, thereby improving the accuracy of transposon insertion identification, making the identification results more reliable and more truly reflecting the insertion of transposons in different haplotypes.
[0042] 2. Based on the physical typing results, the alignment files of the two sets of haplotype genomes are separated, so that the transposon insertion identification can be performed at the haplotype level, thereby more clearly distinguishing the insertion differences of transposons in the two sets of haplotypes, improving the resolution of identification. Compared with the traditional identification on the non-typing genome, the haplotype level transposon insertion identification can provide more detailed genetic information. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 The method flowchart of the embodiment of the present application is shown in the figure.
[0044] Figure 2 The method flowchart of the embodiment of the present application is shown in the figure.
[0045] Figure 3 The method flowchart of the embodiment of the present application is shown in the figure.
[0046] Figure 4 The method flowchart of the embodiment of the present application is shown in the figure.
[0047] Figure 5 The method flowchart of the embodiment of the present application is shown in the figure.
[0048] Figure 6 The result figure of the insertion sequence self-alignment of the embodiment of the present application is shown in the figure.
[0049] Figure 7 The three-primer method verification gel map of the transposon insertion of navel orange at the genomic site V-M4:23817439 is shown in the figure.
[0050] Figure 8 The result schematic diagram of the transposon insertion of navel orange Illumina sequencing sequence detection at the genomic site V-M4:23817439 is shown in the figure. DETAILED DESCRIPTION
[0051] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application.
[0052] With the continuous development of sequencing technology, High Fidelity (HiFi) sequencing has become one of the most advanced sequencing technologies due to its comprehensive advantages in sequencing length and sequencing accuracy, and is widely used in the fields of genome assembly and variant detection. In the process of genome typing assembly, two methods of family typing and physical typing are mainly used. Family typing is to divide the genome of offspring into two sets of haplotypes from the paternal and maternal sources by using the genotype information of the parents; and physical typing is to divide the two sets of haplotypes into different sets of haplotypes by using the genotype difference between the two sets of haplotypes, but the two sets of haplotypes obtained by this method cannot distinguish the genetic information from the paternal and maternal sources. Since the biological parents of most naturally occurring animals and plants are difficult to determine, physical typing has become a commonly used method for genome typing.
[0053] Allelic difference is one of the important factors leading to trait variation, and to detect allelic difference, the first prerequisite is to accurately type the sequencing sequence, that is, to correctly allocate the sequencing sequence to two sets of haplotypes. On this basis, further variant detection and identification at the haplotype level are carried out, and finally the variation between alleles can be obtained. However, there is currently no effective method for physical typing of sequencing sequences. In view of this, the embodiments of the present application provide a method for physical typing of sequencing sequences, aiming to realize accurate variant detection at the haplotype level, and to provide a more powerful method for in-depth study of genome structure and function by comparing sequence differences at the haplotype level.
[0054] As shown in Figure 1 The embodiments of the present application provide a haplotype level transposon insertion identification method based on physical typing, comprising the following steps:
[0055] Step S1: aligning the sequencing data to two sets of haplotype genomes respectively, and obtaining alignment scores of the sequencing data after being aligned to the two sets of haplotype genomes.
[0056] Specifically, as shown in Figure 2 and Figure 3As shown, the embodiments of the present application respectively perform physical typing and variant detection on single-end long reads and paired-end short reads. The single-end long reads and paired-end short reads used in the embodiments of the present application are HiFi sequencing sequences and Illumina sequencing sequences respectively, and the HiFi or Illumina sequencing sequences are aligned to two sets of haplotype genomes to obtain aligned bam files. The alignment score (AS) and the number of mismatches (NM) after the HiFi or Illumina alignment to the two sets of haplotypes are extracted from the bam files.
[0057] Step S2: For single-end long reads, the single-end long reads are assigned to the haplotype with a higher alignment score; for paired-end short reads, the alignment scores of the paired-end short reads are added as the alignment score of each pair of sequences, and each pair of sequences is assigned to the haplotype with a larger alignment score.
[0058] Specifically, for each HiFi sequence, the sequence is assigned to the haplotype with a larger alignment score and a smaller number of mismatches to generate a typed bam file. For Illumina sequences, since Illumina sequencing obtains paired-end sequences, the alignment scores and the number of mismatches of the paired-end sequences are added as the alignment score and the number of mismatches of each pair of sequences, respectively, and each pair of sequences is assigned to the haplotype with a larger alignment score and a smaller number of mismatches.
[0059] It should be noted that the alignment score (AS) and the number of mismatches (NM) are negatively correlated in most cases, i.e., the fewer the mismatches, the higher the alignment score. This is because the alignment algorithm usually imposes a penalty on mismatches, and a positive score is given to matches. However, there are some exceptions, for example, the number of mismatches is low, but there are a large number of insertions or deletions in the sequence, and the alignment score may be significantly reduced due to the cumulative penalty of insertions or deletions. For another example, different tools have different weights for mismatches, insertions / deletions (InDels), and matches, and some tools may impose a higher penalty on long InDels, resulting in a lower AS even if the NM is low.
[0060] In the case of conflict between the alignment score and the number of mismatches, the embodiments of the present application give priority to the alignment score, because AS takes into account the comprehensive weight of, for example, matches, insertions / deletions, and can better reflect the overall alignment quality. If the AS is similar, the haplotype with a lower NM is selected. If neither the AS nor the NM can be distinguished, the sequencing sequence is assigned to both haplotypes.
[0061] Step S3: Based on the physical typing result, the alignment files of the two sets of haplotype genomes are separated, and the transposon insertion identification at the haplotype level is performed based on the alignment files.
[0062] Specifically, according to the sequence ID of the physical typing sequencing sequence, all alignment information aligned to the haplotype to which it belongs is extracted from the bam file to generate two haplotype-level bam alignment files. Using a variation identification software, haplotype-level variations of the physical typing sequencing sequence are identified, including single nucleotide polymorphisms (SNPs), small fragment insertion deletions (InDels), and large fragment structural variations (SVs).
[0063] The embodiment of the present application adopts different transposon insertion identification methods for single-end long reads and double-end short reads. As shown in the following table, the transposon insertion identification of single-end long reads, i.e., HiFi sequencing sequences, includes the following steps: Figure 4
[0064] Step 1: From the haplotype variation detection results after HiFi typing, structural variations (SVs) with an insertion length exceeding a set threshold are screened, which are defined as large fragment insertion type SVs. In the embodiment of the present application, the threshold is set to 2Kb.
[0065] Step 2: The insertion sequence of the above large fragment insertion type SV is extracted from the variation identification results, and the insertion sequence is subjected to self nucleic acid sequence blastn alignment analysis to detect the repeated structure inside the insertion sequence and verify the end sequence integrity of the insertion sequence.
[0066] Step 3: The full-length transcriptome data is aligned to the insertion sequence to identify the coding amino acid sequence, and the coding amino acid sequence is submitted to the Pfam database to determine the domain of the insertion sequence and infer the functional domain of the insertion sequence through the domain.
[0067] Step 4: The extracted insertion sequence is aligned with the transposon database, and the transposon type of the insertion sequence is determined according to the database alignment results, domain and functional domain information. If the insertion sequence is a known transposon, the annotation information in the database is updated, and if it is an unknown transposon, a new family is defined and submitted to the database.
[0068] As shown in the following table, the transposon insertion identification of double-end short reads, i.e., Illumina sequencing sequences, includes the following steps: Figure 5
[0069] Step 1: The target transposon sequence is added to the reference genome as a separate sequence, and the Illumina sequencing sequence is aligned to the reference genome to obtain the aligned bam file.
[0070] Step 2: A self-compiled script is used to identify haplotype-level variations of the aligned bam file, and abnormal alignment sequences are screened in the whole genome range of the alignment file through local depth abnormalities and stacking of abnormal alignment signals (Discordant reads) at the insertion site caused by terminal site sequence duplication (TSD).
[0071] Step 3: Align the abnormal alignment sequence with the target transposon sequence, and take the completely matched site as a candidate insertion site.
[0072] Step 4: Verify the local depth abnormality and abnormal alignment signal of the candidate insertion site by the three-primer method IGV, and confirm the authenticity of the insertion.
[0073] The embodiment of the present application also provides a haplotype level transposon insertion identification system based on physical typing, which is realized based on the above-mentioned haplotype level transposon insertion identification method based on physical typing, and comprises a data input module, a haplotype alignment module, a physical typing distribution module, a haplotype alignment file generation module and a haplotype level transposon insertion identification module.
[0074] The data input module receives original sequencing data, removes low-quality sequences, and distinguishes single-end long reads and double-end short reads.
[0075] The haplotype alignment module aligns the sequencing data to two sets of haplotype genomes respectively, calculates the alignment score of each sequence, and outputs the original alignment file of the two sets of haplotype genomes.
[0076] The physical typing distribution module distributes each sequencing sequence to the haplotype genome according to the alignment score, and outputs the typing result of the sequencing sequence.
[0077] The haplotype alignment file generation module extracts the sequence distributed to each haplotype from the original alignment file according to the typing result, and generates a haplotype-specific alignment file.
[0078] The haplotype level transposon insertion identification module detects the transposon insertion in the haplotype-specific alignment file, and determines the insertion type in combination with structural variation and functional verification.
[0079] The present application will be described in detail below in combination with specific embodiments.
[0080] Embodiment one
[0081] The single-end long reads, i.e. HiFi sequencing sequences of sweet orange, were physically typed and variant detection was performed at haplotype level. The HiFi data of 13 sweet orange varieties were measured and the effective data amount was obtained, and the statistical results are shown in Table 1. Among them, Variety is the variety, Seq num is the sequence number, Seq base is the sequence base number, Min len is the minimum sequence length, Max len is the maximum sequence length, Ave len is the average sequence length. AJTC is Egyptian sugar orange, CARA is Cara Cara navel orange, DH is big red sweet orange, DHWH is big red seedless sweet orange, GNZ is Gannan early navel orange, HSDQC is Washington navel orange, JHBTC is Jin Hong Bingtang orange, JXXC is Jingxu blood orange, LW is Lunwan navel orange, NHE is Newhall navel orange, TLKXC is Tarocco blood orange, TLWT is Trovita sweet orange, WLXY is Valencia sweet orange.
[0082] Table 1 Statistical results of sweet orange HiFi data
[0083]
[0084] The above 13 sweet orange HiFi data were aligned with the sweet orange end-to-end T2T (Telomere-to-Telomere) typing genome (http: / / citrus.hzau.edu.cn / download.php) using the minimap2 software, and two sets of haplotype genomes were obtained. The alignment information of each sweet orange variety, i.e. alignment score (AS) and number of base mismatches (NM), was extracted for each HiFi sequencing sequence using a self-compiled script. The HiFi sequencing sequences were assigned to the haplotype with larger alignment score and smaller number of base mismatches. The typing results of the HiFi sequencing sequences of each sweet orange variety are shown in Table 2. Among them, No. hap1 reads is the number of sequences assigned to haplotype 1; No. hap2 reads is the number of sequences assigned to haplotype 2; No. unphased reads is the number of untyped sequences; Unphased ratio is the proportion of untyped sequences; Median unphased ratio is the median of the proportion of untyped sequences.
[0085] Table 2 Typing results of sweet orange HiFi sequencing sequences
[0086]
[0087] The following command was used in the bcftools software to detect single nucleotide polymorphisms and small fragment insertion and deletion variants for the typing bam files of the two sets of haplotypes:
[0088] bcftools mpileup --threads 48 -O b -m 3 -q 30 -Q 20.
[0089] The detected single nucleotide polymorphisms and small fragment indels were filtered using the following command:
[0090] filter -e 'QUAL<30 || F_MISSING!=0 || FMT / DP<5 || F_PASS(GT="het")=1|| F_PASS(GT="RR")=1 || F_PASS(GT="AA")=1'.
[0091] The two haplotype typing bam files were detected for large fragment structural variations using the structural variation identification software pbsv, and the number of variation detections of 13 sweet orange HiFi sequences is shown in Table 3. Among them, No. SNP hap1 is the number of SNPs in haplotype 1; No. SNP hap2 is the number of SNPs in haplotype 2; No. SNP total is the total number of SNPs; No. InDel hap1 is the number of InDels in haplotype 1; No. InDel hap2 is the number of InDels in haplotype 2; No. InDel total is the total number of InDels; No. SV hap1 is the number of SVs in haplotype 1; No. SV hap2 is the number of SVs in haplotype 2; No. SV total is the total number of SVs.
[0092] Table 3 Number of haplotype-level variation detections after HiFi sequencing sequence typing
[0093]
[0094] The transposon insertion identification of haplotype-level large fragment insertion-type structural variations SV in sweet orange HiFi data includes the following steps:
[0095] Step 1: Extract the insertion sequence with a length of 6936 bp from the identified large fragment insertion-type SV.
[0096] Step 2: Align the extracted sequence with the transposon database, which can be aligned to the MULE type transposon.
[0097] Step 3: Perform self-blastn alignment of the sequence, and the alignment result is shown in Figure 6 As can be seen from Figure 6 , the sequence has four groups of repeated fragments inside.
[0098] Step 4: The full-length transcriptome of sweet orange was aligned to the sequence, and two potential coding regions ORF appeared. The protein sequences of the two ORFs were submitted to the Pfam database to identify the domains of the sequences. The two potential coding regions ORF1 of the sequence had a transmembrane domain, and ORF2 had a transposase, integrase domain.
[0099] Step 5: Based on the structure and coding characteristics of the sequence, the inserted sequence is TE6936, a MULE type DNA transposon.
[0100] Example Two
[0101] The single-end long-read, i.e. HiFi sequencing sequence, detected in the sweet orange individual, was detected as a double-end short-read, i.e. Illumina sequencing sequence, in a large-scale detection of the insertion in the sweet orange population. First, the sweet orange Illumina sequencing sequence was physically typed and the variation at the haplotype level was detected. The effective data amount was obtained by quality control statistics of 10 kinds of sweet orange Illumina data, and the statistical results are shown in Table 4. Among them, AJTC-2 is Egyptian sugar orange; Anliu is dark willow sweet orange; Bailiu NO is white willow navel orange; Banfield NO is banfield navel orange; Blood O-2 is blood orange; Blood O-jingxian is Jingxian blood orange; Blood O-Moro No is Moro blood orange; Blood O-Moro is Moro blood orange; Blood O is blood orange; Blood O-Tarocco Rosso is Tarocco blood orange; No. read pair is the number of double-end sequencing pairs; Length is the length of the sequencing sequence; depth is the sequencing depth.
[0102] Table 4 Statistical results of sweet orange Illumina data
[0103]
[0104] After aligning the above-mentioned 10 kinds of sweet orange Illumina data to the sweet orange T2T typing genome using the alignment software bwa, the alignment information of each pair of Illumina sequencing sequence, i.e. alignment score (AS) and number of mismatches (NM), was extracted using a self-compiled script. The alignment score and the number of mismatches of each pair of Illumina double-end sequencing sequence were added as the alignment score and the number of mismatches of the two sets of haplotype genomes. The Illumina sequencing sequence was assigned to the haplotype with a larger alignment score and a smaller number of base mismatches, and the typing results of the Illumina sequencing sequence are shown in Table 5.
[0105] Table 5 Typing results of sweet orange Illumina sequencing sequence
[0106]
[0107] The following command was used in the bcftools software to detect single nucleotide polymorphisms on the haplotype typing bam files of the two sets:
[0108] bcftools mpileup --threads 48 -O b -m 3 -q 30 -Q 20.
[0109] The following command was used to filter the detected single nucleotide polymorphisms:
[0110] filter -e 'QUAL<30 || F_MISSING!=0 || FMT / DP<5 || F_PASS(GT="het")=1|| F_PASS(GT="RR")=1 || F_PASS(GT="AA")=1'.
[0111] The number of single nucleotide polymorphisms obtained finally is shown in Table 6.
[0112] Table 6 Number of haplotype level variation detections after Illumina sequencing sequence typing
[0113]
[0114] The transposon insertion signal detection was performed on the bam file aligned to the transposon haplotype genome added above, and the abnormal alignment sequencing sequences near the insertion signal site were aligned with the transposon sequence; the site of the abnormal alignment sequence that can be aligned to the transposon sequence was taken as a candidate insertion site, and the transposon insertion signal of the candidate site was manually verified using IGV. Figure 7 The gel map of transposon insertion at the genomic site V-M4:23817439 of navel orange was verified by three-primer method, and a transposon internal primer was added for three-primer amplification. The short band is the genomic band without insertion, and the long band is the genomic band containing transposon insertion. Valencia sweet orange appears a short band as a control, indicating no insertion, and other navel oranges all appear two bands, indicating heterozygous insertion.
[0115] Figure 8 The results of transposon insertion at the genomic site V-M4:23817439 detected by navel orange Illumina sequencing sequence are shown in the schematic diagram. The position of the increased TSD coverage and the broken reads is the site of transposon insertion.
[0116] The technical features of the above embodiments can be combined in any manner. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, and only the preferred embodiments of the present application are expressed, which are described in more detail and in a more specific manner, but it should not be understood as a limitation on the scope of the present application. As long as the combinations of these technical features do not contradict each other, they should be considered as the scope of the present application.
[0117] It should be noted that, for those skilled in the art, there are still several modifications and improvements without departing from the concept of the present application, which are within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for identifying haplotype horizontal transposon insertions based on physical typing, characterized in that, Includes the following steps: Step S1: Compare the sequencing data with the genomes in the publicly available database of the corresponding varieties, and align the sequencing data to two haplotype genomes respectively, and obtain the alignment score and the number of mismatches after aligning the sequencing data to the two haplotype genomes; Step S2: For single-end long read sequencing data, assign the sequencing sequence to the haplotype with the higher alignment score; for paired-end short read sequencing data, add the alignment score and mismatch count of the paired-end short reads as the alignment score and mismatch count for each sequence pair, and assign the sequencing sequence to the haplotype with the higher alignment score; if the alignment scores are the same after aligning the sequencing sequence to two haplotype genomes, assign the sequencing sequence to the haplotype with fewer mismatches; if the alignment scores and mismatch counts are the same after aligning the sequencing sequence to two haplotype genomes, assign the sequencing sequence to both haplotypes simultaneously. Step S3: Based on the physical typing results, the sequencing data is separated and aligned to two sets of haplotype genome alignment files. Based on the alignment files, haplotype-level transposon insertion identification is performed, and different transposon insertion identification methods are used for the single-end long reads and the paired-end short reads. The single-end long-read transposon insertion identification method includes the following steps: Step S301: Perform haplotype-level variant identification on the alignment file, identify structural variants with an insertion length exceeding a set threshold as large-fragment insertion type structural variants, and extract the insertion sequence of the large-fragment insertion type structural variant from the alignment file; Step S302: Perform self-nucleic acid sequence alignment analysis on the inserted sequence to detect repetitive structures within the inserted sequence and verify the integrity of the terminal sequence of the inserted sequence; Step S303: Align the full-length transcriptome data to the inserted sequence, identify potential coding regions, analyze the structural domains of the inserted sequence based on the potential coding regions, and infer the functional domains of the inserted sequence through the structural domains; Step S304: Compare the extracted insertion sequence with the transposon database. Determine the transposon type of the insertion sequence based on the database comparison results, structural domain and functional domain information. If the insertion sequence is a known transposon, update the annotation information in the database. If it is an unknown transposon, define a new family and submit it to the database. The transposon insertion identification method with short readings at both ends Includes the following steps: Step S311: Perform haplotype-level variant identification on the alignment file, and screen for abnormal alignment sequences across the entire genome of the alignment file using local deep anomalies and abnormal alignment signals; Step S312: Align the abnormal alignment sequence with the target transposon sequence, and use the completely matching sites as candidate insertion sites; Step S313: The local depth anomaly and anomaly comparison signal of the candidate insertion site are experimentally verified by the three-primer method to confirm the authenticity of the insertion.
2. The method for haplotype horizontal transposon insertion identification based on physical typing according to claim 1, characterized in that: Step S3, which involves separating two sets of haplotype genome alignment files based on the physical typing results, includes the following steps: extracting sequence alignment records belonging to the two haplotype genomes from the alignment files according to the sequencing sequence number list in the typing results of the sequencing data, thereby forming two sets of typing result files.
3. A haplotype horizontal transposon insertion identification system based on physical typing, implemented based on the haplotype horizontal transposon insertion identification method based on physical typing as described in any one of claims 1 to 2, characterized in that, include: The module includes a data input module, a haplotype comparison module, a physical typing allocation module, a haplotype comparison file generation module, and a haplotype horizontal transposon insertion identification module. The data input module receives raw sequencing data, removes low-quality sequences, and distinguishes between single-end long reads and paired-end short reads. The haplotype alignment module: aligns the sequencing data to two haplotype genomes respectively, calculates the alignment score for each sequence, and outputs the original alignment files for the two haplotype genomes; The physical genotyping allocation module: assigns each sequencing sequence to a haplotype genome based on the alignment score, and outputs the genotyping results of the sequencing sequences; different allocation methods are used for single-end long reads and paired-end short reads: Assign the single-end long read to the haplotype with the higher alignment score; The comparison scores of the short reading lengths of the two ends are added together, and the short reading lengths of the two ends are assigned to the haplotypes with higher total comparison scores; The haplotype alignment file generation module extracts the sequences assigned to each haplotype from the original alignment file based on the genotyping results and generates haplotype-specific alignment files. The haplotype horizontal transposon insertion identification module employs different transposon insertion identification methods for the single-end long read and the double-end short read, detects transposon insertions in haplotype-specific alignment files, and determines the insertion type by combining structural variations and functional verification. The method for identifying transposon insertions in single-end sequencing sequences using the haplotype-level transposon insertion identification module includes the following steps: Step 1: Perform haplotype-level variant identification on the alignment file, identify structural variants with an insertion length exceeding a set threshold as large-fragment insertion type structural variants, and extract the insertion sequence of large-fragment insertion type structural variants from the alignment file; Step 2: Perform self-nucleic acid sequence alignment analysis on the inserted sequence to detect repetitive structures within the inserted sequence and verify the integrity of the terminal sequence of the inserted sequence; Step 3: Align the full-length transcriptome data to the inserted sequence, identify potential coding regions, analyze the structural domains of the inserted sequence based on the potential coding regions, and infer the functional domains of the inserted sequence through the structural domains; Step 4: Compare the extracted insertion sequence with the transposon database. Determine the transposon type of the insertion sequence based on the database comparison results, structural domain and functional domain information. If the insertion sequence is a known transposon, update the annotation information in the database. If it is an unknown transposon, define a new family and submit it to the database. The haplotype horizontal transposon insertion identification module includes the following steps in its method for identifying transposon insertions with short readings at both ends: Step 1: Perform haplotype-level variant identification on the alignment file, and screen for abnormal alignment sequences across the entire genome of the alignment file using local deep anomalies and abnormal alignment signals; Step 2: Align the abnormal alignment sequence with the target transposon sequence, and use the completely matching sites as candidate insertion sites; Step 3: The local depth anomaly and anomaly comparison signal of the candidate insertion site are experimentally verified by the three-primer method to confirm the authenticity of the insertion.
Citation Information
Patent Citations
Method for gene-level haploid typing based on PacBio HiFi three-generation DNA sequencing data
CN120015112A