A method for identifying analysis of herv-k insertion polymorphisms
By using the PacBio-HiFi sequencing platform and reference genome alignment technology, the problems of non-specific alignment and high false positive rate of short-read sequencing technology in detecting HERV-K (HML-2) sequences were solved, and accurate assessment and visualization verification of HERV-K (HML-2) site insertion polymorphism were achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN XINHUA HOSPITAL
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-04
AI Technical Summary
Existing short-read sequencing technologies suffer from problems such as non-specific alignment, low data utilization, high false positive rate, and difficulty in distinguishing insertion polymorphisms when detecting highly repetitive HERV-K (HML-2) sequences.
Long-read sequencing data were acquired using the PacBio-HiFi sequencing platform. Sequence alignment was performed using Host_Ref, LTR_Ref, and Pro_Ref reference genomes. Boolean logic alignment and visualization verification were performed using an automated decision matrix and IGV software to determine the insertion type and polymorphism status of the HERV-K (HML-2) site.
It enables automated assessment of insertion polymorphism at the HERV-K (HML-2) site, improves data utilization, reduces false positive rate, and enhances the accuracy of insertion type through visualization verification.
Smart Images

Figure CN122511366A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics technology, and in particular to a method for identifying and analyzing HERV-K insertion polymorphisms. Background Technology
[0002] Human endogenous retrovirus K, abbreviated as HERV-K, is a class of LTR retrotransposons widely distributed in the human genome, accounting for approximately 8% of the human genome. HERV-K (HML-2) is a repetitive element with over a thousand insertion sites in the human genome, and the high sequence repetition between different HERV-K (HML-2) sites makes precise localization or differentiation of HERV-K (HML-2) in the human genome particularly difficult. For example, the three HERV-K (HML-2) sequences, each 9180 bp in length, located on chromosomes 1q22, 3q27.2, and 5q33.3, share a similarity of up to 99.30%. Current techniques limit the analysis of HERV-K (HML-2). If only specific PCR amplification and Sanger sequencing are used to detect HERV-K (HML-2) sites, the throughput is far from sufficient to achieve a comprehensive understanding of HERV-K (HML-2). Therefore, with the development of next-generation sequencing technologies, such as Illumina PE150, mining information related to HERV-K (HML-2) from whole-genome sequencing data became the main technical method for researchers for a period of time, and various analysis software and workflows emerged, including ERVcaller, RetroSeq, Telescope, etc.
[0003] However, the short read length of whole-genome sequencing inevitably leads to the discarding of many repetitive sequences during alignment analysis due to non-specific alignment, resulting in low data utilization or a high false positive rate due to non-specific alignment. HERV-K (HML-2) loci are widely distributed in the human genome but are only associated with disease development in some individuals. HERV-K (HML-2) polymorphism can explain this phenomenon. However, research on HERV-K (HML-2) polymorphism must be based on a complete understanding of the distribution and sequence information of HERV-K (HML-2) loci in the human genome. Short-read sequencing data cannot achieve this goal and is not effective in assessing insertion polymorphism at HERV-K (HML-2) loci, especially in distinguishing between proviral sequence insertions and LTR sequence insertions. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a method for identifying and analyzing HERV-K insertion polymorphisms, which solves the problems of non-specific alignment, low data utilization, high false positive rate, and difficulty in distinguishing insertion polymorphisms in existing short-read sequencing technologies when detecting highly repetitive HERV-K (HML-2) sequences.
[0005] To achieve the above objectives, the present invention provides the following solution: A method for identifying HERV-K insertion polymorphisms includes: Long-read sequencing data were obtained using the PacBio-HiFi sequencing platform to obtain PacBio-HiFi sequences; the PacBio-HiFi sequences include either ccs.bam or fastq files. Set up a Host_Ref reference genome, an LTR_Ref reference genome, and a Pro_Ref reference genome; the Host_Ref reference genome includes: the host reference genome; the LTR_Ref reference genome includes: the LTR sequence of HML-2 type HERV-K; the Pro_Ref reference genome includes: the internal protein-coding sequence of HML-2 type HERV-K; the host reference genome is GRCh38; By integrating the coordinates of HML-2 HERV-K sites identified on the Host_Ref reference genome, a set of reference site coordinates is obtained; The PacBio-HiFi sequences are aligned to the Host_Ref reference genome, the LTR_Ref reference genome, and the Pro_Ref reference genome based on the reference site coordinate set, respectively, to obtain the total number of sequences aligned to each site, the number of LTR sequences, and the number of Pro sequences; Based on the total number of sequences, the number of LTR sequences, and the number of Pro sequences, the insertion type of each point is determined using a pre-set automated determination matrix to obtain preliminary determination results; The preliminary determination results are compared with the known states of the target site in the reference genome using Boolean logic to obtain the preliminary determination of the polymorphic state; The sites of the soft-clipped sequence were visualized using IGV software. The accuracy of the type of polymorphic insertion initially determined by the soft-clipped sequence was judged. Regions initially determined to be polymorphic without polymorphic insertion were supplemented to obtain the final polymorphic status of HML-2 HERV-K at the whole genome level.
[0006] Preferably, the PacBio-HiFi sequences are aligned to the Host_Ref reference genome, the LTR_Ref reference genome, and the Pro_Ref reference genome according to the reference site coordinate set, respectively, to obtain the total number of sequences aligned to each site, the number of LTR sequences, and the number of Pro sequences, including: The PacBio-HiFi sequences are respectively compared to the Host_Ref reference genome, and the number of sequences within 20 bp upstream and downstream of each site in the reference site coordinate set is extracted to obtain the total number of sequences; Extract the sequences from the PacBio-HiFi sequences that are aligned to each point in the reference site coordinate set within 3kb upstream and downstream to obtain the filtered sequences; The filtered sequences are aligned to the LTR_Ref reference genome, and the number of sequences aligned to each site within 20 bp upstream and downstream of each site in the reference site coordinate set is extracted to obtain the LTR sequence number; The filtered sequences are aligned to the Pro_Ref reference genome, and the number of sequences aligned to each site within 20 bp upstream and downstream of the reference site coordinate set is extracted to obtain the Pro sequence number.
[0007] Preferably, the determination rules of the automated determination matrix include: When the total number of sequences is 0, the preliminary determination result is that no site was detected; When the total number of sequences is not equal to 0, the number of LTR sequences is 0, and the number of Pro sequences is 0, the preliminary determination result is that there is no virus sequence insertion. When the total number of sequences is not equal to 0, the number of LTR sequences is not equal to 0, and the number of Pro sequences is 0, the preliminary determination result is determined to be an LTR sequence insertion; When the total number of sequences is not equal to 0, the number of LTR sequences is not equal to 0, and the number of Pro sequences is not equal to 0, the preliminary determination result is determined to be the insertion of the original virus sequence.
[0008] Preferably, the preliminary determination of the polymorphic state includes any one of the first to sixth states; The first state is a reference non-viral sequence insertion site where the known state is the insertion of a virus-free sequence and the preliminary determination result is the insertion of a virus-free sequence; the second state is a reference LTR sequence insertion site where the known state is the insertion of an LTR sequence and the preliminary determination result is the insertion of an LTR sequence; the third state is a reference protovirus sequence insertion site where the known state is the insertion of a protovirus sequence and the preliminary determination result is the insertion of a protovirus sequence; the fourth state is a polymorphic non-viral sequence insertion site where the known state is either the insertion of an LTR sequence or the insertion of a protovirus sequence and the preliminary determination result is the insertion of a virus-free sequence; the fifth state is a polymorphic LTR sequence insertion site where the known state is either the insertion of a virus-free sequence or the insertion of a protovirus sequence and the preliminary determination result is the insertion of an LTR sequence; the sixth state is a polymorphic protovirus sequence insertion site where the known state is either the insertion of a virus-free sequence or the insertion of an LTR sequence and the preliminary determination result is the insertion of a protovirus sequence.
[0009] Preferably, the sites of the soft-clipped sequence are visualized using IGV software, and the accuracy of the type of polymorphic insertion initially determined by the soft-clipped sequence is judged. Regions initially determined to be free of polymorphic insertions are supplemented to obtain the final polymorphic status of HML-2 HERV-K at the whole-genome level, including: Extract the soft-clipped sequence with a truncated signal and a sequence length of not less than 200 bp from the alignment file, and visualize the soft-clipped sequence in the region of the reference site coordinate set. The insertion type of the polymorphic site corresponding to the soft-clipped sequence was confirmed using IGV software to obtain the final polymorphic status of HML-2 HERV-K at the whole genome level; the final polymorphic status includes: the six states in the preliminary polymorphic status and any one of the seventh to ninth states. The seventh state is the final determination result of a polymorphic non-viral sequence insertion and a heterozygous LTR sequence insertion site for both the non-viral sequence insertion and the LTR sequence insertion; the eighth state is the final determination result of a polymorphic non-viral sequence insertion and a heterozygous LTR sequence insertion for both the non-viral sequence insertion and the original viral sequence insertion; the ninth state is the final determination result of a polymorphic LTR sequence insertion and a heterozygous LTR sequence insertion for both the LTR sequence insertion and the original viral sequence insertion.
[0010] The present invention discloses the following technical effects: This invention provides a method for identifying and analyzing HERV-K insertion polymorphism. By using sequence alignment, automated determination, and Boolean logic alignment, it solves the problems of non-specific alignment, low data utilization, high false positive rate, and difficulty in distinguishing insertion polymorphisms in existing short-read sequencing technologies when detecting highly repetitive HERV-K (HML-2) sequences. It achieves automated assessment of HML-2 type HERV-K site insertion polymorphism. By applying the soft-clipped sequence in the sequencing data alignment results to verify the site insertion type, it solves the problem of low verification accuracy of short-read sequencing sequences and achieves single-molecule level visual verification. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of the identification and analysis process for HERV-K insertion polymorphism provided in an embodiment of the present invention; Figure 2 An automated determination flowchart provided for embodiments of the present invention; Figure 3 A preliminary determination flowchart provided for embodiments of the present invention; Figure 4 A flowchart of the final determination process provided in this embodiment of the invention; Figure 5 The IGV image of the sample HM51 site coordinates chr12:58327458-58336915 provided in the embodiment of the present invention was finally determined to have the polymorphism of Poly_LTR; Figure 6 The IGV image of sample PB09 with coordinates chr12:58327458-58336915 provided in this embodiment of the invention was ultimately determined to have the polymorphism Ref_Pro. Figure 7 The IGV image of the sample PB17 site provided in this embodiment of the invention, with coordinates chr12:58327458-58336915, was ultimately determined to have the polymorphism Poly_LTR&Pro. Figure 8 The IGV image of the sample HM14 site provided in this embodiment of the invention has the final polymorphism determined to be Poly_Pre. (chr3:186865587-186866486) Figure 9The IGV image with the final polymorphism of Poly_Pre<R was determined to be the PB02 site coordinates chr3:186865587-186866486 provided in the embodiments of the present invention. Figure 10 The IGV image with the PB18 site coordinates chr3:186865587-186866486 provided in the embodiment of the present invention was finally determined to have the polymorphism Ref_LTR. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] The purpose of this invention is to provide a method for identifying and analyzing HERV-K insertion polymorphisms, which solves the problems of non-specific alignment, low data utilization, high false positive rate, and difficulty in distinguishing insertion polymorphisms in existing short-read sequencing technologies when detecting highly repetitive HERV-K (HML-2) sequences.
[0015] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0016] Figure 1 This is a schematic diagram of the HERV-K insertion polymorphism identification and analysis process provided in an embodiment of the present invention, as shown below. Figure 1 As shown, this invention provides a method for identifying and analyzing HERV-K insertion polymorphisms, including: Step 100: Obtain long-read sequencing data using the PacBio-HiFi sequencing platform to obtain PacBio-HiFi sequences; the PacBio-HiFi sequences include: ccs.bam files or fastq files; Step 200: Set the Host_Ref reference genome, LTR_Ref reference genome, and Pro_Ref reference genome; the Host_Ref reference genome includes: the host reference genome; the LTR_Ref reference genome includes: the LTR sequence of HML-2 type HERV-K; the Pro_Ref reference genome includes: the internal protein-coding sequence of HML-2 type HERV-K; the host reference genome is GRCh38; Step 300: Integrate the coordinates of the HML-2 HERV-K sites identified on the Host_Ref reference genome to obtain the reference site coordinate set; Step 400: Based on the reference site coordinate set, the PacBio-HiFi sequences are aligned to the Host_Ref reference genome, the LTR_Ref reference genome, and the Pro_Ref reference genome, respectively, to obtain the total number of sequences aligned to each site, the number of LTR sequences, and the number of Pro sequences; Step 500: Based on the total number of sequences, the number of LTR sequences, and the number of Pro sequences, the insertion type of each point is determined using a pre-set automated determination matrix to obtain a preliminary determination result; Step 600: Perform Boolean logic comparison between the preliminary determination result and the known state of the target site in the reference genome to obtain the preliminary determination of polymorphism state; Step 700: Visualize the sites of the soft-clipped sequence using IGV software, use the soft-clipped sequence to accurately determine the type of polymorphic insertion in the preliminary determination of the polymorphic state, and supplement the regions in the preliminary determination of the polymorphic state that do not have polymorphic insertions, to obtain the final determination of the polymorphic state of HML-2 type HERV-K at the whole genome level.
[0017] Specifically, this embodiment proposes an analytical method for identifying HERV-K (HML-2) insertion polymorphisms using PacBio-HiFi sequencing data. The specific steps are as follows: S1: Data Acquisition: Acquire long-read sequencing data based on the PacBio-HiFi sequencing platform.
[0018] S2: Reference Sequence Preparation: This embodiment mainly involves three different reference genomes, including: 1) Host_Ref: Host reference genome (GRCh38). 2) LTR_Ref: Reference sequence containing the HERV-K (HML-2) LTR sequence. 3) Pro_Ref: Reference sequence containing only the protein-coding sequence within HERV-K (HML-2).
[0019] S3: Reference site coordinate preparation: The coordinates of 1037 known HERV-K (HML-2) sites on the reference genome were obtained by searching the literature and compiled into the ref_1037.bed file for subsequent site detection.
[0020] S4: Sequence Alignment: Align the obtained PacBio-HiFi sequences to Host_Ref, and obtain the number of sequences aligned to within 20 bp upstream and downstream of each locus in ref_1037.bed. This number is defined as the total number of sequences that can be aligned to each locus (Total). Specifically, use pbmm2 software to align the obtained PacBio-HiFi sequencing sequence ccs.bam or fastq files to Host_Ref. To ensure specific alignment and prevent multiple alignments from interfering with sequence counts, set the parameters --min-gap-comp-id-perc to 99.0 and -N to 1. Use bedtools software to obtain the number of sequences aligned to within 20 bp upstream and downstream of each locus in ref_1037.bed. This number is defined as the total number of sequences that can be aligned to each locus (Total).
[0021] Considering the large amount of PacBio-HiFi data, in order to improve the efficiency of subsequent analysis, sequences 3kb upstream and downstream of the HERV-K (HML-2) site were extracted for subsequent alignment.
[0022] The filtered sequences are aligned to LTR_Ref, and the number of sequences aligned to each point in ref_1037.bed within 20bp upstream and downstream is obtained using bedtools software. This number is defined as the LTR sequence number that can be aligned to each point (LTR).
[0023] The filtered sequences are aligned to Pro_Ref. The number of sequences aligned to each point in ref_1037.bed within 20bp upstream and downstream is obtained using bedtools software. This number is defined as the number of Pro sequences that can be aligned to each point (Pro).
[0024] S5: Reference Figure 2 Preliminary determination of HERV-K (HML-2) site insertion type: Extract the total number of sequences (Total), LTR sequence number (LTR), and Pro sequence number (Pro) at each site from the three alignment steps mentioned above, and establish an automated determination matrix, as shown in Table 1: Table 1
[0025] 1) If Total ≠ 0, LTR = 0 and Pro = 0, it is determined that there is no virus sequence insertion (Pre-integration, Pre).
[0026] 2) If Total≠0, LTR≠0 and Pro=0, it is determined to be an LTR sequence insertion (LTR integration, LTR).
[0027] 3) If Total≠0, LTR≠0 and Pro≠0, it is determined to be a provirus integration (Pro).
[0028] 4) Reference Figure 3 The above results are compared with the known state (ref_type) of the locus in the reference genome using Boolean logic, and the polymorphism state of the locus is automatically output. Preliminary detectable HERV-K (HML-2) loci are divided into 6 categories (see Table 2), including: ① Ref_Pre loci that are Pre in the reference genome and detected as Pre in the sample; ② Ref_LTR loci that are LTR in the reference genome and detected as LTR in the sample; ③ Ref_Pro loci that are Pro in the reference genome and detected as Pro in the sample; ④ Poly_Pre loci that are not Pre in the reference genome but detected as Pre in the sample; ⑤ Poly_LTR loci that are not LTR in the reference genome but detected as LTR in the sample; ⑥ Poly_Pro loci that are not Pro in the reference genome but detected as Pro in the sample.
[0029] Table 2
[0030] S6: Reference Figure 4 Visual confirmation of HERV-K (HML-2) site insertion type: Soft-clipped sequences with truncation signals and a sequence length of not less than 200 bp are automatically extracted from the alignment file (SAM). For the soft-clipped sequences, IGV software is used to visualize the corresponding HERV-K (HML-2) sites to confirm the polymorphic site insertion type, obtaining the final determination of the polymorphism status of HERV-K (HML-2) sites at the whole genome level. In addition to the 6 initially detectable sites, the following 3 categories may be added (refer to Table 3), including: ⑦ Poly_Pre<R sites where both Pre and LTR are detected in the sample; ⑧ Poly_Pre&Pro sites where both Pre and Pro are detected in the sample; ⑨ Poly_LTR&Pro sites where both LTR and Pro are detected in the sample.
[0031] Table 3
[0032] Wherein, Ref_Pre is the reference non-viral sequence insertion site; Ref_LTR is the reference LTR sequence insertion site; Ref_Pro is the reference proviral sequence insertion site; Poly_Pre is the polymorphic non-viral sequence insertion site; Poly_LTR is the polymorphic LTR sequence insertion site; Poly_Pro is the polymorphic proviral sequence insertion site; Poly_Pre<R is the polymorphic non-viral sequence and LTR sequence heterozygous insertion site; Poly_Pre&Pro is the polymorphic non-viral sequence and proviral sequence heterozygous insertion site; and Poly_LTR&Pro is the polymorphic LTR sequence and proviral sequence heterozygous insertion site.
[0033] refer to Figures 5 to 10 The main functions of IGV visualization can be divided into three categories: 1) Supplementing missed heterozygous polymorphic sites, especially LTR and Pro insertion states that are difficult to distinguish in short-read sequencing sequences. For example, the site chr12:58327458-58336915 has an insertion type of Pro in the reference coordinate set. If Total≠0, LTR≠0, and Pro≠0 are detected in the sample, then automatically, this site is determined to have an insertion state of Ref_Pro. However, visualization can reveal that although Pro≠0 at this site, there may be a heterozygous LTR insertion masked by a Pro insertion. For example... Figures 5 to 7 As shown, the site chr12:58327458-58336915 was initially determined to be Ref_Pro in PB09 and PB17, and to Poly_LTR in HM51. After IGV visualization verification, the final determination of the polymorphic state in PB17 can be corrected to Poly_LTR&Pro.
[0034] 2) To supplement missed detections of polymorphic sites due to large sequence fragment deletions. For example... Figures 8 to 10 As shown, the locus chr3:186865587-186866486 has an insertion type of LTR in the reference coordinate set. In HM14, since Total=0 was obtained by comparing Host_Ref, it was automatically determined that this locus was not detected in sample HM14 and did not proceed to subsequent polymorphism state assessment. In sample PB02, the initial polymorphism state was determined to be Ref_LTR. After IGV visualization verification, it was found that this locus could be located within the 4.3kb deletion region in both samples PB02 and HM14. After correction, the final polymorphism state in HM14 was determined to be Poly_Pre, and the final polymorphism state in PB17 was determined to be Poly_Pre<R.
[0035] 3) Eliminating false positives. Theoretically, all sites with polymorphic insertion states should have corresponding soft-clipped sequences. If a site is initially identified as a polymorphic site but no corresponding soft-clipped sequence is extracted, it is likely a false positive. Such sites can be easily identified through IGV visualization. The reason for this is that PacBio-HiFi sequencing sequences can reach tens of kb in length, and in densely distributed HERV-K (HML-2) sites, such sequencing sequences may cover several HERV-K (HML-2) sites. For example, a 20kb sequencing sequence covers two HERV-K (HML-2) sites, with an LTR insertion type at site A and a Pre insertion type at site B. However, because this sequence contains an LTR sequence at site A, site B will be identified as an LTR insertion site during counting, thus incorrectly evaluating site B as a Poly_LTR. Visualization of the site B region can easily correct this to a Pre site.
[0036] The beneficial effects of this invention are as follows: (1) This invention sets up a whole genome reference sequence (Host_ref.fa), a reference sequence containing HERV-K (HML-2) LTR sequence (LTR_ref.fa), and a reference sequence containing only HERV-K (HML-2) provirus sequence (Pro_ref.fa), and aligns the sequencing sequence to three reference genomes respectively. The insertion type of each HERV-K (HML-2) site can be obtained by calculation alone, which is simpler and more efficient.
[0037] (2) This invention is the first to apply the soft-clipped sequence from the alignment results of PacBio-HiFi sequencing data to the visualization verification of the insertion type of HERV-K (HML-2) sites. Due to the long read length and high accuracy of PacBio-HiFi sequencing sequences, a single soft-clipped sequence can even cover the flanking sequences of the target site and the provirus sequence, thereby achieving visualization verification at the single-molecule level.
[0038] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0039] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for identifying and analyzing HERV-K insertion polymorphism, characterized in that, include: Long-read sequencing data were obtained using the PacBio-HiFi sequencing platform to obtain PacBio-HiFi sequences; The PacBio-HiFi sequence includes either a ccs.bam file or a fastq file; Set up a Host_Ref reference genome, an LTR_Ref reference genome, and a Pro_Ref reference genome; the Host_Ref reference genome includes: the host reference genome; the LTR_Ref reference genome includes: the LTR sequence of HML-2 type HERV-K; the Pro_Ref reference genome includes: the internal protein-coding sequence of HML-2 type HERV-K; the host reference genome is GRCh38; By integrating the coordinates of HML-2 HERV-K sites identified on the Host_Ref reference genome, a set of reference site coordinates is obtained; The PacBio-HiFi sequences are aligned to the Host_Ref reference genome, the LTR_Ref reference genome, and the Pro_Ref reference genome based on the reference site coordinate set, respectively, to obtain the total number of sequences aligned to each site, the number of LTR sequences, and the number of Pro sequences; Based on the total number of sequences, the number of LTR sequences, and the number of Pro sequences, the insertion type of each point is determined using a pre-set automated determination matrix to obtain preliminary determination results; The preliminary determination results are compared with the known states of the target site in the reference genome using Boolean logic to obtain the preliminary determination of the polymorphic state; The sites of the soft-clipped sequence were visualized using IGV software. The accuracy of the type of polymorphic insertion initially determined by the soft-clipped sequence was judged. Regions initially determined to be polymorphic without polymorphic insertion were supplemented to obtain the final polymorphic status of HML-2 HERV-K at the whole genome level.
2. The method for identifying and analyzing HERV-K insertion polymorphism according to claim 1, characterized in that, Based on the reference site coordinate set, the PacBio-HiFi sequences are aligned to the Host_Ref reference genome, the LTR_Ref reference genome, and the Pro_Ref reference genome, respectively, to obtain the total number of sequences aligned to each site, the number of LTR sequences, and the number of Pro sequences, including: The PacBio-HiFi sequences are respectively compared to the Host_Ref reference genome, and the number of sequences within 20 bp upstream and downstream of each site in the reference site coordinate set is extracted to obtain the total number of sequences; Extract the sequences from the PacBio-HiFi sequences that are aligned to each point in the reference site coordinate set within 3kb upstream and downstream to obtain the filtered sequences; The filtered sequences are aligned to the LTR_Ref reference genome, and the number of sequences aligned to each site within 20 bp upstream and downstream of each site in the reference site coordinate set is extracted to obtain the LTR sequence number; The filtered sequences are aligned to the Pro_Ref reference genome, and the number of sequences aligned to each site within 20 bp upstream and downstream of the reference site coordinate set is extracted to obtain the Pro sequence number.
3. The method for identifying and analyzing HERV-K insertion polymorphism according to claim 1, characterized in that, The determination rules of the automated determination matrix include: When the total number of sequences is 0, the preliminary determination result is that no site was detected; When the total number of sequences is not equal to 0, the number of LTR sequences is 0, and the number of Pro sequences is 0, the preliminary determination result is that there is no virus sequence insertion. When the total number of sequences is not equal to 0, the number of LTR sequences is not equal to 0, and the number of Pro sequences is 0, the preliminary determination result is determined to be an LTR sequence insertion; When the total number of sequences is not equal to 0, the number of LTR sequences is not equal to 0, and the number of Pro sequences is not equal to 0, the preliminary determination result is determined to be the insertion of the original virus sequence.
4. The method for identifying and analyzing HERV-K insertion polymorphism according to claim 3, characterized in that, The preliminary determination of the polymorphic state includes any one of the first to sixth states; The first state is a reference non-viral sequence insertion site where the known state is the insertion of a virus-free sequence and the preliminary determination result is the insertion of a virus-free sequence; the second state is a reference LTR sequence insertion site where the known state is the insertion of an LTR sequence and the preliminary determination result is the insertion of an LTR sequence; the third state is a reference protovirus sequence insertion site where the known state is the insertion of a protovirus sequence and the preliminary determination result is the insertion of a protovirus sequence; the fourth state is a polymorphic non-viral sequence insertion site where the known state is either the insertion of an LTR sequence or the insertion of a protovirus sequence and the preliminary determination result is the insertion of a virus-free sequence; the fifth state is a polymorphic LTR sequence insertion site where the known state is either the insertion of a virus-free sequence or the insertion of a protovirus sequence and the preliminary determination result is the insertion of an LTR sequence; the sixth state is a polymorphic protovirus sequence insertion site where the known state is either the insertion of a virus-free sequence or the insertion of an LTR sequence and the preliminary determination result is the insertion of a protovirus sequence.
5. The method for identifying and analyzing HERV-K insertion polymorphism according to claim 4, characterized in that, The sites of the soft-clipped sequence were visualized using IGV software. The accuracy of the type of polymorphic insertion initially determined by the soft-clipped sequence was assessed. Regions initially determined to lack polymorphic insertions were supplemented to obtain the final polymorphic status of HML-2 HERV-K at the whole-genome level, including: Extract the soft-clipped sequence with a truncated signal and a sequence length of not less than 200 bp from the alignment file, and visualize the soft-clipped sequence in the region of the reference site coordinate set. The insertion type of the polymorphic site corresponding to the soft-clipped sequence was confirmed using IGV software to obtain the final polymorphic status of HML-2 HERV-K at the whole genome level; the final polymorphic status includes: the six states in the preliminary polymorphic status and any one of the seventh to ninth states. The seventh state is the final determination result of a polymorphic non-viral sequence insertion and a heterozygous LTR sequence insertion site for both the non-viral sequence insertion and the LTR sequence insertion; the eighth state is the final determination result of a polymorphic non-viral sequence insertion and a heterozygous LTR sequence insertion for both the non-viral sequence insertion and the original viral sequence insertion; the ninth state is the final determination result of a polymorphic LTR sequence insertion and a heterozygous LTR sequence insertion for both the LTR sequence insertion and the original viral sequence insertion.