Lentivirus integration junction analysis
The method addresses the need for rapid and accurate lentiviral integration analysis by shearing genomic DNA, using inverse PCR and rolling circle amplification, to determine precise integration junctions and copy number in the human genome.
Patent Information
- Application Number
- JP2025030146
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-29
- Filing Date
- 2025-02-27
- Publication Date
- 2025-09-10
AI Technical Summary
Current methods for evaluating the quality of lentiviral delivery and infection in gene therapy lack rapid and accurate analysis of proviral location and integration in the host human genome.
A method involving shearing of genomic DNA, followed by inverse PCR and rolling circle amplification using primers derived from the lentiviral LTR region, and bioinformatic analysis to determine precise integration junctions and copy number.
Enables high-resolution analysis of lentiviral integration in the human genome, providing precise genomic location and optional copy number information.
Smart Images

Figure 2025133083000001_ABST
Abstract
Description
[Technical Field]
[0001] Lentiviral vector delivery has advantages over other gene therapy methods due to its high efficiency of infection of both dividing and non-dividing cells, long-term stable expression of transgenes, and low immunogenicity. Lentiviruses have been used to induce immune responses against tumor antigens. However, as with most current gene therapy experiments, rapid and accurate methods for evaluating the quality of viral delivery and infection are required.
[0002] Rapid and efficient analysis of lentiviral proviral location and integration in the host human genome is essential for assessing the quality of viral transduction used in gene therapy.
[0003] This invention describes a method for assessing the location and optional copy number of lentiviral integration in the host human genome based on shearing of extracted genomic DNA followed by sequencing by inverse PCR and rolling circle amplification using a primer set derived from the lentiviral LTR (long terminal repeat) region. The sequencing results are then analyzed by a customized bioinformatic algorithm to locate the precise integration junction between the lentiviral provirus and the host genome.
[0004] Summary of the Invention The object of the present invention is to provide a method for obtaining the precise genomic integration location and, optionally, copy number information of the transduced and integrated lentivirus (provirus) with higher resolution than known techniques.
[0005] The method involves PCR preamplification of a fragment of genomic DNA containing the junction sequences of the 5' and 3' LTR regions and host genomic DNA, and rolling circle amplification of the PCR product.
[0006] Means to solve the problem The object of the present invention is a method for obtaining the genomic location of a proviral sequence embedded in a DNA strand by two LTR regions, comprising the following steps: a. Fragmenting the DNA strand into multiple strands having a length of 50-1000 bp, thereby obtaining a mixture of strands containing at least one LTR region of the proviral sequence and a DNA strand junction region having a length of 20-100 bp, and strands that do not contain the LTR region; b. converting the strand into a circularized strand; c. providing PCR primers P1 and P2 at the 3' and 5' ends of at least one LTR region within the circularized strand; d. Doubling the circularized strand provided with PCR primers by inverse PCR, thereby obtaining multiple linear strands in which at least one LTR region and junction region are filled in by PCR primers; e. converting the linear strand with the embedded LTR and junction region into a circularized strand and amplifying the circularized strand ronically by RCA; f. Obtaining sequence information of the rolonyi gene, thereby obtaining sequence information of the junction region; g. aligning the sequence information of the junction region with the sequence information of the DNA strand, thereby obtaining the genomic location of the proviral sequence; A method characterized by
[0007] The method according to the invention is particularly suitable for obtaining the genomic location of proviral sequences within DNA strands that are human genomic sequences of human chromosomes.
[0008] Proviral sequences can be derived from any type of virus, such as a lentivirus. The terms "proviral sequence" and "lentiviral sequence" are used interchangeably.
[0009] In addition to obtaining the genomic location of the proviral sequence, the copy number of the proviral sequence within the DNA strand can be obtained.
[0010] For this purpose, the sequence information of the DNA strand should be obtained either before or after the method.Since the human genome sequence is known, this information is accessible to those skilled in the art.However, the sequence information of the DNA strand can also be obtained by standard DNA sequencing methods. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 illustrates the overall workflow of the method of the present invention. [Figure 2] A diagram showing the circularized strand with PCR primers P1 and P2 provided in at least one LTR region (step c). [Figure 3] FIG. 1 shows the PCR-doubled linear strand obtained in step d). [Figure 4] FIG. 4 shows the exact DNA sequences of the PCR primer constructs used in the inverse PCR reactions described in FIGS. 2 and 3. [Figure 5] FIG. 1 shows the bioinformatic workflow (step g) for aligning junction region information with the sequences of the DNA strands. [Figure 6] FIG. 1 shows the bioinformatic workflow (step g) for aligning junction region information with the sequences of the DNA strands. [Figure 7] Figure 1 shows the location of lentiviral insertion sites along the lentiviral genome. All reads map to the 5' and 3' terminal regions of the LTR region, as expected in a lentiviral integration sample. [Figure 8] FIG. 1 shows the distribution of unique insertion sites along the host cell genome in test experiments performed with lentiviral integration samples (GFP+). [Figure 9]Figure 1 shows an overview of proof-of-concept experiments from two different methods (long primer blunt end ligase and short primer Circligase) for positive control GFP+ and negative control GFP- samples. IS merged count is the total number of samples found to be integrated within a particular gene. [Figure 10] FIG. 1 shows detailed results (from four individual sequencing datasets) from two different methods (long primer blunt end ligase and short primer Circligase) against the positive control GFP+. [Figure 11] FIG. 1 shows a method for determining proviral (integrated lentiviral) copy number per cell.
[0012] MODE FOR CARRYING OUT THE INVENTION The overall method of the present invention is shown in Figure 1, which is now described in detail. The method is used in lentivirally transduced cells. Such transduction processes are known to those skilled in the art.
[0013] First, genomic DNA is extracted from the transduced cells (Figure 1A).
[0014] Next, the genomic DNA is preferably mechanically or enzymatically sheared to an average size of 100-300 base pairs (bp), such as 200 bp (Figure 1C).
[0015] In a first embodiment of the present invention, the strands are denatured into single-stranded DNA strands, which are then converted into circularized single-stranded DNA, which can be achieved by mechanical denaturation of the denatured single-strands, cold shock, and single-stranded DNA ligation (e.g., with Circligase).
[0016] In a different approach, the sheared genomic DNA is end-polished and the blunt ends are ligated with T4 DNA ligase (FIG. 1C). Thus, in a second embodiment of the present invention, after step a), the strands are provided with blunt ends by ligase to obtain a plurality of double-stranded DNA strands, which are then converted into a plurality of circularized double-stranded DNA strands, preferably by T4 ligase.
[0017] Next, the circularized genomic DNA containing the LTR sequence is amplified by polymerase chain reaction using inverse PCR primers with adapters bearing fixed DNA sequences P1 and P2 that bind to the LTR region of the lentivirus (Figure 1D and Figure 1E).
[0018] Two sets of PCR primers may be used. Because the 5'LTR and 3'LTR do not coexist in the same sheared DNA molecule, one set is specific for the 5'LTR and the other set is specific for the 3'LTR. The orientation of the PCR primers relative to the 5'LTR region is such that primer P1 faces toward the 5' end of the 5'LTR (Figure 2A).
[0019] In a third embodiment of the method, the circularized strands provided with PCR primers comprise a mixture of a first circularized strand having a junction region 3' to the LTR region and a second circularized strand having a junction region 5' to the LTR region, and the first and second circularized strands are independently doubled by inverse PCR, thereby obtaining a linear strand having a junction region 3' to the LTR region or a linear strand having a junction region 5' to the LTR region.
[0020] In a fourth embodiment of the method, PCR primers P1 and / or P2 are provided at the 3' and / or 5' ends of at least one LTR region, respectively, in the circularized strand, such that PCR primer P1 faces towards the 5' end of the 5' LTR or 3' LTR region, and PCR primer P2 faces towards the 3' end of the 5' LTR or 3' LTR region.
[0021] In this embodiment, the orientation of the PCR primers relative to the 3' LTR region is such that primer P1 faces toward the 3' end of the 3' LTR (Figure 2B). In this way, both genomic insertion junctions of the provirus can be simultaneously mapped. The orientation of the 5' LTR PCR product (Figure 3A) and the 3' LTR PCR product (Figure 3B) are shown.
[0022] Preferably, step e) is carried out by ligation with T4 DNA ligase using a DNA splint bridging oligonucleotide that brings the P1 and P2 ends together (Figure 1F).
[0023] The circularized strand is then amplified by rolling circle amplification (RCA) (Figure 1G) into so-called rolony, which contains multiple linear concatemers of the circularized strand.
[0024] The resulting RCA product (ROL) is then NGS sequenced using a 5' LTR junction sequencing primer that binds to the 5' end of the LTR and 8 bases upstream of the genomic junction, and a 3' LTR junction sequencing primer that binds to the 3' end of the LTR and 8 bases upstream of the genomic junction, so that it is possible to distinguish between the two ends (Figure 4). NGS sequencing of LTR and genomic DNA junctions can be analyzed using the analytical workflow described below.
[0025] Step g) of the method of the present invention can be performed by mapping the sequence information of the junction region to the sequence information of the DNA strand, and the region of the DNA strand that aligns with high confidence (MAPQ ≥ 30) to the sequence information of the junction region is designated as the genomic location of the proviral sequence. The alignment can include merging of overlapping regions and annotation by the crossover gene or the nearest gene if no crossover gene is identified.
[0026] Example This example describes a method for assessing the location and copy number of lentiviral integration in the host human genome according to the present invention.
[0027] The method involves PCR preamplification of a fragment of genomic DNA containing the junction sequences of the 5' and 3' LTR regions and host genomic DNA, and rolling circle amplification of the PCR product.
[0028] Transduction / doubling step Genomic DNA was extracted from lentiviral vector-transduced SUP-T1 cells (a human T lymphoblastoid cell line from ATCC) and used as a negative control against genomic DNA extracted from non-transduced cells. As an example, a lentiviral vector construct derived from HIV, a highly efficient vehicle for in vivo gene delivery, was used. This construct carries green fluorescent protein (GFP), which exhibits green fluorescence when exposed to light in the blue-to-ultraviolet range, serving as a marker for positive transduction and integration of the provirus into the host genome.
[0029] First, genomic DNA was extracted from lentiviral-transduced (GFP+ plus) and non-transduced (GFP- minus) SUP-T1 cells (Figure 1A).
[0030] Next, the genomic DNA samples were mechanically sheared by sonication (Covaris Ultrasonicator) with the desired size set to 200 base pairs (Figure 1C). The concentration of sheared DNA was 200 nanograms in a 130 μL volume of 1X Tris / EDTA (1.54 ng / μL).
[0031] In one method, 50 nanograms of sheared genomic DNA was then denatured at 95°C for 10 minutes, cold-shocked at 4°C, and placed on ice. Single-strand circle ligation was performed with Circligase at 60°C for 60 minutes and inactivated at 80°C for 10 minutes. The ligation reaction was cleaned up to remove non-circularized molecules by incubation with exonucleases I and III at 37°C for 60 minutes and purified with Spry Beads or a size-exclusion column.
[0032] Alternatively, the fragmented genomic DNA was end-polished and the blunt ends were ligated with T4 DNA ligase (Figure 1C).
[0033] Next, the circularized genomic DNA containing the LTR sequence is amplified by polymerase chain reaction using inverse PCR primers with adapters bearing fixed DNA sequences P1 and P2 that bind to the LTR region of the lentivirus (Figure 1D and Figure 1E).
[0034] Two sets of PCR primers are used. One set is specific for the 5'LTR and the other set is specific for the 3'LTR, because the 5'LTR and 3'LTR do not coexist in the same 200 base pair sheared DNA molecule. The orientation of the PCR primers relative to the 5'LTR region is such that primer P1 faces toward the 5' end of the 5'LTR (Figure 2A).
[0035] In contrast, the orientation of the PCR primers for the 3' LTR region is such that primer P1 faces toward the 3' end of the 3' LTR (Figure 2B). In this way, both proviral junctions can be mapped simultaneously. The orientation of the 5' LTR PCR product (Figure 3A) and the 3' LTR PCR product (Figure 3B) are shown.
[0036] The PCR amplification product can then be circularized using a splint bridge oligonucleotide primer that brings the P1 and P2 ends together (Figure 1F).
[0037] The circles are then amplified by rolling circle amplification (RCA) (Figure 1G).
[0038] The resulting RCA products were then NGS sequenced using a 5' LTR junction sequencing primer that binds to the 5' end of the LTR and 8 bases upstream of the genome junction, and a 3' LTR junction sequencing primer that binds to the 3' end of the LTR and 8 bases upstream of the genome junction, allowing for differentiation between the 3' and 5' ends (Figure 4).
[0039] The genomic DNA / LTR junction can then be sequenced and the junction region can then be defined relative to a genomic DNA reference sequence.
[0040] Run the bioinformatic insertion site analysis workflow to determine the integration site (coordinates on the chromosome). Figures 5 and 6 show the bioinformatic workflow for analyzing integration site (IS) sequencing results.
[0041] Perform read pre-processing by removing PCR products and potential mispriming of adapters. Trim low-quality bases and ensure a minimum length of 30 base pairs.
[0042] The sequence is then mapped to a human reference sequence to identify the IS.
[0043] To validate the IS analysis process, sequences are also mapped to the lentiviral vector sequence.
[0044] Next, the overlapping ISs are merged and the ISs are annotated.
[0045] Finally, generate and visualize the report.
[0046] Unique ISs are identified with high confidence MAP Q≧30 (mismapping probability≦0.001%).
[0047] Ambiguous ISs are also identified (reads that map equally well to multiple locations).
[0048] Figure 6 shows the bioinformatic workflow for analyzing integration site (IS) sequencing results.
[0049] Figures 7-10 show sample results generated from proof-of-concept experiments performed with lentiviral integrated samples (GFP+ plus) and negative controls (GFP- minus).
[0050] Figure 7 shows that all lentiviral insertion site mappings resulted in reads that mapped to the 5' and 3' end regions of the LTR region, as expected for lentiviral integration samples (GFP+). The number of reads at the 3' and 5' ends was very close to each other (987,736 vs. 987,859), indicating that the method performed equally well at both the 3' and 5' ends. The fact that reads mapped only to the LTR and adjacent regions, but not to internal viral regions, suggests that the method is performing well.
[0051] Figure 8 shows the distribution of unique insertion sites along the host cell genome in lentiviral integration samples (GFP+). Both the long primer blunt-end ligation method and the short primer circular ligase method showed a wide distribution along the host chromosome in different GFP+ samples.
[0052] Figure 9 shows a summary of proof-of-concept experiments from two different methods (long primer blunt-end ligase and short primer circular ligase) for positive control GFP+ and negative control GFP- samples. IS merge count is the total number of samples found to be integrated within a specific gene. IS crossover count is the total number of samples found to cross over a specific gene. Unique crossover count is the total number of unique genes crossing over the viral insertion site. IS outside count is the total number of samples found outside any gene. Unique nearest neighbor count is the total number of unique genes close to the "IS outside gene." Total unique gene count is the total number of genes crossing or nearest to the IS (viral insertion site).
[0053] Figure 10 shows detailed results (from four individual sequencing datasets) from two different methods (long primer blunt-end ligase and short primer circular ligase) against the positive control GFP+. This shows which chromosome the lentivirus integrated into (chromosome). This shows the exact location on the gene where the sequencing reads were mapped (IS_chromosome start and IS_chromosome end). It also shows the distance between the integration site and the specific gene.
[0054] Figure 11 shows a method for determining proviral copy number. To determine the relative copy number of LTR-proviral copies, a single-copy internal control gene, such as CD3, is used as a reference. First, an RCA is generated from LTR-specific iPCR and single-copy gene-specific (e.g., CD3) RCA. Second, NGS sequencing is performed on both the LTR region and the CD3 gene to determine the number of reads. Since the single-copy gene (CD3) has two copies per cell (diploid), the number of fully primed cells is divided by 2 (e.g., 10,000 reads divided by 2 = 5,000 copies). The number of reads from the LTR is divided by the single-copy gene CD3 (15,000 / 5,000 = 3). There are three copies of provirus per cell. Alternatively, quantitative PCR can be used to determine the copy number for the reference gene and the LTR region of the provirus.
Claims
1. A method for obtaining the genomic location of a proviral sequence embedded by two LTR regions in a DNA strand, comprising the steps of: a. Fragmenting the DNA strand into a plurality of strands having a length of 50-1000 bp, thereby obtaining a mixture of strands containing at least one LTR region of a proviral sequence and a junction region of the DNA strands having a length of 20-100 bp, and strands not containing the LTR region; b. converting said strand into a circularized strand; c. providing PCR primers P1 and P2 at the 3' and 5' ends of said at least one LTR region within said circularized strand; d. Doubling the circularized strand provided with the PCR primers by inverse PCR, thereby obtaining a plurality of linear strands in which the at least one LTR region and the junction region are filled in by the PCR primers; e. Converting the linear strand with embedded LTR and junction regions into a circularized strand and amplifying the circularized strand ronically by RCA; f. Obtaining sequence information of the Rolony, thereby obtaining sequence information of the junction region; g. aligning the sequence information of the junction region with the sequence information of the DNA strands, thereby obtaining the genomic location of the proviral sequence; A method characterized by
2. 2. The method of claim 1, wherein after step a), the plurality of strands is denatured into a plurality of single-stranded DNA strands, which are then converted into a plurality of circularized single-stranded DNA strands.
3. 3. The method of claim 1 or 2, characterized in that after step a), the strands are provided with blunt ends by a ligase to obtain a plurality of double-stranded DNA strands, which are then converted into a plurality of circularized double-stranded DNA strands.
4. 4. The method according to claim 1, wherein step e) is carried out by ligation with T4 DNA ligase using a DNA splint bridging oligonucleotide that brings together the P1 and P2 ends.
5. 5. The method of claim 1, wherein the circularized strands provided with the PCR primers comprise a mixture of a first circularized strand having a junction region in the 3' direction of the LTR region and a second circularized strand having a junction region in the 5' direction of the LTR region, and the first and second circularized strands are independently doubled by inverse PCR, thereby obtaining a linear strand having a junction region in the 3' direction of the LTR region or a linear strand having a junction region in the 5' direction of the LTR region.
6. 6. The method of claim 1, wherein the PCR primers P1 and / or P2 are provided at the 3' end and / or the 5' end of each of the at least one LTR region in the circularized strand, such that the PCR primer P1 faces toward the 5' end of the 5'LTR or the 3'LTR region, and the PCR primer P2 faces toward the 3' end of the 5'LTR or the 3'LTR region.
7. 7. The method according to any one of claims 1 to 6, characterized in that the copy number of the proviral sequence in the DNA strand is obtained.
8. 8. The method according to claim 1, wherein the sequence information of the DNA strand is obtained.
9. 9. The method according to claim 1, wherein step g) is performed by mapping the sequence information of the junction region to the sequence information of the DNA strand, and the region of the DNA strand that aligns with high confidence (MAPQ ≥ 30) to the sequence information of the junction region is designated as the genomic location of the proviral sequence.