A method for analyzing integration sites based on multi-primer design
By designing multiple primers in a recombinant adeno-associated virus vector and combining them with bioinformatics algorithms to process sequencing data, the problems of low detection sensitivity and redundancy caused by the integration of incomplete sequences into AAV vectors were solved, enabling accurate detection of integration sites and drug safety assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI WEIKE BIOTECHNOLOGY CO LTD
- Filing Date
- 2025-06-27
- Publication Date
- 2026-07-31
AI Technical Summary
Existing conventional integration site analysis methods cannot effectively capture cases where recombinant adeno-associated virus (rAAV) vectors integrate into the host genome with incomplete sequences, resulting in reduced detection sensitivity and problems with redundancy and base loss in sequencing results.
Primers at multiple locations were designed, and a redundancy removal bioinformatics algorithm and primer optimization process were combined with multi-primer sequencing data. Integration site sequences were captured by multi-primer amplification, and the sequencing results were processed by redundancy removal and optimization algorithms to ensure the accuracy and sensitivity of the detection.
It improves the detection sensitivity in the case of AAV vector terminal ITR loss, eliminates the redundancy of sequencing results, and ensures accurate statistics of integration sites and drug safety assessment.
Smart Images

Figure CN120766765B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics technology. Specifically, it belongs to the application direction of second-generation sequencing data analysis in the field of bioinformatics. It discloses an integration site analysis method based on multi-primer design, which aims to carry out accurate detection and analysis of integration sites in recombinant adeno-associated virus (rAAV) vectors. Background Technology
[0002] Recombinant adeno-associated virus (rAAV) is one of the commonly used vectors for integration analysis. Adeno-associated virus (AAV) is a non-enveloped virus. After being designed and edited, recombinant adeno-associated virus (rAAV) can be used as a vector to deliver sequences to target cells. It has the characteristics of simple structure, high safety and low immunogenicity. Therefore, rAAV has become one of the most widely used vectors in the field of gene therapy.
[0003] Integration site analysis (ISA) refers to the analysis of the location and functional annotation of the integration site after the vector integrates into the host genome, including preclinical and clinical gene therapy. ISA can predict whether the vector will cause adverse reactions after integration into the target cell genome and is key to assessing the biosafety of gene therapy vectors.
[0004] Compared to other types of vectors such as lentiviral vectors, AAV vectors do not insert the complete viral sequence when integrating into the host genome. Instead, only a portion of the sequence is integrated into the host genome, usually losing the ITR (Inverse Tandem Repeats) at both ends of the vector. In this case, conventional analysis methods based on rAAV ITR sequences for primer amplification may fail to identify valid primer sequences during sequence capture, resulting in missing data.
[0005] Conventional integration site analysis methods, when designing primers (e.g., Chinese patent document CN109554447A), design only a single primer at the end of the viral vector sequence, using the capture of the LTR region as the criterion for viral sequence capture. In subsequent bioinformatics analysis steps for integration site detection, the single primer sequence is also used as the basis for sequencing result identification and integration site statistics. Furthermore, during primer matching, a allowed mismatch is set to address the issue of minor base loss at the primer ends during integration (e.g., Chinese patent document CN116403647A).
[0006] However, as Figure 1As shown, current conventional experimental methods only design a single primer at the end of the vector sequence. Since the AAV vector sequence integrates into the host genome in the form of a non-complete sequence, it is impossible to effectively capture the integration sequence when the integration sequence does not contain a terminal primer, which reduces the sensitivity of detection. Summary of the Invention
[0007] To address the aforementioned problem of AAV vector sequence breaks leading to undetectable integration sites by single primers, this invention, through research, has discovered that by designing primers at multiple locations during the experimental design phase, the problem of ineffective capture of the integration site sequence can be solved when the AAV vector integrates into the host genome with an incomplete sequence during the experiment. Therefore, the first objective of this invention is to provide a method for integration site analysis based on multi-primer design. The second objective of this invention is to provide primers for the aforementioned analysis method. The third objective of this invention is to provide a computer-readable storage medium. The fourth objective of this invention is to provide a server system for executing the integration site analysis method based on multi-primer design. The fifth objective of this invention is to provide a cloud-based service platform.
[0008] To address the above problems, the present invention adopts the following technical solution:
[0009] As a first aspect of the present invention, a method for integration site analysis based on multi-primer design includes the following steps:
[0010] S1. Primers are inserted at multiple positions in the AAV vector sequence to form a multi-primer vector;
[0011] S2. Integrate the multi-primer vector into the host genome, and capture the sequence containing the integration site by primer amplification;
[0012] S3. Sequencing is performed on the sequence containing the integration site to obtain sequencing data;
[0013] S4. Sequencing data redundancy removal: Through a multi-primer sequencing data redundancy removal bioinformatics algorithm, the redundancy of sequencing results caused by multi-primer amplification is eliminated, and accurate integration site detection results are obtained.
[0014] According to the present invention, the multi-primer sequencing data redundancy removal bioinformatics algorithm includes the following steps:
[0015] Obtain all sequencing sequences that support integration sites and label the primer information contained in each sequence;
[0016] Compare the primer information of different sequences. If the primer information of one sequence is a subset of the primer information of another sequence, retain the sequence corresponding to the subset and filter out redundant sequences.
[0017] Redundancy removal was performed on all integration sites to obtain accurate statistical results of integration sites.
[0018] According to the present invention, the primer design includes the following steps:
[0019] Design multiple primers based on the ITR region of the AAV vector sequence;
[0020] The effectiveness and specificity of the primers were verified through experiments to ensure that the primers could effectively amplify the target sequence.
[0021] Furthermore, the nucleotide sequences of the primers are shown in SEQ ID No. 1, SEQ ID No. 2 and SEQ ID No. 3.
[0022] According to the present invention, the redundancy removal process for sequencing data includes the following steps:
[0023] S41. Use the k-mer algorithm to enumerate candidate primer subsequences;
[0024] S42. Count the occurrence frequency of each candidate primer subsequence and select the subsequence with the highest occurrence frequency as the optimal primer;
[0025] S43. Extract and integrate sequences based on the optimal primers to maximize the retention of sequencing results.
[0026] As a second aspect of the present invention, a primer for the analytical method described above, the nucleotide sequence of which is shown in SEQ ID No. 1, SEQ ID No. 2 and SEQ ID No. 3.
[0027] As a third aspect of the invention, a computer-readable storage medium stores a computer program for performing an integration site analysis method based on multi-primer design, the computer program including bioinformatics algorithms for redundancy removal and primer selection, the computer program being capable of running on a server or cloud platform.
[0028] As a fourth aspect of the present invention, a server system for performing an integration site analysis method based on multi-primer design is characterized in that it includes a data receiving module, a data analysis module, and a result output module, wherein the data analysis module is configured with bioinformatics algorithms for redundancy removal and primer selection, and the system is capable of running on a cloud platform.
[0029] As a fifth aspect of the present invention, a cloud-based service platform is provided for executing an integration site analysis method based on multi-primer design. The service platform includes a data upload module, a data analysis module, and a result download module. The data analysis module is configured with bioinformatics algorithms for redundancy removal and primer selection. The platform supports concurrent access by multiple users.
[0030] As a sixth aspect of the present invention, an apparatus for performing an integration site analysis method based on multi-primer design includes a gene disruptor, a PCR instrument, a sequencer, and a data analysis server, the data analysis server being configured with bioinformatics algorithms for redundancy removal and primer selection.
[0031] As a seventh aspect of the present invention, an integration site analysis method based on multi-primer design is applied to drug safety assessment. By detecting the integration site of the AAV vector in the host genome, the potential genotoxicity of the drug to the host genome is assessed to optimize drug safety.
[0032] As an eighth aspect of the present invention, an integration site analysis method based on multi-primer design is applied to gene therapy vector design. By detecting the integration site of the AAV vector in the host genome, the integration efficiency and safety of the gene therapy vector are evaluated to optimize the design of the gene therapy vector.
[0033] Compared to existing technologies, the beneficial effect of this invention is that, through specific optimization methods, the sensitivity for detecting integration sites in vectors with randomly lost terminal ITRs is greatly improved. Specifically, this is reflected in:
[0034] 1. Multi-primer vector experimental scheme: The experimental design steps of this invention solve the problem that the integration site sequence cannot be effectively captured when the AAV vector integrates into the host genome with an incomplete sequence during the experiment by designing primers at multiple positions.
[0035] Specifically, such as Figure 2 As shown in the experimental design section, primers are inserted at multiple locations within the AAV vector sequence. Under this premise, when the AAV vector breaks into shorter sequences and integrates into the host genome, the integration sequence can be detected using the remaining primers in the shorter sequences. This approach effectively avoids the problem of ineffective capture of the integration site sequence caused by primer loss during experiments with single-primer vector sequences, thereby improving detection sensitivity.
[0036] 2. Multi-primer redundancy removal scheme: To address the redundancy problem in sequencing results of multi-primer integrated sequences, this invention designs a multi-primer sequencing data redundancy removal bioinformatics algorithm to eliminate the redundancy caused by multiple primers amplifying the same integrated sequence, thereby obtaining accurate quantitative results.
[0037] Specifically, when an integration sequence contains multiple primers, each primer generates a separate sequence after amplification, resulting in redundancy in the sequencing results. Therefore, it is necessary to filter out redundant results amplified from the same integration sequence. For sequences at the same integration site, the primer inclusion relationship is examined. When multiple sequences have primers that are mutually inclusive, the sequence corresponding to the included sequence is retained. This approach can accurately remove redundant terms from the sequencing results, reducing the bias in the integration site statistical results caused by redundancy from multiple primer amplification.
[0038] 3. Primer Subsequence Selection Scheme: Primer sequences are one of the important bases for extracting and integrating sequencing results. Since a small number of bases may be lost at the primer ends, this invention selects the optimal primer subsequences to ensure they appear in the most frequent sequencing sequences during the analysis of actual sequencing data. This maximizes the effective sequencing results and solves the problem of undetectable ITR deletion sites on primers.
[0039] Specifically, to address the problem of reduced sequencing data matching range due to the loss of terminal bases in primer sequences, this invention establishes a primer sequence optimization process. When analyzing actual sequencing data, it is necessary to screen the optimal primer sub-sequences from the actual data so that they appear in the most sequencing sequences, thereby maximizing the effective sequencing results obtained. Attached Figure Description
[0040] Figure 1 This is a schematic diagram of the experimental design for a conventional integration site. In experiments using conventional integration sites, only one primer is designed at the end of the vector sequence. When the vector sequence integrates into the host genome as an incomplete sequence, the integration sequence without a primer cannot be detected, resulting in lost results.
[0041] Figure 2 This is a schematic diagram of the experimental design for the integration site of the present invention. The novel experimental method inserts primers at multiple locations in the vector sequence, so that even if the vector sequence is not a complete sequence when it integrates into the host genome, it can still be detected by the primers in that region.
[0042] Figure 3 This is a schematic diagram illustrating the integration of multiple primer vector sequences according to the present invention. When the AAV vector sequence integrates into the host genome, sequence breaks occur, resulting in an incomplete vector sequence ultimately integrated into the host genome. Since primers exist at multiple locations within the vector sequence, the integrated sequence may also contain multiple primers, each of which will amplify a result sequence.
[0043] Figure 4This diagram illustrates the principle behind sequencing result redundancy. Primers are designed at multiple locations within the vector sequence. After the vector sequence integrates into the host genome, the integrated sequence may still contain multiple primers. Each primer triggers an amplification, resulting in multiple sequencing results from the same integrated sequence, thus generating redundant results.
[0044] Figure 5 The flowchart shows the process of removing redundancy from multi-primer sequencing data using the bioinformatics algorithm of this invention.
[0045] Figure 6 This is an example of sequence-primer association. Each element Si in the sequence list S is a sequence, and each element Vi in the primer information vector V is the primer number contained in Si. The sequence list S and the primer information vector V have the same length, and Si and Vi correspond one-to-one.
[0046] Figure 7 A schematic diagram of the process for selecting the best primer sequence.
[0047] Figure 8 Example of selecting the best primer sequence.
[0048] Figure 9 To verify the effectiveness and specificity of the primers, amplification was performed on plasmid standards. From left to right, the primers are Marker, ITR1, ITR2, and ITR3.
[0049] Figure 10 To construct a library for triple-well targeted amplification of samples using designed AAV multiple primers, a diffuse library structure was obtained, which could then be sequenced. Detailed Implementation
[0050] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0051] 1. The specific meanings of the terms used in this invention are as follows:
[0052] AAV: Adeno-Associated Virus;
[0053] Barcode: A short and unique DNA sequence tag used to identify different samples or experimental conditions;
[0054] Primer: A primer is a base sequence used to initiate a DNA amplification reaction and is closely linked to a barcode.
[0055] Fastq files, abbreviated as fq, are a file format for storing biological sequences (usually nucleic acid sequences) and their sequencing quality score information;
[0056] k-mer: A subsequence consisting of k consecutive bases in a DNA or RNA sequence;
[0057] Integration site: The integration site refers to the specific location where the vector sequence is inserted into the host genome.
[0058] 2. Experimental Procedure
[0059] (1) In this invention, 1.5 μg of genomic DNA was cut to an average length of 500 ± 50 bp using a Covaris M220 gene disruptor (Tapestation 4150). The fragmented DNA was purified, and primer extension was performed using AAV vector-specific biotin-modified primers. The reaction system was 50 μl, containing 0.01 mM dNTPs and 1.25 U Taq enzyme. The cycling program was 95°C pre-denaturation for 2 minutes followed by one cycle (95°C for 45 seconds, 58°C for 45 seconds, and 72°C for 5 seconds extension). The extension product was purified using Ampure XP magnetic beads, then captured by Dynabeads M-280 biotin affinity magnetic beads, and washed with PBS-BSA solution and LiCl buffer to construct the DNA-magnetic bead complex. Subsequently, using a ligation system containing 25% PEG8000 and 1 mM hexammine cobalt(III) chloride, a 5′ phosphorylated, 3′ end ddC-modified single-stranded linker (10 pmol) was ligated to the unknown end of the target DNA. The ligation reaction was incubated overnight (16–24 hours) at room temperature with shaking at 300 rpm, followed by elution and purification. The ligation product was divided into two aliquots. One aliquot was used for a first nested PCR with biotinylated target fragment primers and ligation adapter primers at a primer concentration of 8.35 pmol. The cycling program consisted of 10 cycles of denaturation at 95°C followed by 25 cycles of denaturation at 90°C, for a total of 35 cycles. The other aliquot was stored as an intermediate product in a library. The biotinylated PCR product was captured by magnetic beads, enriched, and washed. Half of the elution buffer was used as a template for a second exponential PCR, amplified using the inner vector and adapter primers, and ligated with the barcode and sequencing adapter required for sequencing library construction. The PCR products were qualitatively controlled by 2% gel electrophoresis; diffuse bands indicated successful amplification. The remaining products were purified using Qubit 4.0 to control product concentration, and mixed with samples at 50 ng / sample. Next-generation sequencing was then performed on a MiSeq sequencer (Illumina). Detailed experimental procedures can be found in the section on High-resolution insertion-site analysis by linear amplification–mediated PCR (LAM-PCR).
[0060] (2) After amplifying the AAV vector fragment using multiple primers in the experimental steps, since the break position of the AAV vector sequence is not fixed, the integrated sequence obtained after integration may contain multiple primers. Each primer will trigger a sequence amplification, which will eventually lead to one integrated sequence containing multiple primers producing multiple result sequences (e.g., Figure 4As shown in the figure, these result sequences yielded multiple sequencing sequences, resulting in redundancy. To address this, this invention designs a multi-primer sequencing data redundancy removal bioinformatics algorithm. This algorithm filters the data based on the primer information of the sequencing data to obtain accurate quantitative results. Figure 5 The flowchart illustrates the process, and the specific implementation method of this process is as follows:
[0061] 1) Obtain paired-end Fastq files through sequencing, and analyze them using integration site detection tools to obtain the detection results of redundant integration sites, including the specific location of the integration site on the host genome and the sequencing sequence that supports the integration site. It should be noted that the sequencing sequence that supports the integration event theoretically consists of two parts: one part is the host genome part and the other part is the terminal part of the AAV genome.
[0062] 2) For any integration site, obtain all sequencing sequences S supporting that integration site. For each sequence in S, locate the primer sequences contained within its AAV genome terminal portion, label the presence of these primer sequences, and store the labeling information in vector V. Specifically, if the AAV genome terminal portion contains primer sequence A, record 'A' in vector V; if the AAV genome terminal portion contains primer sequences A, B, and C simultaneously, record 'ABC' in vector V. After recording, the length of vector V is equal to that of S, and it corresponds one-to-one with each sequence in S (e.g., ...). Figure 4 (As shown).
[0063] 3) For any integration site, obtain all sequencing sequences S supporting that integration site and a vector V indicating the presence of primers. Iterate through all sequences, comparing the information stored in the vector V between any two sequences. Assuming i and k are the i-th and k-th sequences in sequence S, if the primer information Vi corresponding to a sequence Si is a subsequence of the primer information Vk corresponding to Sk (e.g., 'A' is a subsequence of 'ABC'), then the sequence Si corresponding to the included sequence is retained, and the other sequence Sk is discarded. Perform the above filtering operation on all sequences corresponding to integration sites to obtain the final statistical results of integration sites, which no longer contain redundant information, thus obtaining accurate quantitative results. See the detailed flowchart for more details. Figure 5 .
[0064] (3) To address the issue of sequencing data containing only incomplete primer sequences due to primer terminal base loss, this invention introduces a novel bioinformatics algorithm to automatically identify new primers from actual sequencing data and uses the kmer algorithm to statistically select the optimal primer sequences, thereby maximizing the retention of sequencing results. Figure 7 The diagram illustrates the primer sequence optimization process, and the specific implementation method is as follows:
[0065] 1) Given a Fastq file obtained after sequencing, traverse the sequences in the file, extract a fixed-length base sequence in the primer position region, and obtain a candidate primer list.
[0066] 2) Use the kmer algorithm to enumerate all subsequences of candidate primers with a length not less than k. That is, from each candidate primer sequence, take out k consecutive bases, k+1 bases, k+2 bases, etc. until multiple subsequences of the full length of the candidate primer sequence to obtain a list of candidate primer subsequences.
[0067] 3) Count the frequency of each primer subsequence in the candidate primer subsequence list and take out the primer subsequence with the highest frequency. If multiple primer subsequences have the same frequency, take out the longest primer subsequence to obtain a unique candidate primer.
[0068] 4) Check if the frequency of the unique candidate primer sequence passes the threshold test. In this method, the threshold for determining the candidate primer sequence is 0.6, meaning that the candidate primer sequence should appear in at least 60% of the Fastq file sequences. If it passes the threshold test, a new primer sequence is obtained; otherwise, the primer region is reselected and steps 1-3 are repeated.
[0069] Thus, a new primer sequence is obtained. Using this new primer sequence as a basis for extracting the integrated sequence can achieve the maximum preservation of the result data.
[0070] Example 1
[0071] This embodiment mainly describes the design of multiple primers based on the ITR sequence in the AAV vector sequence, and the results of library construction and sequencing based on the primer sequence.
[0072] In this embodiment, the specific implementation process is as follows:
[0073] 1. AAV vector primer design:
[0074] The ITR sequence is as follows:
[0075] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCG, (SEQ ID No. 1);
[0076] Based on conventional primer design principles, the following multi-primer system was ultimately determined:
[0077] ITR1:GAACCCCTAGTGATGGAGTTGGC, (SEQ ID No. 2);
[0078] ITR2:ATGGAGTTGGCCACTCCCTC,(SEQ ID No.3);
[0079] ITR3:GAGTTGGCCACTCCCTCTCTG,(SEQ ID No.4);
[0080] 2. Primer verification was performed based on the primers described above, and the results were obtained. Figure 9 After confirming the primers, targeted amplification and library construction were performed on the DNA from AAV-infected samples, and the following library quality control results were finally obtained. Figure 10 .
[0081] Example 2
[0082] This embodiment mainly describes a bioinformatics algorithm for removing redundancy from multi-primer sequencing data, and compares the statistical results before and after redundancy removal of one data set, aiming to demonstrate that the algorithm can effectively remove multiple amplified sequences caused by multiple primers. Figure 5 A flowchart illustrating the redundancy removal algorithm for multi-primer sequencing data.
[0083] In this embodiment, the input data that needs to be prepared includes:
[0084] (1) One pair of paired-end Fastq files: One human cell sample infected with AAV was infected using the AAV2 serotype viral vector. The infection conditions were set at MOI (Multiplicity of Infection) = 10,000 vg / cell, and cells were collected after 72 hours of culture. After extracting genomic DNA, PCR amplification was performed using primers specific to the AAV vector sequence, followed by high-throughput sequencing (NGS) to obtain one pair of paired-end Fastq files (150bp).
[0085] In this embodiment, the specific analysis steps are as follows:
[0086] 1. Detection of integration sites: For the paired-end Fastq files obtained from sequencing, integration site detection tools (such as GENIS, INSPIRED, etc.) are used for data analysis to obtain the original integration site detection results. At this time, the original integration site detection results obtained contain redundancy.
[0087] 2. Association between integration sites and sequencing sequences: The original integration site detection results include the specific location of the integration site on the host genome and the sequencing sequences supporting the integration site. It should be noted that the sequencing sequences supporting the integration event theoretically consist of two parts: (1) a part of the host genome and (2) a part of the AAV genome terminal portion, and the AAV genome portion will contain at least one primer sequence. For any integration site, all sequencing sequences supporting the integration site are obtained, and the sequence list is denoted as S. In this embodiment, taking the integration site chr6:100405492 as an example, the sequence list S of this integration site contains a total of 647 sequences.
[0088] 3. Sequence and Primer Association: For each sequence in sequence list S, locate the primer sequences contained within the AAV genomic terminal portion of the sequence, mark the presence of these primer sequences, and store the marking information in vector V. Taking the integration site chr6:100405492 in this embodiment as an example, its sequence list S contains a total of 647 sequences. The first sequence S1 contains only primer A (primer A represents the sequence GTACCGTAAAGTTGAGAGTC), and the second sequence S2 contains both primer A and primer B (primer B represents the sequence CGTAACAAAGGTCCCGATAG). Therefore, V1 = 'A' and V2 = 'AB'. Perform the above operation on all sequences in sequence list S. After recording, the length of vector V is equal to that of S and corresponds one-to-one with each sequence in S (e.g., ...). Figure 6 (As shown).
[0089] 4. Redundant Information Filtering: For any integration site, obtain all sequencing sequences S = [S1, S2, S3...Sn] that support the integration site, and a vector V = [V1, V2, V3...Vn] recording primer annotation information. The number of elements in both vectors is equal and they correspond one-to-one. Iterate through all sequences in S and compare the primer annotation information of every two sequences. Taking the integration site chr6:100405492 in this embodiment as an example, the primer information vector V1 = 'A' corresponding to the first sequence S1, and the primer information vector V2 = 'AB' corresponding to the second sequence S2. Since V1 is a true subset of V2, S1 is retained, and S2 is filtered out.
[0090] 5. Perform the above filtering operation on the sequences corresponding to all integration sites to obtain the final statistical results of integration sites, which no longer contain redundant information, thus obtaining accurate quantitative results.
[0091] Table 1 shows the sequence count supported by the integration sites before and after redundancy removal. Three integration sites are selected for display. After redundancy removal, the sequence count of all sites decreased, proving that the multi-primer sequencing data redundancy removal algorithm achieved data filtering and made the quantitative results more accurate.
[0092] Table 1. Statistics on the number of sequences supported by the integration site before and after redundancy removal.
[0093]
[0094] Example 3
[0095] In this embodiment, a primer selection algorithm was used to screen for new primers on one dataset, and the number of integrated sequences extracted using the complete primer sequence and the new primer sequence were compared to demonstrate that the algorithm can maximize the preservation of sequencing results. Figure 7 A schematic diagram of the process for selecting the best primer sequence.
[0096] (1) One pair of paired-end Fastq files: One human cell sample infected with AAV was infected using the AAV2 serotype viral vector. The infection conditions were set at MOI (Multiplicity of Infection) = 10,000 vg / cell, and cells were collected after 72 hours of culture. After extracting genomic DNA, PCR amplification was performed using primers specific to the AAV vector sequence, followed by high-throughput sequencing (NGS) to obtain one pair of paired-end Fastq files (150bp).
[0097] In this embodiment, the specific analysis steps are as follows:
[0098] 1. Extract candidate primer sequences: For the paired-end Fastq file obtained after sequencing, traverse the sequencing sequences and extract candidate primer sequences according to the position of the primer sequence in the Fastq file. In this embodiment, the length of the candidate primer sequence is 20 bases. After traversal, all candidate primer sequences are output to the candidate primer list P, resulting in a total of 743,006 candidate primer sequences.
[0099] 2. Extracting candidate primer subsequences using the k-mer algorithm: The k-mer algorithm is used to enumerate all subsequences of length k or longer than the primer sequence. In this example, k = 18. Specifically, multiple subsequences of consecutive lengths of 18, 19, and 20 bases are extracted from each candidate primer sequence to obtain a candidate primer subsequence list Q (e.g., ...). Figure 8 (as shown in a).
[0100] 3. Selection of candidate primer subsequences: For each candidate primer subsequence in the candidate primer subsequence list Q, its frequency of occurrence in the primer list P is calculated, i.e., what percentage of the primer sequences contains that subsequence. The candidate primer subsequence with the highest frequency is obtained. For example... Figure 8 As shown in b, the subprimer sequence ACGCCTGCTTGAAAGTGCC has the longest length and appears the most times in the primer list P, at 572233 times, therefore it is selected as the new primer sequence.
[0101] 4. Check if the occurrence frequency of this unique candidate primer sequence meets the 60% threshold test. The check revealed that it appeared in 77.02% of primer sequences, thus meeting the release criteria.
[0102] Thus, a new primer sequence is obtained. The new primer sequence selected by the primer selection method will be included in the most sequencing sequences. Therefore, using this new primer sequence as the basis for extraction of the integration sequence can achieve the maximum preservation of the result data.
[0103] Table 2 shows the number of sequences extracted using the original primer sequence and the optimized new primer sequence, respectively. It can be seen that using the new primer sequence for sequence extraction increased the analyzable data by 20.34%, proving that the algorithm can select new primer sequences that are more applicable to actual sequencing data, thus maximizing the retention of sequencing data.
[0104] Table 2. Statistics on the number of sequences obtained by extracting using the primer sequence.
[0105]
[0106] In summary, this invention addresses the problem of ineffective capture of the integration site sequence when the AAV vector integrates into the host genome with an incomplete sequence during the experimental design phase. Furthermore, during sequencing, if the rAAV breakpoint happens to be on the primer sequence, a small number of terminal bases may be lost. Directly using the reference primer sequence for sequence matching would result in the loss of matching data. Even if a small number of primer sequence mismatches are allowed during matching, the matching logic is not entirely consistent with the logic of terminal base loss. Therefore, this invention designs a bioinformatics algorithm to select the optimal primer sequence from subsequences based on actual sequencing data to maximize detection results.
[0107] Furthermore, based on the experimental design of inserting multiple primers into the AAV vector sequence, after the AAV vector sequence integrates into the host genome, if an integrated sequence contains multiple primers, each primer will trigger a sequence amplification. Ultimately, an integrated sequence containing multiple primers will amplify multiple different result sequences, meaning that the original result sequences contain data redundancy. Therefore, this invention designs a bioinformatics algorithm to eliminate the result redundancy caused by multiple primers amplifying the same integrated sequence, obtaining accurate quantitative results.
[0108] The above are merely a few preferred embodiments of the present invention, described in a relatively specific and detailed manner, but they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A method for integration site analysis based on multi-primer design, characterized in that, Includes the following steps: S1. Primers are inserted at multiple positions in the AAV vector sequence to form a multi-primer vector; S2. Integrate the multi-primer vector into the host genome, and capture the sequence containing the integration site by primer amplification; S3. Sequencing is performed on the sequence containing the integration site to obtain sequencing data; S4. Sequencing data redundancy removal: Through a multi-primer sequencing data redundancy removal bioinformatics algorithm, the redundancy of sequencing results caused by multi-primer amplification is eliminated, and accurate integration site detection results are obtained. The multi-primer sequencing data redundancy removal bioinformatics algorithm includes the following steps: Obtain all sequencing sequences that support integration sites and label the primer information contained in each sequence; Compare the primer information of different sequences. If the primer information of one sequence is a subset of the primer information of another sequence, then retain the sequence corresponding to the subset and filter out the redundant sequence. Redundancy removal was performed on all integration sites to obtain accurate statistical results of integration sites.
2. The integration site analysis method of claim 1, wherein, The primer design includes the following steps: Design multiple primers based on the ITR region of the AAV vector sequence; The effectiveness and specificity of the primers were verified through experiments to ensure that the primers could effectively amplify the target sequence.
3. The integration site analysis method of claim 2, wherein, The nucleotide sequences of the primers are shown in SEQ ID No. 1, SEQ ID No. 2 and SEQ ID No.
3.
4. The integration site analysis method according to any one of claims 1 to 2, wherein, The redundancy removal process for the sequencing data includes the following steps: S41. Use the k-mer algorithm to enumerate candidate primer subsequences; S42. Count the occurrence frequency of each candidate primer subsequence and select the subsequence with the highest occurrence frequency as the optimal primer; S43. Extract and integrate sequences based on the optimal primers to maximize the retention of sequencing results.
5. A computer readable storage medium, characterized in that, The computer program stores a method for performing integration site analysis based on multi-primer design as described in any one of claims 1-4, the computer program including bioinformatics algorithms for redundancy removal and primer selection.
6. A server system for performing a multi-primer design based integration site analysis method according to any one of claims 1 to 4, characterized in that, It includes a data receiving module, a data analysis module, and a result output module. The data analysis module is equipped with bioinformatics algorithms for redundancy removal and primer selection.
7. An apparatus for performing a multi-primer design based integration site analysis method as claimed in any one of claims 1 to 4, characterized in that, It includes a gene disruptor, a PCR instrument, a sequencer, and a data analysis server, wherein the data analysis server is configured with bioinformatics algorithms for redundancy removal and primer selection.
8. The application of an integration site analysis method based on the multi-primer design as described in any one of claims 1-4 in drug safety assessment or gene therapy vector design.