Long-reading long-load quality control method based on UMI consensus and system and application thereof

By combining single-end UMI labeling and three rounds of PCR amplification with a UMI consensus-based approach, the problems of full-length integrity, quantification of a few components, and monitoring of packaging sequence recombination in the quality control of integrated lentiviral vectors were solved, achieving high accuracy and stable quality control results.

CN121496044APending Publication Date: 2026-02-10INST OF HEMATOLOGY & BLOOD DISEASES HOSPITAL CHINESE ACADEMY OF MEDICAL SCI & PEKING UNION MEDICAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511689746.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing integrated lentiviral vector quality control methods are insufficient in terms of full-length integrity, quantification of sub-percentage minority components, and monitoring of packaging sequence recombination. In particular, they are difficult to effectively suppress non-specific transmission and chimeras in the context of the human genome, and the accuracy of long-read platforms is insufficient.

Method used

A single-end UMI labeling strategy combined with semi-nested and fully nested amplification methods was adopted. Through three rounds of PCR amplification and a UMI consensus method, it was ensured that only labeled target templates were amplified. Data analysis was performed using long-read sequencing technology to achieve full-length determination, counting of a few components, and monitoring of packaging sequence recombination.

Benefits of technology

It significantly suppressed nonspecific propagation and chimeras, increased the proportion of full-length sequences to 99%, and reduced the mismatch and insertion/deletion event rate to 10⁻⁴–10⁻⁵ bp⁻¹, achieving highly accurate molecule counting and recombination monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121496044A_ABST
    Figure CN121496044A_ABST
Patent Text Reader

Abstract

The invention discloses a long-reading long-load quality control method based on UMI consensus and a system and application thereof, and belongs to the technical field of gene detection and biological medicine. According to the method, three rounds of single-ended UMI marking-semi-nested-nested PCR (Polymerase Chain Reaction) system are adopted. Obtained products are subjected to Nano or PacBio sequencing, grouping is carried out according to UMI, UMI-by-UMI consensus sequences are generated, and the UMI-by-UMI consensus sequences are compared with reference sequences to output full-length integrity rate, rare component quantification and packaging sequence recombination detection results. Experimental results show that the specificity of the method is improved by about 100 times, the error rate can be reduced to 104-10-5 (Q is approximately equal to 40) under the condition of more than or equal to 5 read segments / UMI, the full-length coverage rate is approximately equal to 99%, 0.1% of low-abundance pollution and packaging recombinant molecules can be quantitatively detected, and the method is used for safety and consistency evaluation of gene and cell therapy vectors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of gene detection and biomedicine, specifically to a long-read vector quality control method based on UMI consensus and its reagent kit and system. Background Technology

[0002] Currently, integrated lentiviral vectors (LVs) are widely used in gene and cell therapy. With the expansion of clinical applications, genetic QC for CMC / release and batch-to-batch comparability mainly focuses on three types of issues: (i) full-length integrity and sequence accuracy: whether the integrated provirus maintains end-to-end (typically 5–10 kb) and structural stability; (ii) sub-percentage fractional component quantification: whether a count-calibrated proportion and confidence interval can be provided for components ≤0.1% in a mixed system; and (iii) packaging sequence recombination monitoring: whether there are molecular splices with packaging-derived sequences such as gag / pol / rev / VSV-G, serving as molecular clues to RCL risk. Meanwhile, commonly used solutions each have their shortcomings. For example, qPCR / Sanger only covers local sites; short-read sequencing struggles to span entire provirus and repetitive modules (such as LTR and WPRE), and in multi-sample scenarios, it is susceptible to index (barcode) mismatches / crosstalk, making it difficult to reliably estimate sub-percentage components. While long-read platforms can cover the full length, the original reads contain mismatches / insertions (indels) and regional biases, making them insufficient for accurate interpretation of low-frequency events when used directly. Furthermore, upstream long-fragment PCR is further biased by length / GC dependence and template switching, resulting in "read count ≠ molecule count". To suppress amplification and sequencing errors, the UMI strategy can "fold" reads into UMI consensus molecules and directly count the original molecules, which theoretically helps improve accuracy and quantitative reliability. However, the published paired-end UMI + secondary non-selective amplification route has inherent limitations in the context of eukaryotic genomes: the second round amplifies the non-specific products generated in the first round, and the nesting makes it difficult to achieve the selectivity of "only amplifying the target template again", leading to non-specific propagation and chimera problems; related work is mostly validated in small genome / bacterial scenarios, and it is difficult to reproduce stably in the complex host background of integrated LV.

[0003] Therefore, there is an urgent need for a UMI-long-read method for integrated latent gene mutations (LVs). This method should be designed to amplify only the labeled target template after the first round of single-end UMI labeling, using semi-nested / fully nested site constraints to suppress non-specific propagation at the source. Downstream, mismatches / indels should be reduced to a usable level for mutation interpretation using a per-UMI consensus mechanism, and the following should be completed within a unified dataset according to auditable thresholds: (a) full-length determination (e.g., "continuous coverage ≥95% / continuous coverage ≥95% or "self-length ≥90% continuous alignment"); (b) molecular counts and Clopper-Pearson confidence intervals for a few components; and (c) setting a positive rule of "≥100bp continuous alignment and ≤5bp indels" for packaging gene references (gag / pol / rev / VSV-G) for recombination monitoring. These requirements and criteria have been repeatedly raised in recent QC scenarios, but existing publicly available solutions have not yet addressed the integrated LV context with single-end UMI + semi-nested selective re-amplification combined with per-UMI methods. The consensus is to simultaneously satisfy the auditability of the three outputs and the daily turnover target. Summary of the Invention

[0004] The purpose of this invention is to provide a quality control method, system and application for long read long carriers based on UMI consensus.

[0005] In a first aspect, the present invention provides a long-read sequencing method for quality control of integrated lentiviral vectors, the method comprising the following steps: a) Perform the first round of PCR, using forward primer F1 and reverse primer R1 to amplify the long fragment product, wherein the forward primer F1 contains a 5′ constant tail, a unique molecular identifier (UMI) sequence and a sequence targeting the 5′ LTR, and the reverse primer R1 targets the 3′ LTR sequence. b) Perform a second round of PCR using forward primer F2 and reverse primer R2 for semi-nested amplification, wherein forward primer F2 binds to the 5′ constant tail and reverse primer R2 is located inside the reverse primer R1, thereby selectively amplifying the product carrying the constant tail and UMI. c) Perform a third round of PCR using forward primer F3 and reverse primer R3 for fully nested amplification, wherein forward primer F3 contains the sample barcode and binds the 5′ constant tail, and reverse primer R3 is located inside reverse primer R2 to obtain indexed amplicon; d) Perform long-read sequencing on the indexed amplicon to obtain sequencing reads; e) Based on the 5′ constant tail location parsing UMI, the read segments are grouped according to UMI, and multi-sequence consistency calculation is performed on the groups with ≥5 read segments per UMI to generate a per-UMI consensus sequence; f) Compare the UMI consensus sequence with the carrier reference sequence and classify them based on the following rules: f1) Completely in the target sequence: The UMI consensus sequence continuously covers ≥95% of the reference sequence length, or the consensus sequence itself is continuously aligned to the reference sequence for ≥90% of its length; f2) Packaging and Reassembly Sequence: The UMI consensus sequence has a continuous alignment block of ≥100 bp (CIGAR M / = / X) with the gag / pol / rev / VSV G reference and allows small indels of ≤5 bp; f3) Contaminated sequences or non-packaged recombination sequences: Iterative classification is performed based on 30 nt motifs of 5′Ψ and 3′WPRE, where the edit distance of a single motif is ≤3; f4) Non-specific sequences: sequences that do not belong to any of the above categories; g) Output the full-length proportion, minority component proportion, and packaging recombination conclusions based on UMI molecular counting.

[0006] Preferably, the number of cycles in the first round of PCR is 1-3.

[0007] Preferably, the length of the UMI is 18 nt, and the degeneracy mode is HBNHVNBDNHVNBDNHBD.

[0008] Preferably, the multi-sequence consistency calculation is performed using partial sequence ordered alignment or an equivalent method.

[0009] Preferably, the proportion of the minority components is given a 95% confidence interval using the Clopper-Pearson method.

[0010] Preferably, the long-read sequencing is Oxford Nanopore or PacBio sequencing, and the mismatch event rate is no higher than 2 × 10⁻⁵ reads per UMI. -4 bp -1 The insertion or missing event rate is no higher than 8 × 10⁻⁶. -5 bp -1 .

[0011] Secondly, the present invention provides a primer composition for the above-described method, the primer composition comprising: Forward primer F1, which contains a 5′ constant tail, a UMI sequence, and a sequence targeting the downstream specific region of the 5′LTR; Reverse primer R1, which targets the 3′LTR specific region; Forward primer F2, which binds to the 5′ constant tail; Reverse primer R2, which is located inside the reverse primer R1; Forward primer F3, which binds to the 5′ constant tail and contains the sample barcode; Reverse primer R3, which is located inside the reverse primer R2.

[0012] Thirdly, the present invention provides a data processing system for the above-described method, characterized in that the data processing system comprises: The data splitting module is used to split the reading segments based on the sample barcode and unify the reading direction; The UMI parsing module is used to parse UMIs with constant tails as anchors. The UMI grouping and consensus module is used to perform multi-sequence consensus comparison on groups with ≥5 UMI read segments to generate a UMI consensus sequence per group. The classification and interpretation module is used to execute the following rules: The criteria for determining whether the UMI consensus sequence is completely within the target sequence are: the UMI consensus sequence continuously covers ≥95% of the length of the reference sequence, or the consensus sequence itself is continuously aligned to the reference sequence for ≥90% of its length. The rules for determining contaminated or recombinant sequences are as follows: iterative classification is performed based on 30 nt motifs of 5′Ψ and 3′WPRE, where the edit distance of a single motif is ≤3; The rule for determining non-specific sequences: sequences that do not belong to any of the above classifications; The mutation detection module is used to detect mutation events entirely within the target consensus sequence; The packaging recombination detection module is used to align contaminated or recombinant sequences to the packaging gene reference panel. If there is a continuous alignment block of ≥100bp (CIGAR M / = / X) and small indels of ≤5bp are allowed, it is considered positive. The reporting module is used to output the overall length percentage, error rate, percentage of minor components, and packaging recombination conclusions.

[0013] Fourthly, the present invention provides a computer-readable medium, characterized in that the computer-readable medium stores program instructions, which, when executed by a processor, cause a device to implement the module flow of the data processing system as described in claim 8.

[0014] Fifthly, the present invention provides the application of the above method in the assessment of the full-length integrity of integrated lentiviral vectors, the quantification of ≤0.1% minority components, and the monitoring of packaging sequence recombination.

[0015] The beneficial effects of this invention are as follows: Compared to the generic UMI-long read route that does not distinguish between on-target and off-target amplification products, this method significantly suppresses non-specific propagation and chimeras in the human genome context through single-end UMI markers and site-restricted semi / full nesting; under consensus conditions of ≥5 per UMI, the mismatch and indel event rate can be reduced to 10. -4 -10 -5 bp -1 (Approximately Q≈40), the proportion of full-length sequences completely on the target is approximately 99%, and it can stably detect and quantify a small number of components at the 0.1% level by molecular count within the same dataset, while simultaneously completing the positive determination and breakpoint localization of packaging sequence recombination; the method exhibits consistent platform-independent accuracy under Nanopore and PacBio inputs, and supports same-day turnaround. Attached Figure Description

[0016] Figure 1 (A–F) are schematic diagrams of the LUNa-seq front-end amplification process and site constraints; Figure 1 In the diagram, A: First round (UMI labeling), F1 = "5′ constant tail + UMI + 5′ LTR specific region", R1 = "3′ LTR", completing single-end UMI labeling; B: Second round (semi-nested), F2 = "binding constant tail", R2 = "located inside R1", products carrying only the constant tail can be further amplified; C: Third round (fully nested / indexed), F3 = "binding constant tail + barcode", R3 = "located inside R2"; D: Non-specific initiation interception - if the first round fails to initiate at the predetermined site (e.g., mismatch or template discontinuity), due to the lack of the constant tail and lentivirus-specific binding site, the fragment will not be amplified in subsequent rounds (marked with "X"); Control method illustration (double-end UMI) - the double-end UMI process failed to avoid non-specific amplification in the second round; E: a single master band is obtained according to the three-round nested strategy of this invention; F: lack of site constraints will produce multiple non-specific bands / tails (red arrows).

[0017] Figure 2 Here is a flowchart of the LUNa-seq data analysis process; Figure 2The process is as follows: Sample splitting → UMI parsing and orientation → Grouping by UMI → POA / equivalence consensus generation per-UMI → Alignment and classification with reference (fully on target (full length / truncated) / recombination / contamination / non-specific) → Variation analysis of fully on-target sequences / detection of a few components and molecule counting (95% CI, Clopper–Pearson) → Report (judgment criteria: ≥5 per UMI; fully on target = "covers reference ≥95%" or "self-length ≥90% continuous alignment"; packaging recombination = "continuous alignment with gag / pol / rev / VSV-G ≥100bp and ≤5bp indel"). Figure 3 A detection graph to improve read integrity and coverage uniformity for UMI consensus; Figure 3 In the diagram, A: IGV comparison chart: Raw vs UMI consensus; B: Percentage of targets with ≥95% continuous coverage (%), mean ± SD, n=6 (biological replicates), Welch's t-test, significance threshold: ns (p≥0.05); judgment criteria are the same. Figure 2 ; Figure 4 The result of reducing background errors for UMI consensus; Figure 4 In the middle, A: Mismatch (×10) -4 bp -1 B: Indel (×10) -4 bp -1 Mean ± SD; groups include pLV and gLV, Raw vs UMI consensus, n=6 (biological replicates), Welch's t-test, significance threshold: ****(p< 0.0001) ; Figure 5 A lineage diagram of SNVs before integration of viruses after the UMI consensus; Figure 5 In the middle, A: Comparison of SNV rates per base, for pLV control and gLV integrator; SNV rate after background subtraction (gLV − pLV) is expressed as the number of events per base (×10). -4 bp -1 (Mean ± Standard Deviation, n = 5): gLV: 2.856 ± 1.529 × 10 -4 bp -1 The gLV–pLV after background subtraction is: 1.036 ± 0.541 × 10⁻⁶ -4 bp -1B: Comparison of insertion / deletion rates per base, after background subtraction, units are the same: gLV: 1.082 ± 0.2896 ×10 - 4 bp -1 The gLV–pLV after background subtraction is: 0.5266 ± 0.0678 × 10⁻⁶. -4 bp -1 C: Single base mutation profile, G→A substitution accounted for 62.4%, consistent with known retroviral mutation bias. Other substitution categories included: C→T (12.6%), A→G (6.8%), T→C (6.2%), and other substitution categories (11.9%); these percentages represent the average proportion of each substitution category across five biological replicates (n = 5); the sector percentages in the pie chart represent the mean of each substitution category. Figure 6 The results of detection and quantification of Spike-in with 0.1% contamination sequence are shown in the figure. Figure 6 The table shows the proportions of each category: full-length sequence, truncated sequence, non-packaging recombination, packaging recombination, contaminated sequence, and non-specific sequence; the measured proportion of contaminated sequences is 0.243% (K=3, N=1234), with a 95% confidence interval (Clopper–Pearson) of 0.050%–0.709%; ≥5 per UMI; Figure 7 The results of localization detection of 0.1% packaged recombinant Spike-in; Figure 7 Top: Alignment coverage plot / breakpoint illustration with gag / pol / rev / VSV-G reference (≥100bp continuous alignment, ≤5bp indel); Bottom: Frequency bar (packaging recombination frequency 0.10% (K=1, N=1000; 95% confidence interval (Clopper–Pearson) is 0.00253%–0.5559%)). Figure 8 A cross-platform full-length aspect ratio comparison (Nanopore vs PacBio) result chart; Figure 8 In the target, the reference sequence is continuously covered by ≥95% (%), with a mean ± SD, and n=3; the analytical caliber and judgment threshold are the same. Figure 2 Significance threshold: ns (p≥0.05) ; Figure 9 A chart showing the results of cross-platform accuracy comparison; Figure 9 In the middle, the (A–C) mismatch Q-score (Phred, Q=-10 log) 10Error varies with minimum read threshold per UMI (≥3, ≥5, ≥10, ≥20, ≥50, ≥100 reads / UMI), for standard (n=6), ultra-long, and high GC components (n=3), respectively. Both Nanopore and PacBio platforms reach approximately Q40 at ≥5 reads / UMI, with no significant difference between groups. (D–F) Comparison of insertion / deletion Q scores under the same conditions: PacBio is slightly higher at lower thresholds (≥3); when the threshold is ≥5, both platforms approach Q40 for standard / ultra-long components and approximately ≈Q30 for high GC components, with the platform difference narrowing. Bars represent the mean, and the error bar is ±SD. Two-way ANOVA (platform × UMI threshold), multiple comparisons corrected using Benjamini–Hochberg FDR; significance threshold: ns (p≥0.05), * (p<0.05), ** (p<0.01), *** (p< 0.001)、****(p<0.0001) . Detailed Implementation

[0018] The representative flow of the method of the present invention is described below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments are for illustration and not for limiting the present invention; equivalent modifications made by those skilled in the art without departing from the concept of the present invention fall within the protection scope of the present invention.

[0019] The sequences involved in this invention are described as follows:

[0020] The corresponding gene sequence is as follows:

[0021] Example 1 PCR Procedure and UMI Introduction (1) Sample preparation Cells containing an integrated lentiviral vector (LV) were selected (HT1080 cell line, primary T cells, or CD34 cells were selected in this example). + Cells were used to extract high molecular weight genomic DNA. This embodiment uses pre / post PCR partitioning and disposable consumables throughout the process, and sets up negative control (NTC) and positive control (pLV: known LV plasmid incorporation into background gDNA) to monitor contamination and specificity.

[0022] (2) Primer and site design F1 (SEQ ID NO:1) consists of three segments: a 5′ constant tail (used for subsequent anchoring and indexing) + a UMI degenerate region (18nt in this embodiment; IUPAC: H / B / D / N) + a 5′ LTR downstream conserved region specific sequence; R1 (SEQ ID NO:2): 3′LTR conserved region specific sequence; F2 (SEQ ID NO:3): complementary / paired with the 5′ constant tail of F1; R2 (SEQ ID NO:4): located inside R1; F3 (SEQ ID NO:5-20): Paired with the 5′ constant tail of F1 and carries the sample barcode; R3 (SEQ ID NO:21): Located inside R2.

[0023] The relative positions of F2 / R2 and F3 / R3 ensure that the second / third rounds can structurally amplify only the target template carrying the constant tail / UMI, thereby suppressing the non-specific fragments generated in the first round from entering subsequent amplification (see...). Figure 1 (D "X" logo).

[0024] The conditions and objectives of the three-round nested amplification method adopted in this invention are as follows: Round 1 (UMI marker; see) Figure 1 A): Amplification with F1 / R1; 1–3 cycles (2 cycles in this example), annealing temperature set according to F1 / R1 Tm, extension 68–72℃ (68℃ in this example), 5–9 min (8 min in this example).

[0025] Objective: To introduce a single-ended UMI and constant tail in one go, and to limit the number of loops to reduce early bias and template switching.

[0026] Second round (semi-nested; see...) Figure 1 B): Amplification with F2 / R2; 16-20 cycles (20 cycles in this example), extension at 68-72℃ (68℃ in this example), 5-9 min (8 min in this example).

[0027] Objective: F2 binds to the constant tail, and R2 is located inside R1, thus only amplifying the target product carrying the constant tail / UMI, significantly inhibiting non-specific propagation.

[0028] The third round (fully nested / indexed; see...) Figure 1 C): Amplification with F3 / R3; 18–22 cycles (21 cycles were chosen in this example), extension as above.

[0029] Objective: F3 combines a constant tail and introduces a barcode, with R3 located inside R2, to complete library indexing and final amplification.

[0030] Enzyme and system: High-fidelity long-fragment polymerase and matching buffer system are preferred (KOD Multi & Epi are selected in this example; to reduce length / GC bias, the extension time is extended (e.g., ≥8 min) and compatible additives (e.g., appropriate amount of DMSO / glycerol) can be added).

[0031] Process quality control: Agarose gel electrophoresis / fragment analysis can be performed after each round. The target band should be close to the expected full length (usually 5–10 kb). No specific bands should appear in NTC; if they do, the batch should be terminated and the zoning and consumables checked. The agarose gel electrophoresis results of this invention show no visible non-specific bands, while the control method shows obvious non-specific bands (see...). Figure 1 EF).

[0032] (3) Long read sequencing and data acquisition (see Figure 1 A) Library construction and sequencing • Oxford Nanopore: Perform end repair / A addition on the third-round product (if needed, required in this example), connect the adapter and load the sequencing chip. Select a suitable protocol and chip for long amplicon sequencing; run the program to acquire raw current data and complete base identification, outputting FASTQ.

[0033] •PacBio: Prepare an SMRTbell library from the third-round product and obtain CCS reads (high-fidelity mode is selected in this embodiment).

[0034] Both platforms recommend setting up positive / negative controls and balancing sample volume / readings in the same batch.

[0035] Basic Data Requirements Regardless of the platform, the original read segment FASTQ and sample barcode information should be output; the read segment should include a locatable constant tail and UMI. To avoid anchoring failure, long homopolymers and high-order structures should be avoided in primer and adapter design, and a small-scale trial run should be performed on several representative samples before the test.

[0036] Example 2 UMI parsing and consensus building per-UMI (see...) Figure 2 ) (1) Sample splitting and segment preprocessing The data is split at the sample level according to the barcode; low-complexity / joint sequence removal and length filtering are performed on the read segments. The read segment direction is unified to 5′LTR→3′LTR; if a barcode / constant tail / UMI is detected in the reverse direction, the reverse complement is taken and the direction change is recorded.

[0037] (2) UMI positioning and fault tolerance strategy Using the constant tail sequence as an anchor, fuzzy matching (allowing a difference of ≤1nt) is performed in the neighborhood of the 5′ end of the read segment to locate the previous anchor point; a 10nt sequence after the UMI (18nt in this embodiment) is extracted from the previous anchor point with a fixed offset, and it is checked whether it is the subsequent anchor point (allowing a difference of ≤1nt). If both the previous and subsequent anchor point requirements are met, the sequence in between is determined as the UMI sequence.

[0038] (3) Grouping by UMI and threshold Reads with the same UMI are clustered together. Consensus is generated and reads are evaluated only for groups with "≥5 reads per UMI"; reads below the threshold are marked as low support and removed from downstream statistics or reported separately. This threshold balances error suppression capability with molecule utilization and is a prerequisite for subsequent Q≈40 accuracy and full-length ≈99%.

[0039] (4) Generation of consensus per UMI Within each UMI group, partial sequence ordered alignment (POA) or its equivalent multi-sequence consensus is performed to obtain the UMI consensus sequence.

[0040] After consensus was reached, the end-truncation problem in the original read segments was significantly resolved, ultimately achieving uniform coverage. This indicates that after UMI consensus processing, end-truncation of read segments was effectively eliminated, achieving a relatively consistent coverage pattern (see...). Figure 3 Through UMI clustering and consensus processing, background errors were also significantly suppressed, specifically as the mismatch rate in the original Nanopore reads (10) was reduced. -4 bp -1 The mean ± standard deviation (MS / SD) of the consensus sequence decreased from 43.83 ± 1.47 for pLV samples and 44.40 ± 1.85 for gLV samples to 2.28 ± 1.04 and 4.43 ± 2.34, respectively; the indel rate decreased from 59.38 ± 2.39 for pLV samples and 58.85 ± 3.13 for gLV samples to 0.42 ± 0.26 and 0.78 ± 0.29, respectively, providing the necessary accuracy for subsequent full-length determination / minority component quantification / packaging recombination identification. The consensus sequence and read consistency statistics will be incorporated into the subsequent classification and reporting module (see...). Figure 3 ).

[0041] (5) Chimera and pseudo-amplification control Because the site constraints in the second / third rounds only amplify the target template with constant tails / UMI, the propagation of early non-specific products is significantly suppressed. In actual operation, the chimeric read rate obtained according to this procedure can be significantly lower than the baseline of non-nested / non-selective protocols.

[0042] Note: The UMI consensus sequence set generated in this embodiment will be used as the input for the next embodiment, "Classification and Quantification" (corresponding to the threshold of claims f1 / f2 / f3: continuous alignment of the target sequence with continuous coverage of the reference ≥95% or its own length ≥90%). Packaging reassembly = continuous alignment with gag / pol / rev / VSV G ≥100bp and ≤5bp indel; Minority components = counted as UMI molecules and given Clopper–Pearson 95% CI).

[0043] This embodiment does not repeat the threshold; it only illustrates that consensus serves as a prerequisite for sufficient accuracy in interpretation and as a data interface.

[0044] Example 3 Sequence alignment and classification analysis (see Figure 2 , Figure 6 –8) Reference sequence preparation (1) Establish a reference database for interpretation, which shall include at least: a) The reference full sequence of the target lentiviral vector; b) Packaging gene panel: sequences derived from helper plasmids such as gag, pol, rev, and VSV G; c) (Optional) Host genome reference (for identifying integration junctions).

[0045] The reference sequences are derived from process design and production documents; the panel allows for equivalent version updates but retains gene boundary annotations to ensure comparability of recombination localizations.

[0046] (2) Comparison with the baseline threshold Perform long read length comparisons on each UMI consensus sequence (using tools such as minimap2 / BLASR, etc.) and apply the following default thresholds: • Comparison parameters: For example, in minimap2 v2.28, use the parameter "-a -x asm5 -z 1000,200 --secondary=no"; • Continuous Alignment Block: Calculates the length of the longest matching block that is in the same direction, referenced, and without soft / hard shearing based on CIGAR; • Target: The determination is always based on the “UMI consensus sequence” (not the original read segment).

[0047] (3) Categories and Judgment Rules 1) Completely On Target: A target is considered completely on target if either of the following conditions is met: (i) The consensus sequence continuously covers the reference length ≥ 95%; or (ii) The consensus sequence itself is continuously aligned to the carrier reference for ≥90% of its length.

[0048] •Full length: The consensus sequence continuously covers ≥95% of the reference length; •Truncation: The portion that satisfies the condition of being completely on the target but not the full-length sequence; Note: Small differences (SNV / small indel) are included in the variation statistics and do not change the f1 criterion.

[0049] 2) Perform order-guided iterative classification on the residual set that has not reached f1: • Seed identification: If the consensus simultaneously contains 5′ Ψ (30nt) and 3′ WPRE (30nt) and their respective edit distances are ≤3 and the direction is correct, then it is defined as a seed sequence; • Allocation and Iteration: The residual set is compared with the seed sequence, and the sequence belonging to the template is determined using the target sequence classification criteria. After removing the belonging entries, if a new seed sequence threshold exists, the above process is repeated to identify potential recombination / contamination sequences. Termination Handling: To prevent the program from running indefinitely, a maximum number of iterations can be set as a stopping condition.

[0050] • Host sequence extension (optional record) If one end of the consensus extends to the host sequence (across the integration junction), the vector portion is still determined according to the above rules; additional records of host coordinates and the integration sequence are provided for integration site analysis.

[0051] • Contamination: Potential recombinant / contamination sequences are manually verified using IGV. Other known sequences (such as process-related vectors or plasmids) aligned with non-target vectors / non-packaging panels are recorded as contamination sequences.

[0052] • Recombination: The read contains lentivirus-specific sequences at both ends, but an indel event of ≥50bp occurs in the middle.

[0053] Packaging recombination: There are continuous aligned blocks of ≥100bp on the packaging panel, and the total number of indels of each block is ≤5bp; record the corresponding genes, coordinates and breakpoint maps.

[0054] Non-packaging restructuring: The part of restructuring other than packaging restructuring.

[0055] 3) Non-specific: Those that do not meet any of the above categories, or whose alignment quality is below the threshold.

[0056] • Final Tags and Statistics Each consensus sequence was assigned one of the six labels mentioned above (full-length sequence / truncated / non-packaging recombination / packaging recombination / contaminated sequence / non-specific), and the proportions of each category and the 95% CI (Clopper–Pearson) were calculated. The proportion of "complete and continuous coverage ≥95% in the target" was recorded as the full-length proportion. IGV snapshots show that the end-truncation problem in the original reads was significantly resolved after consensus generation, ultimately achieving uniform coverage. This indicates that after UMI consensus processing, end-truncation of reads was effectively eliminated, achieving a relatively consistent coverage pattern. In this embodiment, read integrity was quantitatively evaluated: the proportion of "full-length reads" was used as the read integrity indicator. After UMI clustering and consensus processing, the full-length read proportions of gLV and pLV increased to 98.8% and 99.3%, respectively. (See...) Figure 3 ) In this embodiment, a specific example (0.1% contamination with Spike in) is as follows: full-length sequence 92.16%, truncation 0.39%, non-packaging recombination / structural variation 5.88%, packaging recombination 0.00%, contamination sequence 0.20%, nonspecificity 1.37% (see...) Figure 6 This caliber is used for intra-batch release and is comparable between batches.

[0057] In this embodiment, a specific example (0.1% packaged recombinant Spike in) is as follows: the recombinant sequence was detected at a rate of 0.1%, demonstrating the sensitivity of the method of the present invention, and supporting the output of a breakpoint diagram (see...). Figure 7 ).

[0058] Example 4 Low abundance variation and mutation load analysis (see Figure 5 ) (1) Callable scope and filtering SNV / small indel (≤50bp) analysis was performed entirely within the target set: • Exclude 5bp from the end of each read segment; • Use consensus-based majority (or quality-weighted) site selection and require site coverage ≥5; (2) Indicators and presentation • Mutation load: expressed as the number of variants per 10kb or the average number of variants per genome; • Lineage: Provide the transition / transversion configuration (e.g., G→A dominance) and indicate the number of samples and the total number of callable bases; • Hotspot sites: Outputs the variation frequency of sites of interest.

[0059] In this embodiment, by comparing the SNV rate per base of the pLV control group and the gLV integrin, and after background subtraction (gLV–pLV), the number of events per base (×10) was obtained. -4 bp -1 The results showed that the SNV rate of the gLV integrase was 2.856 ± 1.529 × 10⁻⁶. -4 bp -1 The gLV–pLV SNV rate after background subtraction was 1.036 ± 0.541 × 10⁻⁶. -4 bp -1 (n = 5). This result indicates a high mutational burden in the gLV integrator, and the difference after background subtraction further confirms the mutational characteristics of the pre-integration virus (see...). Figure 5 A).

[0060] Furthermore, the comparison of insertion / deletion (indel) rates showed that the indel rate for gLV was 1.082 ± 0.2896 × 10⁻⁶. -4 bp -1 The gLV–pLV indel rate after background subtraction was 0.5266 ± 0.0678 × 10⁻⁶. -4 bp -1 (n = 5). This data indicates that the gLV integrons have a significantly higher incidence of insertion / deletion mutations than the background level, further supporting the genetic characteristics of this virus (see...). Figure 5 B).

[0061] In the analysis of the single-base mutation profile, G→A substitution accounted for 62.4%, a result consistent with known retroviral mutation biases. Other substitution categories included C→T (12.6%), A→G (6.8%), T→C (6.2%), and other substitution types (11.9%). These data reflect the mutation profile characteristics of the pre-integration virus, further revealing the regularity of viral mutations (see...). Figure 5 C).

[0062] Example 5 Cross-platform validation and method robustness (see Figure 8-9 ) Nanopore and PacBio analyses were performed on samples from the same batch under a unified criteria (≥5 UMIs per UMI; f1 / f2 / f3 as above): • Integrity (complete on target and coverage ≥95%): Nanopore 96.6%, PacBio 96.0% (mean ± SD, n see Figure 8 ;ns); • Error rate (×10) -4bp -1 The mismatches of the two platforms are basically the same; Indel PacBio is lower than Nanopore, but both meet the e1 / e2 / e3 accuracy requirements (see...). Figure 9 ); • Robustness: When only 50% of the planned readings are available, the critical few components and packaging remodeling interpretations remain unchanged (point estimate fluctuations are within CI); for external noise incorporation, the pipeline classifies it as nonspecific or contamination rather than on-target events.

[0063] Unless otherwise specified, in the statistical methods used in this invention, bar charts represent mean ± SD, where n = number of biological replicates; comparisons between two groups were performed using Welch's t-test, with significance thresholds of ns (p ≥ 0.05), * (p < 0.05), ** (p < 0.01), *** (p < 0.001), and **** (p < 0.0001). Consistent with the figure captions.

[0064] Example 6 A LUNa-seq method The core of this method lies in a closed-loop process of single-end UMI tagging + (semi-nested → fully nested) selective re-amplification + per-UMI consensus + triple interpretation caliber, specifically including: 1. Single-ended UMI tagging and three-round nested amplification: First round (UMI tagging): The UMI + target 5'LTR sequence is attached to the 5' constant tail of the forward primer F1 and amplified with the 3'LTR reverse primer R1, so that each target molecule obtains a single-end UMI with a constant tail; Second round (semi-nested): Using F2 (combined with the constant tail) and R2 (located inside R1), only the on-target template carrying the constant tail / UMI is amplified again, suppressing the non-specific propagation generated in the first round; The third round (fully nested / indexed): using F3 (including sample barcodes, combined with constant tails) and R3 (located inside R2) to complete indexing and library construction.

[0065] 2. Consensus on Long Read Sequencing and UMI-by-UMI: Sequencing was performed on Nanopore or PacBio platforms, and UMIs were resolved based on constant tail localization. After grouping by UMI, multiple sequence consistency calculations were performed only on groups with ≥5 reads per UMI to obtain per-UMI consensus sequences, in order to reduce the impact of mismatches / indels and end truncation on read interpretation.

[0066] 3. Triple interpretation criteria (for the same dataset): Full-length integrity: The consensus sequence is compared with the carrier reference. If it continuously covers the reference length ≥95%, or the consensus sequence itself is ≥90% of the length, it is considered to be completely on target. Among them, the sequence that continuously covers the reference length ≥95% is defined as full-length sequence. Packaging sequence recombination monitoring: A sequence with a continuous alignment block of ≥100bp and ≤5bp indel with the gag / pol / rev / VSV G reference is considered positive, and the coordinates and breakpoint map are output. Quantification of a few components: Molecular counting was performed using the UMI group number and Clopper–Pearson 95% confidence intervals were given.

[0067] 4. Primer / Reagent Kit and System Implementation: Provide primer compositions (F1: 5′ constant tail + UMI + 5′ LTR specific region; R1: 3′ LTR specific region; F2 / F3 binding constant tail / barcode site; R2 / R3 located inside the outer primer, SEQ ID NO: 1-21) and kits (KOD Multi&Epi (Toyobo, KME-101)) that match the above process, as well as a data processing system or computer-readable medium (https: / / github.com / yangyang045200 / LUNa-seq) containing modules such as sample splitting, UMI parsing and orientation, UMI grouping and consistency, classification interpretation and variant detection.

[0068] Without departing from the technical solution of this invention, the method provided by this invention can be adapted to different LV backbone constructs (5–10kb), different cell sources (such as HT1080, primary T cells, CD34⁺ cells), and other integrated systems (γ retroviruses, transposons, etc.); POA consensus and long read alignment tools can be replaced with equivalent implementations; UMI degeneracy mode and amplification conditions can be optimized according to the range listed in the sequence listing and examples.

Claims

1. A long-read sequencing method for quality control of integrated lentiviral vectors, characterized in that, The method includes the following steps: a) Perform the first round of PCR, using forward primer F1 and reverse primer R1 to amplify the long fragment product, wherein the forward primer F1 contains a 5′ constant tail, a unique molecular identifier (UMI) sequence and a sequence targeting the 5′ LTR, and the reverse primer R1 targets the 3′ LTR sequence. b) Perform a second round of PCR using forward primer F2 and reverse primer R2 for semi-nested amplification, wherein forward primer F2 binds to the 5′ constant tail and reverse primer R2 is located inside the reverse primer R1, thereby selectively amplifying the product carrying the constant tail and UMI. c) Perform a third round of PCR using forward primer F3 and reverse primer R3 for fully nested amplification, wherein forward primer F3 contains the sample barcode and binds the 5′ constant tail, and reverse primer R3 is located inside reverse primer R2 to obtain indexed amplicon; d) Perform long-read sequencing on the indexed amplicon to obtain sequencing reads; e) Based on the 5′ constant tail location parsing UMI, the read segments are grouped according to UMI, and multi-sequence consistency calculation is performed on the groups with ≥5 read segments per UMI to generate a per-UMI consensus sequence; f) Compare the UMI consensus sequence with the carrier reference sequence and classify them based on the following rules: f1) Completely in the target sequence: The UMI consensus sequence continuously covers the reference sequence length ≥ 95%, or the consensus sequence itself is continuously aligned to the reference sequence for ≥ 90% of its length; f2) Packaging and Reassembly Sequence: The UMI consensus sequence has a continuous alignment block of ≥100bp (CIGARM / = / X) with the gag / pol / rev / VSV G reference and allows small indels of ≤5bp; f3) Contaminated sequences or non-packaged recombination sequences: Iterative classification is performed based on 30 nt motifs of 5′Ψ and 3′WPRE, where the edit distance of a single motif is ≤3; f4) Non-specific sequences: sequences that do not belong to any of the above categories; g) Output the full-length proportion, minority component proportion, and packaging recombination conclusions based on UMI molecular counting.

2. The method according to claim 1, characterized in that, The number of cycles for the first round of PCR is 1-3.

3. The method according to claim 2, characterized in that, The length of the UMI is 18 nt, and the degenerate mode is HBNHVNBDNHVNBDNHBD.

4. The method according to claim 3, characterized in that, The multi-sequence consistency calculation uses partial sequence ordered alignment.

5. The method according to claim 4, characterized in that, The proportions of the minority components are given a 95% confidence interval using the Clopper-Pearson method.

6. The method according to claim 5, characterized in that, The long-read sequencing is performed using Oxford Nanopore or PacBio sequencing, and the mismatch event rate is no higher than 2 × 10⁻⁵ reads per UMI. -4 bp -1 The insertion or missing event rate is no higher than 8 × 10⁻⁶. -5 bp -1 .

7. A primer composition for use in the method according to any one of claims 1-6, characterized in that, The primer composition comprises: Forward primer F1, which contains a 5′ constant tail, a UMI sequence, and a sequence targeting the downstream specific region of the 5′LTR; Reverse primer R1, which targets the 3′LTR specific region; Forward primer F2, which binds to the 5′ constant tail; Reverse primer R2, which is located inside the reverse primer R1; Forward primer F3, which binds to the 5′ constant tail and contains the sample barcode; Reverse primer R3, which is located inside the reverse primer R2.

8. A data processing system for the method according to any one of claims 1-6, characterized in that, The data processing system includes: The data splitting module is used to split the reading segments based on the sample barcode and unify the reading direction; The UMI parsing module is used to parse UMIs with constant tails as anchors. The UMI grouping and consensus module is used to perform multi-sequence consensus comparison on groups with ≥5 UMI read segments to generate a UMI consensus sequence per group. The classification and interpretation module is used to execute the following rules: The criteria for determining whether the UMI consensus sequence is completely within the target sequence are: the UMI consensus sequence continuously covers ≥95% of the length of the reference sequence, or the consensus sequence itself is continuously aligned to the reference sequence for ≥90% of its length. The rules for determining contaminated or recombinant sequences are as follows: iterative classification is performed based on 30 nt motifs of 5′Ψ and 3′WPRE, where the edit distance of a single motif is ≤3; The rule for determining non-specific sequences: sequences that do not belong to any of the above classifications; The mutation detection module is used to detect mutation events entirely within the target consensus sequence; The packaging recombination detection module is used to align contaminated or recombinant sequences to the packaging gene reference panel. If there is a continuous alignment block of ≥100bp (CIGAR M / = / X) and small indels of ≤5bp are allowed, it is considered positive. The reporting module is used to output the overall length percentage, error rate, percentage of minor components, and packaging recombination conclusions.

9. A computer-readable medium, characterized in that, The computer-readable medium stores program instructions that, when executed by a processor, cause the device to implement the modular flow of the data processing system as described in claim 8.

10. The application of the method as described in any one of claims 1-6 in the assessment of full-length integrity of integrated lentiviral vectors, quantification of ≤0.1% minority components, and monitoring of packaging sequence recombination.