HPV DNA integration event detection method, system, equipment and medium

By employing a hybrid sequencing detection method guided by HPV genotyping, non-overlapping DNA libraries are constructed and high-throughput sequencing is performed. This solves the problems of high cost and low throughput in HPV integration detection, achieving low-cost and high-efficiency detection of HPV DNA integration events with broad-spectrum capture capability and cost-effectiveness.

CN121109657APending Publication Date: 2025-12-12NANJING DRUM TOWER HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511291688.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-08-10
Filing Date
2025-09-10
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing integrated HPV testing methods are costly and have low throughput, making them difficult to apply on a large scale for screening and risk stratification of HPV-related cancers.

Method used

By reusing clinical pre-screening data and using HPV genotyping-guided hybrid sequencing detection methods, a DNA library containing non-overlapping HPV genotypes was constructed. This library was then hybridized with HPV whole-genome probes for high-throughput sequencing, using HPV genotypes as natural identifiers to trace HPV DNA integration events.

Benefits of technology

It achieves low-cost, high-throughput detection of HPV DNA integration events, reducing the cost of single-sample testing by 60%, and improving cost-effectiveness in large-scale HPV genotyping and integration detection. It has significant non-target hybridization efficiency and broad-spectrum capture capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121109657A_ABST
    Figure CN121109657A_ABST
Patent Text Reader

Abstract

The invention discloses a method, a system, equipment and a medium for detecting an HPV (human papilloma virus) DNA (deoxyribonucleic acid) integration event, and relates to the technical field of gene detection. According to the method, HPV types and virus loads (Ct values) pre-screened by qPCR are used as natural identifiers to guide grouping and mixing of samples, and HPV integrated event detection and sample tracing are synchronously realized through single sequencing. Clinical verification shows that compared with qPCR, the comprehensive sensitivity reaches 97.1%, and the cost is reduced by 60%. And a built-in targeted retest mechanism improves the reliability of a detection result. The invention provides a novel, traceable, economical and feasible solution, is used for high-throughput HPV-host genome integration feature analysis, balances cost, expandability and clinical accuracy, and provides a high-throughput and low-cost solution for HPV-related cancer screening.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of clinical molecular diagnosis, and particularly relates to a detection method, system, device and medium for HPV DNA integration event. BACKGROUND

[0002] Persistent infection with high-risk HPV (hrHPV) is a core cause of malignancies such as cervical cancer. Integration of HPV DNA into the host genome marks the beginning of malignant transformation and is a key biomarker for risk stratification. Current clinical integration detection faces two major bottlenecks:

[0003] 1. Technical bottleneck: traditional methods (such as DIPS-PCR) have low throughput and limited sites; NGS can achieve whole genome integration analysis, but the cost of a single sample is too high (about $200 / sample), hindering large-scale screening applications;

[0004] 2. Identification bottleneck: existing hybrid sequencing schemes rely on synthetic barcodes (such as CN104573407B) or host gene markers, requiring additional library construction steps and not suitable for HPV clinical pathways.

[0005] Persistent infection with high-risk human papillomavirus (hrHPV) is an important cause of the development of cervical cancer and head and neck cancer. Integration of HPV DNA into the host genome marks a key transition from transient infection to malignant transformation, which is closely related to disease progression and treatment resistance. Unlike transient HPV infection, which usually resolves, integration events are considered markers of transforming infection and are closely related to genomic instability and cancer development. Current clinical guidelines increasingly emphasize the need for biomarkers to classify HPV-positive patients into different risk categories, with integration status as a prognostic indicator for targeted therapy. Therefore, identifying HPV integration status can be a potential tool to distinguish truly high-risk individuals and reduce unnecessary diagnostic procedures and overtreatment. However, existing integration detection methods are limited in their application to large-scale screening and risk stratification due to technical complexity, site specificity, or low throughput.

[0006] Traditionally, the evaluation of HPV integration mainly relies on methods such as DIPS-PCR and APOT, which are not only time-consuming, but also specific to certain gene sites, making it difficult to be applied to clinical large-scale. With the development of high-throughput next-generation sequencing (NGS) technology, it is now possible to conduct comprehensive analysis of viral integration at the whole genome range. By using HPV whole genome probes for targeted capture sequencing, HPV sequences can be enriched, and integration breakpoints in the human genome can be detected. This method not only enables accurate genotyping of HPV, but also systematically maps integration events, thereby providing a more comprehensive understanding of the interaction between HPV and the host. However, the high cost of NGS per sample is still the main obstacle to its application in population-level HPV screening or large-scale epidemiological studies.

[0007] Therefore, the pain point to be solved by the present application is: how to reuse existing HPV typing data (qPCR output) as a natural identifier to achieve "zero additional cost" mixed sample detection. SUMMARY

[0008] (I) Technical problems solved

[0009] In view of the technical bottleneck of high cost and low throughput of existing HPV integration detection methods (such as single-sample NGS), the present application provides a detection method, system, device and medium for HPV DNA integration events. By reusing clinical pre-screening data, low-cost high-throughput detection is achieved, with a 60% reduction in single-sample detection cost under the premise of ensuring 97.1% comprehensive sensitivity, and the reliability is improved through targeted retesting mechanism, providing an economic and efficient solution for HPV-related cancer screening.

[0010] (II) Technical solutions

[0011] To achieve the above-mentioned purpose, the present application provides a detection method for HPV DNA integration events for non-diagnostic purposes, which is a mixed sequencing detection method guided by HPV genotyping, comprising the following steps:

[0012] Data reuse: obtain the HPV genotype and viral load of the sample to be tested;

[0013] Intelligent grouping: construct a DNA library of the sample to be tested containing HPV genotypes that do not overlap with each other;

[0014] Mixed sequencing: hybridize and capture the DNA library with HPV whole genome probes, and perform high-throughput sequencing;

[0015] Integration analysis: use HPV genotype as a natural identifier to trace the HPV DNA integration event.

[0016] In one embodiment, whether the sample to be tested is a sample with HPV DNA integration event is judged based on the following criteria:

[0017] If (1) the integration breakpoint detected by SVdetect software is supported by split-reads≥3 and / or discordant read-pairs≥5; (2) verified by CREST algorithm and manually checked the consensus sequence at the integration site, the virus-host junction sequence is clear and unique alignment, then the sample to be tested is the sample of HPV DNA integration event;

[0018] Otherwise, the sample to be tested is the sample without HPV DNA integration event.

[0019] In one embodiment, the maximum viral load of the sample to be tested in the DNA library is 1-20 times the minimum value.

[0020] In one embodiment, the viral load calculation formula is: Copy=(1 / 2)Ct.

[0021] In one embodiment, the Ct value range is 12-35.

[0022] In one embodiment, the DNA library contains 2-10 samples to be tested.

[0023] In one embodiment, the DNA library contains 2-4 samples to be tested.

[0024] In one embodiment, the detection method further comprises error correction verification: when high-throughput sequencing does not detect the HPV genotype of a sample to be tested, the sample is detected alone.

[0025] In one embodiment, the HPV genotype includes one or a combination of HPV16, 18, 31, 33, 35, 39, 45, 51, 52, 56, 58, 59, 66 or 68.

[0026] In one embodiment, the HPV genotype is detected by one or more of the following detection methods: reverse dot blot (RDB), real-time fluorescent quantitative PCR, multiplex PCR+capillary electrophoresis or gene chip.

[0027] In one embodiment, the sample to be tested refers to a composition obtained or derived from a target subject, which contains cell entities and / or other molecular entities to be characterized and / or identified, for example, based on physical, biochemical, chemical and / or physiological characteristics.

[0028] In one embodiment, the sample to be tested includes, but is not limited to, tissue, blood, serum, plasma, blood-derived cells, lymph fluid, synovial fluid, cerebrospinal fluid, pleural fluid, peritoneal fluid, bladder irrigant, secretion (e.g., breast secretion), mouth rinse, swab (e.g., mouth swab), touch preparation, fine needle aspirate, cell extract, and combinations thereof.

[0029] In one embodiment, the sample to be tested is derived from the subject's cervical exfoliated cells.

[0030] In one embodiment, the subject includes mammals and non-mammals.

[0031] In one embodiment, examples of the mammal include, but are not limited to, any member of the mammalian class: humans, non-human primates such as chimpanzees and other apes and monkey species; farm animals such as cows, horses, sheep, goats, swine; domestic animals such as rabbits; dogs and cats; laboratory animals including rodents such as rats, mice and guinea pigs, and the like. Examples of non-mammals include, but are not limited to, birds, fish or other non-mammalian animals, and the like.

[0032] In one embodiment, the subject is a human.

[0033] In one embodiment, the subject is healthy.

[0034] In one embodiment, the subject is non-healthy.

[0035] In one embodiment, the subject's cervical exfoliated cells are HPV genotype negative.

[0036] In one embodiment, the subject's cervical exfoliated cells are HPV genotype positive.

[0037] In one embodiment, the HPV DNA belongs to one or a combination of HPV16, 18, 31, 33, 35, 39, 45, 51, 52, 56, 58, 59, 66, 68, 6, 11, 26, 30, 32, 34, 40, 42, 43, 44, 53, 54, 55, 61, 62, 67, 70, 71, 73, 74, 81, 82, 83, 84, 86, 87, 89, 90, 91, and 101.

[0038] In one embodiment, the detection method specifically includes:

[0039] Data multiplexing: obtaining HPV genotype and viral load of the sample to be tested;

[0040] Intelligent grouping: constructing DNA library of the sample to be tested containing non-overlapping HPV genotypes;

[0041] Hybrid sequencing: the DNA library is hybridized with HPV whole genome probes for capture, and high-throughput sequencing is performed;

[0042] Integration analysis: using HPV genotype as a natural identifier, trace the HPV DNA integration event;

[0043] Error correction verification: when high-throughput sequencing does not detect the HPV genotype of a certain sample to be tested, the sample is detected alone.

[0044] In one embodiment, the detection method specifically comprises:

[0045] Data multiplexing: obtain the HPV genotype and viral load of the sample to be tested, and the viral load calculation formula is: Copy=(1 / 2)Ct;

[0046] Intelligent grouping: 2-4 samples to be tested are mixed to construct a DNA library, and the DNA library only contains samples with non-overlapping HPV genotypes and / or the maximum viral load of the sample to be tested in the DNA library is 1-20 times the minimum value;

[0047] Hybrid sequencing: the DNA library is hybridized with HPV whole genome probes for capture, and high-throughput sequencing is performed;

[0048] Integration analysis: using HPV genotype as a natural identifier, trace the HPV DNA integration event;

[0049] Error correction verification: when high-throughput sequencing does not detect the HPV genotype of a certain sample to be tested, the sample is detected alone.

[0050] In another aspect, the present application provides a detection system for HPV DNA integration events, comprising:

[0051] Data multiplexing module: obtain the HPV genotype and viral load of the sample to be tested;

[0052] Intelligent grouping module: construct a DNA library containing samples to be tested with non-overlapping HPV genotypes;

[0053] Hybrid sequencing module: the DNA library is hybridized with HPV whole genome probes for capture, and high-throughput sequencing is performed;

[0054] Integration analysis module: using HPV genotype as a natural identifier, trace the HPV DNA integration event;

[0055] Error correction verification module: when high-throughput sequencing does not detect the HPV genotype of a certain sample to be tested, the sample is detected alone.

[0056] In one embodiment, the maximum viral load of the sample to be tested in the DNA library is 0-20 times the minimum value.

[0057] In one embodiment, the viral load calculation formula is: Copy = (1 / 2)Ct.

[0058] In one embodiment, the Ct value ranges from 12 to 35.

[0059] In one embodiment, the DNA library comprises 2-10 samples to be tested.

[0060] In one embodiment, the DNA library comprises 2-4 samples to be tested.

[0061] In one embodiment, the detection method further comprises error correction verification: when high-throughput sequencing does not detect the HPV genotype of a sample to be tested, the sample is detected alone.

[0062] In one embodiment, the HPV genotype includes one or a combination of HPV16, 18, 31, 33, 35, 39, 45, 51, 52, 56, 58, 59, 66 or 68.

[0063] In one embodiment, the HPV genotype is detected by one or more of reverse dot blot (RDB), real-time fluorescent quantitative PCR, multiplex PCR + capillary electrophoresis or gene chip.

[0064] In another aspect, the present application also provides an electronic device comprising a memory and a processor, wherein the memory is used to store program instructions; and the processor is used to call the program instructions, and when the program instructions are executed, the above-mentioned detection method is realized.

[0065] In another aspect, the present application also provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned detection method is realized.

[0066] (III) Beneficial effects

[0067] The present application provides a detection method, system, device and medium for HPV DNA integration events. Compared with the prior art, the present application has the following beneficial effects:

[0068] 1. The detection method provided by the present application identifies 45 unique genotypes, which indicates that there is significant non-target hybridization efficiency or possible cross-reaction due to the conserved regions between HPV types. This highlights the broad capture ability of the hybrid capture sequencing method, and implies the potential for epidemiological monitoring outside the predefined high-risk targets.

[0069] 2. Human papillomavirus genotypes detected by NGS were compared to qPCR reference results in 175 pooled samples. Results showed that 135 cases (77.1%) were completely concordant, 34 cases (19.4%) were partially missing, 6 cases (3.4%) were completely missing, and no over-detection was found ( Figure 3 A). Cohen's Kappa coefficient was 0.696 (p<0.001), indicating moderate to high agreement between the two methods, demonstrating that NGS technology based on pooled samples can reliably replicate the results of qPCR genotyping in most cases.

[0070] 3. Compared to standard single-sample NGS, costs can be reduced by up to 60%. These results suggest that HPV Pool-Seq has the potential to improve cost-effectiveness in large-scale HPV genotyping and integration detection applications. BRIEF DESCRIPTION OF DRAWINGS

[0071] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the description of the embodiments or the prior art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0072] Figure 1 is a schematic diagram of the prediction method.

[0073] Figure 2 is a schematic diagram of the prediction program.

[0074] Figure 3 is a prediction result verification diagram.

[0075] Figure 4 is a cost-effectiveness schematic diagram of the prediction method.

[0076] Figure 5 is a sample analysis result diagram.

[0077] Figure 6 is a factor analysis diagram of the prediction method.

[0078] Figure 7 is a sequencing quality control indicator diagram of different batches. DETAILED DESCRIPTION

[0079] ​In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts should fall within the protection scope of the present application.

[0080] Terms and definitions

[0081] As used herein, the term "test sample" refers to a composition obtained or derived from a subject of interest that comprises cellular entities and / or other molecular entities to be characterized and / or identified, e.g., based on physical, biochemical, chemical, and / or physiological characteristics. The sample can be obtained from a subject's blood and other fluid samples and tissue samples of biological origin, such as biopsy tissue samples or tissue cultures or cells derived therefrom. The source of the tissue sample can be solid tissue, such as from a fresh, frozen and / or preserved organ or tissue sample, a biopsy tissue or aspirate; blood or any blood component; a body fluid; cells from an individual at any time of gestation or development; or plasma.

[0082] As used herein, the terms "host" and "subject" are used interchangeably. In one embodiment, the subject includes mammals and non-mammals.

[0083] Examples of mammals include, but are not limited to, any member of the mammalian class: humans, non-human primates such as chimpanzees and other apes and monkey species; farm animals such as cows, horses, sheep, goats, swine; domestic animals such as rabbits, dogs, and cats; laboratory animals including rodents, such as rats, mice and guinea pigs, and the like. Examples of non-mammals include, but are not limited to, birds, fish or other non-mammalian animals, and the like.

[0084] In one embodiment, the subject is a human

[0085] As used herein, the term "different HPV genotypes" refers to at least one HPV genotype, e.g. at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten or more genotypes. HPV genotypes that can be diagnosed in the context of the present application are preferably the alpha genus mucosal HPV genotypes 6, 11, 13, 16, 18, 26, 30, 31, 32, 33, 34, 35, 39, 42, 43, 44, 45, 51, 52, 53, 54, 55, 56, 58, 59, 61, 62c, 64, 66, 67, 68 (subtypes A and B), 69, 70, 71 72, 73, 74, 81, 82 (including IS39 and MM4), 83, 84, 85c, 86c, 87c, 89c / Cp6108, 90c, 97, 102 and 106; the alpha genus cutaneous HPV genotypes 2, 3, 7, 10, 27, 28, 29, 40, 57, 77, 91c and 94, more preferably 6, 11, 13, 16, 18, 26, 30, 31, 32, 33, 35, 39, 42, 43, 44, 45, 51, 52, 53, 54, 56, 58, 59, 66, 67, 68 (subtypes A and B), 69, 70, 72, 73, 82 (including IS39 and MM4), 89c / Cp6108, most preferably the putative high-risk HPV genotypes 26, 53, 66 and the high-risk HPV genotypes 16, 18, 31, 33, 35, 39, 45, 51, 52, 56, 58, 59, 68, 73 and 82. Preferably different HPV genotypes are predicted simultaneously.

[0086] As used herein, the term "viral load" preferably refers to the amount of HPV DNA in a sample of a subject. The assessment of the viral load of a specific HPV genotype present in a sample is preferably achieved by determining the amount of amplified polynucleotides specific for the different HPV genotypes. Methods for determining the amount of polynucleotides are well known in the art and described, e.g., in the Examples. Preferred methods for determining the amount of polynucleotides are real-time PCR, real-time RT-PCR, real-time NASBA, or signal amplification methods, including hybrid capture II, bDNA, rolling circle amplification (RCA) and dendrimer. By determining the viral load, the severity of the HPV infection can be assessed.

[0087] As used herein, the term "assessing the severity of an HPV infection" preferably means distinguishing between a severe form of an HPV 16 infection and a mild form of an HPV 16 infection. As will be appreciated by the skilled person, such an assessment is not intended to be 100% correct for the subject to be diagnosed. However, the term requires that a statistically significant fraction of the subjects can be assessed correctly. Whether a fraction is statistically significant can be determined by the skilled person using a variety of well-known statistical evaluation tools, such as determining confidence intervals, p-value determination, Student's t-test, Mann-Whitney test, and the like, without undue burden. See Dowdy and Wearden, Statistics for Research, John Wiley & Sons, New York 1983. Preferred confidence intervals are at least 90%, at least 95%, at least 97%, at least 98%, or at least 99%. P-values are preferably 0.1, 0.05, 0.01, 0.005, or 0.001. Preferably, the probabilities contemplated by the present application are such that the assessment is correct for at least 60%, at least 70%, at least 80%, or at least 90% of the subjects in a given group or cohort.

[0088] As used herein, the term "mild form of an HPV infection" preferably means a form of HPV infection which is histologically classified as normal cervical tissue or CINl (minimal or mild cervical dysplasia), or which is cytologically classified as NIL / M (no intraepithelial lesion or malignancy) or LSIL (low-grade squamous intraepithelial lesion). Thus, a mild form of an HPV infection preferably encompasses benign cervical lesions.

[0089] As used herein, the term "severe form of an HPV infection" preferably means a form of HPV infection which is histologically classified as CIN2 (moderate cervical epithelial dysplasia) or CIN3 (severe cervical dysplasia) or cancerous (in situ or invasive). Thus, the term "severe form of an HPV 16 infection" preferably means a form of HPV 16 infection which is cytologically classified as HSIL (high-grade squamous intraepithelial lesion) or cancerous. Thus, a severe form of an HPV infection preferably encompasses malignant cervical lesions.

[0090] As used herein, "comprising" or "containing" or "including" or "having" includes "consisting essentially of" and "consisting of"; "consisting essentially of" and "consisting of" are subsumed into "comprising" or "containing" or "including" or "having".

[0091] The experimental methods used in the following examples are routine methods unless otherwise specified, and the reagents, methods and apparatus used are routine reagents, methods and apparatus in the art unless otherwise specified.

[0092] The present application provides a prediction method for predicting HPV DNA integration host genome based on HPV typing guided hybrid sequencing, which is used for comprehensive detection of HPV integration events. The core of the method is to use the intrinsic specificity of HPV genotype as a biological barcode, so as to realize sample mixing without molecular index. First, HPV positivity is confirmed by quantitative PCR (q-PCR), and the genotype is determined, which is a routine step before any subsequent integration analysis. According to the HPV genotype and viral load (Ct value), samples of different genotypes are grouped into mixed samples to balance the sequencing depth and reduce the interference between samples. After completing the target capture and whole genome sequencing of HPV, a computational decoding algorithm is used to reassign the integration events to each sample through the genotype track and alignment score. Figure 1

[0093] Example 1 Screening of samples to be tested:

[0094] From July 2022 to July 2024, women who received cervical screening (including cytological examination or HPV detection) at Nanjing Drum Tower Hospital and Nanjing Maternal and Child Health Hospital were prospectively included in the study. After routine cytological examination (ThinPrep cytological test, TCT) or human papillomavirus detection, residual cervical exfoliated cells in HPV positive samples were collected. One or more of the 14 human papillomavirus types were detected in all samples. The study obtained informed consent and was approved by the institutional ethics committee.

[0095] The basic information of the subjects is shown in Table 1.

[0096] Table 1

[0097]

[0098]

[0099] Example 2 DNA extraction and quantitative PCR for HPV typing and viral load estimation:

[0100] DNA was extracted from cervical exfoliated cells using the TIANamp Genomic DNA Kit from Tiangen (Beijing, China). The fluorescent quantitative PCR (q-PCR) detection method based on TaqMan probe technology (HPV DNA detection kit, Hongmeng Microbiology Technology, Jiangsu, China) was used for genotyping and quantitative analysis of 14 human papillomavirus (14hrHPV), including types 16, 18, 31, 33, 35, 39, 45, 51, 52, 56, 58, 59, 66 and 68.

[0101] ​Amplification and result interpretation were performed according to the manufacturer's instructions. Results were considered positive when Ct value was < 38 and amplification curve was S-shaped. The relationship between Ct value and viral copy number was roughly calibrated using a mixed plasmid standard containing 14 human papillomavirus types as reference. Viral copy number was estimated by Ct value using the formula Copy = 1 / 2Ct, assuming 100% PCR amplification efficiency.

[0102] This method provides a relative comparison, not an absolute quantification. In order to improve the normality of statistical analysis, the data were log-transformed.

[0103] Sample pooling strategy and library preparation of Example 3:

[0104] Samples were grouped according to different HPV genotypes detected by qPCR. Each mixed pool (DNA library) was composed of samples infected with non-overlapping HPV genotypes. The volume of mixed pool was adjusted according to Ct value to ensure the uniform distribution of HPV viral load in each pool. Each mixed pool was processed according to standard Illumina-compatible protocol for library construction as a single sample.

[0105] Example 1: Pooling two samples

[0106] (1) Select samples of different infection types, calculate viral load;

[0107] Sample A: detected as HPV16 positive, Ct value = 21.11.

[0108] Sample B: detected as HPV51 positive, Ct value = 22.33.

[0109] The same batch of HPV16 type reference was detected with the same amount of loading detection, 100000 copies detected ct value = 22.87.

[0110] The same batch of HPV51 type reference was detected with the same amount of loading detection, 1000 copies detected ct value = 30.53.

[0111] Viral load calculation: according to the formula △ct = Ct_reference - Ct_sample.

[0112] Viral load of sample A = 2^△ct*100000.

[0113] Viral load of sample B = 2^△ct*1000.

[0114] Therefore:

[0115] Viral load of sample A = 2^(22.87-21.11) = 2^1.67*100000 = 338698.

[0116] Viral load of sample B = 2^(30.53-22.33) = 2^8.20*1000 = 294067.

[0117] (2) Mixing strategy:

[0118] Volume adjustment: To make the contribution of the two samples to the viral genome in the mixed pool similar, adjust the input DNA volume so that the product of "volume * relative viral load" is equal.

[0119] (The sample with lower viral load in the pool is used as the baseline for calculating the volume ratio, and the volume of DNA input by each sample in the mixed pool is calculated).

[0120] In the above samples A and B, the sample with lower viral load is sample B, and the volume ratio of sample B is set as the baseline, i.e. BDNA (DNA of sample B) volume ratio = 1.

[0121] According to: Viral load of sample A * ADNA (DNA of sample A) volume = Viral load of sample B * BDNA volume.

[0122] Then the ADNA volume ratio of sample A = 294067 / 338698 = 0.87.

[0123] In actual operation, the total volume after mixing is set to 20ul.

[0124] Therefore,

[0125] The BDNA volume of sample B = 20 / (ADNA volume ratio + BDNA volume ratio) = 20 / (1+0.87) = 10.7ul.

[0126] Then the ADNA volume of sample A = BDNA volume * 0.87 = 9.3ul.

[0127] Actual operation: Mix 9.3ul of genomic DNA of sample A with 10.7ul of genomic DNA of sample B, and mix well by vortexing to form a mixed pool.

[0128] Example 2: Mixing three samples

[0129] (1) Select samples of different infection types and calculate viral load

[0130] Sample A: detected as HPV16 positive, Ct value = 27.20.

[0131] Sample B: detected as HPV51 positive, Ct value = 23.23.

[0132] Sample C: detected as HPV39 positive, Ct value = 26.37.

[0133] The same batch of HPV16 type reference product was detected with the same amount of sample, and the ct value of 100000 copies was detected. =27.11.

[0134] The same batch of HPV51 type reference product was detected with the same amount of sample, and the ct value of 1000 copies was detected. =29.16.

[0135] The same batch of HPV39 type reference product was detected with the same amount of sample, and the ct value of 1000 copies was detected. =32.85.

[0136] Viral load calculation: according to the formula △ct=Ct_ reference-Ct_sample.

[0137] Viral load of sample A=2^△ct*100000.

[0138] Viral load of sample B=2^△ct*1000.

[0139] Viral load of sample C=2^△ct*1000.

[0140] Therefore:

[0141] Viral load of sample A=2^(27.11-27.20)=2^-0.09*100000=93925.

[0142] Viral load of sample B=2^(29.16-23.23)=2^5.93*1000=60969.

[0143] Viral load of sample C=2^(32.85-26.37)=2^6.48*1000=89264.

[0144] (2) Mixed strategy:

[0145] Volume adjustment: In order to make the viral genome contribution of two samples in the mixed pool similar, adjust the input DNA volume to make the product of "volume*relative viral load" equal.

[0146] (With the sample with lower viral load in the pool as the baseline for calculating the volume ratio, calculate the input DNA volume of each sample in the mixed pool).

[0147] Among the above samples A, B and C, the sample with lower viral load is sample B, and the volume ratio of sample B is set as the baseline, that is, BDNA volume ratio=1.

[0148] According to: viral load of sample A*sample A DNA volume=viral load of sample B*sample B DNA volume=viral load of sample C*sample C DNA(C DNA) volume.

[0149] The sample ADNA volume ratio = 60969 / 93925 = 0.65.

[0150] The sample CDNA volume ratio = 60969 / 89264 = 0.68.

[0151] In actual operation, the total volume after mixing is set to 20ul.

[0152] Therefore, the sample BDNA volume = 20 / (ADNA volume ratio + BDNA volume ratio + CDNA volume ratio) = 20 / (1 + 0.65 + 0.68) = 8.5ul.

[0153] Therefore, the sample ADNA volume = BDNA volume * 0.65 = 5.6ul.

[0154] The sample CDNA volume = BDNA volume * 0.68 = 5.9ul.

[0155] Actual operation: Take 5.6ul of genomic DNA of sample A, 8.5ul of genomic DNA of sample B, and 5.9ul of genomic DNA of sample C, mix them well, and constitute a mixed pool.

[0156] Example 4 Targeting HPV capture and next-generation sequencing:

[0157] The HPV probe is designed according to the full-length genome of 32 HPV types, provided by MyGenostics, Baltimore, Maryland, USA, for hybrid capture enrichment.

[0158] After hybrid capture of the DNA library constructed in Example 2, paired-end sequencing was performed on the Illumina platform of Illumina, San Diego, California, USA. The enriched library was sequenced on the Illumina platform (NovaSeq 6000) using 150bp paired-end reads (PE150), which provided sufficient coverage and resolution for mapping of integration sites.

[0159] Example 5 Sequencing data processing and HPV integration analysis:

[0160] Based on the human genome hg38 database and 514 reference genomes of 32 HPV types, bioinformatics analysis was performed.

[0161] First, the sequencing data obtained in Example 4 was preprocessed to remove adapter sequences, low-quality reads, and short sequences with length less than 40 base pairs.

[0162] Subsequently, the pre-processed data was aligned to all HPV reference genomes using BWA. Data with average sequencing depth over 10 and coverage over 45% at >4x depth were filtered to obtain HPV subtype information. According to the subtype results, the best HPV reference genome was selected and the data was aligned to hg38 and the selected reference genome using BWA to generate BAM files.

[0163] Structural variations were analyzed by SVdetect to identify the preliminary integration site region. The analysis method includes the following steps:

[0164] (1) The sequence alignment file (BAM format) obtained after alignment was input into SVdetect;

[0165] (2) The `links` command of SVdetect was called to identify structural variation breakpoints that may be caused by integration events by scanning discor dant read pairs in the whole genome range;

[0166] (3) The parameters were set as follows (according to the parameters used in the actual analysis, this is an example): the sliding window size --window-size) was 500bp, the support read pair threshold --min-num-reads) was 3, and the read with alignment quality value below --min-mapQ) 20 was ignored;

[0167] (4) Running SVdetect generated a list of candidate virus-host genome integration site intervals.

[0168] The identification method refers to Wang J, Mullighan CG, Easton J, Roberts S, Heatley SL, Ma J, Rusch MC, Chen K, Harris CC, Ding L, Holmfeldt L, Payne-Turner D, Fan X, Wei L, Zhao D, Obenauer JC, Naeve C, Mardis ER, Wilson RK, Downing JR, Zhang J. 2011. CREST maps somatic structural variation in cancer genomes with base-pair resolution. Nat Methods 8:652-4, the entire contents of which are incorporated herein), and CREST is used to validate these integration sites, the results of CREST are defined as high confidence, and the results not detected by SVdetect are classified as low confidence.

[0169] Finally, the consensus sequence near the integration site is extracted and collated (finally, the consensus sequence near the integration site is extracted and collated, and the specific steps are as follows:

[0170] (1) Determine the precise breakpoint and flanking sequence range: based on the high-confidence integration breakpoint coordinates verified by CREST (for example, host chromosome Chr6: 12345678 and HPV16 genome nt 1250), use the slop command of BEDTools (version 2.30.0) to extend 50 bases upstream and downstream of each breakpoint, and obtain a total of 101 bp of candidate sequence intervals.

[0171] (2) Extract raw sequencing reads: use the view and fasta commands of samtools (version 1.15) to extract all sequencing reads successfully aligned with the above candidate intervals (such as Chr6: 12345628-12345728) from the final generated BAM alignment file. The command used is as follows: samtools view -b -L candidate_regions.bed sample.bam | samtools fasta -> flanking_reads.fasta.

[0172] (3) Consensus sequence generation and trimming: The extracted read file (flanking reads.fasta) was inputted into the multiple sequence alignment software MAFFT (version 7.505) for alignment with default parameters to obtain the initial consensus sequence. Subsequently, the alignment results were trimmed using TrimAl (version 1.4.1) to remove flanking low-quality or insufficiently covered bases (--gt 0.8), and finally output the high-quality consensus sequence flanking the integration site.

[0173] (4) Result output and annotation: The final trimmed consensus sequence was outputted in FASTA format, and its corresponding sample ID, integration site genomic coordinates (host and virus), sequence length, and sequencing depth, etc. information were annotated in the file.

[0174] Methods refer to Xu S, Shi C, Zhou R, Han Y, Li N, Qu C, Xia R, Zhang C, Hu Y, Tian Z, Liu S, Wang L, Li J, Zhang Z. 2024. Mapping the landscape of HPV integration and characterising virus and host genome interactions in HPV-positive oropharyngeal squamous cell carcinoma. Clin Transl Med 14: e1556. The entire contents of which are incorporated herein by reference.

[0175] Among the 175 sequencing samples passing quality control, 45 unique HPV genotypes were detected. These genotypes included 14 high-risk types (HPV16, 18, 31, 33, 35, 39, 45, 51, 52, 56, 58, 59, 66, and 68) and 31 low-risk or uncertain risk types (HPV6, 11, 26, 30, 32, 34, 40, 42, 43, 44, 53, 54, 55, 61, 62, 67, 70, 71, 73, 74, 81, 82, 83, 84, 86, 87, 89, 90, 91, and 101), as shown in B of Figure 5

[0176] ​Among the high-risk types, HPV16, HPV52 and HPV58 were the most frequently detected, which is consistent with their known prevalence in the target population. It is noteworthy that despite the capture probe set design was intended for 32 full-length HPV reference genomes, the prediction method provided by the present invention identified 45 unique genotypes, suggesting significant non-target hybridization efficiency or possible cross-reactivity due to conserved regions among HPV types. This highlights the broad capture capability of the hybrid capture sequencing method and implies the potential for epidemiological surveillance beyond the predefined high-risk targets.

[0177] qPCR-derived Ct values exhibited a broad dynamic range (median Ct = 21.7, interquartile range 22.9-34.1), indicating the presence of both high and low copy numbers in the samples. Figure 5 The distribution of Ct values shown in FIG. 2C presents a bimodal pattern, reflecting the heterogeneity of viral load in the clinical samples.

[0178] High-throughput sequencing collectively generated 164 usable sample libraries. The raw sequencing data for each sample averaged 5.35 GB bases, and after adapter trimming and quality filtering, there were 5.02 Gb of clean data on average. The average aligned read data was 2.52 Gb, with an alignment rate of 52.15%. The duplication rate was generally low, with an average duplication rate of 10.75% (range: 0.4%-31.6%), indicating fewer artifacts generated during the PCR amplification process. A summary of sequencing quality indicators is listed in Table 2, including the minimum and maximum aligned data volume, alignment rate, and duplication statistics for each sample.

[0179] Table 2

[0180]

[0181]

[0182]

[0183]

[0184]

[0185]

[0186] Example 6 Development of a web-based HPV pooling strategy analysis tool:

[0187] To improve the applicability and reproducibility of the prediction method based on the present invention, a user-friendly Shiny application was developed using the R (v4.3.1) language.

[0188] The application contains two core functional modules: automatic pooling design and sample decoding. Users need to upload two Excel files: one contains HPV genotypes and Ct values of qPCR detection (for pooling design), and the other contains sequencing-derived HPV genotype results and pool allocation information (for decoding). Figure 2

[0189] The pooling module utilizes HPV genotypes as natural biological barcodes. According to the sample's HPV genotypes and estimated viral load (Ct values), samples with different genotypes are grouped into corresponding pools, ensuring an even number of samples in each pool. The decoding module is responsible for matching the HPV genotypes detected by sequencing back to the original sample ID in each pool.

[0190] The tool is implemented using R language and packages such as shiny, dplyr, readxl, etc., with error handling, interactive data upload, and result export functions. The complete source code and demonstration files are available for local deployment or customization.

[0191] Example 7 Verification of the prediction method:

[0192] To evaluate the agreement between the prediction method provided by the present application and the qPCR-based genotyping reference, a classification consistency analysis was performed, focusing on 14 human papillomavirus (hrHPV) genotypes. This selection was based on two reasons: first, qPCR detection is designed only to detect 14 clinically important hrHPV genotypes (HPV16, 18, 31, 33, 35, 39, 45, 51, 52, 56, 58, 59, 66, and 68); second, only integration events involving high-risk genotypes are recognized as having clinical significance for the occurrence and development of cervical cancer.

[0193] For each sample, the HPV genotypes detected by the prediction method provided by the present application were compared with the genotypes obtained by qPCR (the gold standard). A matching code variable was defined to classify the comparison results into four categories: complete match, i.e., all NGS-detected genotypes are completely consistent with qPCR-detected genotypes, both in content and quantity; partial omission, i.e., NGS missed one or more qPCR-detected genotypes but did not detect additional types; complete omission, i.e., NGS failed to detect any types present in qPCR (including samples that could not be detected); and over-detection, i.e., NGS detected additional genotypes not present in qPCR.

[0194] The results showed that 135 cases (77.1%) were completely consistent, 34 cases (19.4%) were partially omitted, 6 cases (3.4%) were completely omitted, and no over-detection was found ( Figure 3 ​Cohen's Kappa coefficient was 0.696 (p < 0.001), indicating a moderate to high agreement between the two methods, demonstrating that NGS technology based on mixed samples can reliably replicate the results of qPCR genotyping in most cases.

[0195] To further investigate the cases of inconsistency, 14 mixed samples were selected for single-sample NGS retesting. The results showed that 9 samples (64.3%) recovered a human papillomavirus genotype that was not detected in the previous mixed sample test. Notably, completely missed types such as HPV16, 33, 39, and 52 were successfully detected in single testing. At the same time, 9 cases (64.3%) also revealed new integration events (e.g., HPV39, 52, 66) that were not detected in the mixed sample setting, as shown in Table 3.

[0196] Table 3

[0197]

[0198]

[0199] These results indicate that technical limitations, particularly in low-quality batches (such as batch 3), can lead to under-detection, while targeted retesting can effectively resolve critical cases.

[0200] Example 8 Impact factor analysis of the prediction method:

[0201] To explore the factors that may lead to detection errors, statistical analysis was performed on multiple variables, including viral load (roughly assessed by Ct value), batch-specific effects, mixed pool size, HPV genotype, and multiple infection status. Ct values were converted to log10 viral copy numbers for analysis.

[0202] As shown in B of Figure 6 , the Ct values of partially or completely undetected samples were significantly higher than those of completely matched samples (Bonferroni-corrected p < 0.01), confirming that viral load is a major factor affecting detection sensitivity. Notably, five of the six completely undetected samples came from the same experimental batch, indicating a batch-specific technical error that may have caused detection failure during library preparation or hybridization (as shown in Table 3).

[0203] The pool size (i.e., the number of samples combined in each pool) was also evaluated for its impact on consistency. Pool sizes varied from 2 to 4 samples and were similar across groups. Kruskal-Wallis test showed no significant difference between groups (χ 2 = 1.21, df = 2, p = 0.546) (Figure 6 pool size did not negatively affect the accuracy of genotyping. This demonstrated the robustness and scalability of HPV Pool-Seq under different pooling strategies. Further sequencing quality control metrics showed that the number of aligned reads for batch 3 was significantly reduced ( Figure 7 pool size did not negatively affect the accuracy of genotyping. This demonstrated the robustness and scalability of HPV Pool-Seq under different pooling strategies. Further sequencing quality control metrics showed that the number of aligned reads for batch 3 was significantly reduced (

[0204] Example 9 Cost-effectiveness of the prediction method:

[0205] To evaluate the economic impact of implementing the HPV Pool-Seq strategy, the cost of sequencing 1000 HPV-positive samples under different pooling schemes was simulated. Although qPCR genotyping and computational decoding increased the cost, the pooling-based approach significantly reduced the number of sequencing libraries and sequencing lanes required.

[0206] As shown in Figure 4 , the cost saving was the largest when pooling 3 to 4 samples per library, and the cost could be reduced by up to 60% compared to standard single-sample NGS. These results suggest that HPV Pool-Seq has the potential to improve cost-effectiveness in large-scale HPV genotyping and integration detection applications.

[0207] It should be noted that the relational terms herein such as first and second and the like are used solely to distinguish one from another entity or action, without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0208] The above examples are only used to illustrate the technical solutions of the present application, rather than limit it; although the present application has been described in detail with reference to the foregoing examples, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for detecting HPV DNA integration events for non-diagnostic purposes, characterized in that, The detection method includes: Data reuse: Obtain the HPV genotype and viral load of the sample to be tested; Intelligent grouping: Constructing a DNA library containing test samples with non-overlapping HPV genotypes; Hybrid sequencing: The DNA library is hybridized with HPV whole genome probes for capture and high-throughput sequencing; Integration analysis: Using HPV genotypes as natural identifiers to trace HPV DNA integration events.

2. The detection method according to claim 1, characterized in that, The following criteria are used to determine whether a sample to be tested has undergone HPV DNA integration events: If the number of split-reads ≥ 3 and / or discordant read-pairs ≥ 5, and the virus-host border sequence is clear and uniquely aligned, then the sample to be tested is a sample in which an HPV DNA integration event has occurred.

3. The detection method according to claim 1, characterized in that, The DNA library contains 2 to 10 samples to be tested.

4. The detection method according to claim 1, characterized in that, The maximum viral load of the sample to be tested in the DNA library is 1 to 20 times the minimum value. The viral load is calculated using the formula: Copy = (1 / 2)Ct.

5. The detection method according to claim 4, characterized in that, The value of Ct ranges from 12 to 35.

6. The detection method according to claim 1, characterized in that, The detection method also includes error correction verification: when high-throughput sequencing fails to detect the HPV genotype of a sample, the sample is tested separately.

7. The detection method according to claim 1, characterized in that, The HPV genotypes include one or a combination of HPV16, 18, 31, 33, 35, 39, 45, 51, 52, 56, 58, 59, 66 or 68.

8. A detection system for HPV DNA integration events, characterized in that, It includes: Data reuse module: Obtains the HPV genotype and viral load of the sample to be tested; Intelligent grouping module: Constructs a DNA library containing test samples with non-overlapping HPV genotypes; Hybrid sequencing module: Hybridizes the DNA library with HPV whole genome probes for high-throughput sequencing; Integration analysis module: Utilizes HPV genotype as a natural identifier to trace HPV DNA integration events; Error correction and verification module: When high-throughput sequencing fails to detect the HPV genotype of a sample, test that sample separately.

9. An electronic device, characterized in that, It includes memory and processor; The memory is used to store program instructions; The processor is used to invoke program instructions, which, when executed, are used to perform the detection method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is provided with a computer program, which, when executed by a processor, implements the detection method according to any one of claims 1-7.

Citation Information

Patent Citations

  • A search method for species-specific endogenous barcodes and its application in multi-sample mixed sequencing

    CN104573407B