Method for identifying variant sites that differentiate HPV risk levels, and use thereof
By constructing a guide tree and a kinship evolution tree and screening HPV characteristic mutation sites, the problems of low efficiency and high cost of HPV subtype detection in the prior art are solved, and accurate identification and rapid distinction of HPV risk levels are achieved.
Patent Information
- Application Number
- PCT/CN2025/076515
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-29
- Filing Date
- 2025-02-08
- Publication Date
- 2025-09-04
AI Technical Summary
The existing HPV subtype detection methods rely on the reference genome of known high-risk HPV subtypes, ignoring genetic variations between viral strains, resulting in low detection efficiency and high cost, and the inability to quickly and accurately distinguish HPV risk levels.
By constructing a guide tree and affinity evolution tree, marking HPV risk levels, screening out characteristic mutation sites, using the principle of maximum simplicity to determine the risk dividing point, calculating the distribution frequency of the mutation sites, and screening the mutation sites that meet the conditions for identification.
Accurate identification of HPV risk levels has been achieved, which can quickly distinguish high-risk, possible high-risk and low-risk HPV subtypes, reduce detection costs and improve detection efficiency.
Smart Images

Figure CN2025076515_04092025_PF_FP_ABST
Abstract
Description
A method for identifying variant sites for differentiating HPV risk levels and its application Technical Field
[0001] The present invention relates to the field of bioinformatics, and relates to a method for identifying variant sites for distinguishing HPV risk levels and its application. Background Art
[0002] Papillomaviruses are a ubiquitous group of nonenveloped, double-stranded, circular DNA viruses that infect basal keratinocytes of the skin and mucosal epithelium of a variety of animal host species, including humans. This adaptive evolution of papillomaviruses, which allows them to adapt and utilize diverse organ and tissue types in different host species, leaves traces of genetic variation in their genome sequences. Traditionally, papillomaviruses have been classified into different genera and species based primarily on differences in the nucleotide sequence of the papillomavirus capsid-encoding gene, L1.
[0003] Members of each genus share >60% L1 nucleotide sequence similarity, while different species within the same genus share 60%-70% sequence similarity. Within the genus and species, it can be further subdivided into different subtypes. If the newly isolated papillomavirus has a similarity of <90% to any other previously defined papillomavirus type, it can be defined as a new papillomavirus type. To date, approximately 450 subtypes of human papillomaviruses (HPVs) that infect humans alone have been defined, and these HPVs are distributed in five genera: Alpha, Beta, Gamma, Mu, and Nu. Interestingly, HPV subtypes that infect mucosal epithelium are only found in AlphaPV, while HPV subtypes that infect skin epithelium are commonly distributed in all five genera.
[0004] Epidemiological studies have shown that many HPV subtypes are clinically asymptomatic or only cause benign tumors such as warts, and are therefore considered low-risk (LR). These include HPV6, HPV11, HPV40, HPV42, HPV43, HPV44, HPV54, HPV61, HPV72, and HPV81. However, some HPV subtypes are closely associated with the development and progression of malignancies such as cervical cancer and oropharyngeal cancer and are considered high-risk (HR) or probable high-risk (pHR) HPV subtypes. High-risk HPVs include HPV16, HPV18, HPV31, HPV33, HPV35, HPV39, HPV45, HPV51, HPV52, HPV56, HPV58, HPV59, HPV68, and HPV82. Potentially high-risk HPV types include HPV26, HPV53, and HPV66. Furthermore, in clinical diagnosis and treatment, infection with these high-risk HPV subtypes has become an important reference for tumor diagnosis and treatment. Therefore, identifying and differentiating HPV risk biomarkers and their detection methods have important clinical applications and potential translational value.
[0005] Currently, there are several main methods for detecting and identifying HPV subtypes. One involves determining the full sequence of the isolated HPV genome or the HPV-L1 gene. These sequences are then compared with standard reference sequences of different HPV subtypes in HPV reference genome databases (such as PaVE and HPV center) and assigned to the subtype using traditional classification methods. This method involves multiple experimental and analytical steps and is not suitable for rapid clinical testing. The second method is based on fluorescence quantitative PCR, with representative products including Roche AG's Cobas 4800 / 5800 / 6800 / 8800 HPV typing assays. The third method is based on RNA probe hybridization PCR, with representative products including Qiagen NV's HC-2 high-risk HPV assay. Currently, many domestic molecular testing companies have also developed similar high-risk HPV detection kits based on fluorescence quantitative PCR or RNA probe hybridization techniques. For example, Guangzhou Hybrid, a Chinese company, has developed a series of detection kits based on methods such as fluorescence PCR, PCR-membrane hybridization, and PCR-flow-through hybridization.
[0006] It is worth noting that the above-mentioned existing methods are all designed based on the reference genomes of currently known representative high-risk HPV subtypes (such as HPV16, HPV18, etc.), and have the following shortcomings: 1) The relevant detection probes or reagents need to simultaneously integrate multiple nucleic acid fragment information of multiple known high-risk HPV subtypes, and the preparation process is complicated and costly; 2) The genetic variations between different virus strains within the same subtype are ignored, so there are inherent deficiencies in dealing with existing and emerging virus mutations.
[0007] In summary, the existing nucleic acid fragments and probe designs used for HPV molecular typing-related detection are completely dependent on known high-risk HPV subtypes identified by epidemiological surveys, and lack the understanding and targeted detection of specific genetic variation sites that are truly associated with the high pathogenicity of HPV. Summary of the Invention
[0008] In order to address at least one of the deficiencies in the prior art for detecting and identifying HPV subtypes, the present invention provides a method for identifying variant sites that distinguish HPV risk levels, which can identify characteristic genetic variant sites associated with HPV risk, thereby enabling subsequent typing detection to achieve targeted and precise prevention and control.
[0009] In one aspect, the present invention provides a method for identifying variant sites that distinguish HPV risk levels, comprising the following steps:
[0010] Constructing a guide tree: Align each genome sequence in the standard HPV genome to be analyzed with the standard reference genome to obtain the first alignment result; constructing a guide tree based on the filtered first alignment result;
[0011] Constructing a phylogenetic tree: Based on the guide tree, all genomic sequences in the standard HPV genome to be analyzed and the standard reference genome are subjected to a whole-genome multiple alignment to obtain a second alignment result; and constructing a phylogenetic tree based on the filtered second alignment result;
[0012] Marking the phylogenetic tree: Based on known HPV risk classification information, mark the HPV risk level corresponding to each branch in the phylogenetic tree;
[0013] Determine the risk cutoff point: Based on the principle of maximum parsimony and the topological structure of the phylogenetic tree, find the differentiation nodes of the phylogenetic tree of HPVs at different risk levels;
[0014] Variant site risk distribution frequency analysis: Extract the variant site information from the filtered second alignment results and calculate the distribution frequency of the wild-type and mutant genotypes corresponding to each variant site in HPVs of different risk levels in the phylogenetic tree;
[0015] Screening for variant sites that meet the pre-determined distribution frequency criteria and extracting the corresponding mutant genotype states provides characteristic variants that can be used to distinguish different risk levels of HPV. In some embodiments, these are characteristic variants that can be used to distinguish high-risk, potentially high-risk, and / or low-risk HPVs.
[0016] In some embodiments, the first and second alignment results are both results of multiple alignment.
[0017] The phylogenetic tree is also called phylogenetic tree, phylogenetic tree, evolutionary tree, evolutionary tree, etc.
[0018] The principle of maximum parsimony is based on the hypothesis that evolution requires the fewest number of genotype, nucleotide, amino acid, or trait substitutions. Based on this principle, researchers calculate the number of mutations or trait substitutions required for all possible evolutionary pathways and ultimately select the pathway with the fewest required mutations or trait substitutions as the most parsimonious explanation of the possible evolutionary pathways, deeming this pathway the most likely. This principle is an application of Ockham's razor to evolutionary biology. While maximum parsimony is often used to construct phylogenetic trees, its underlying logic applies to any evolutionary biology analysis.
[0019] In some embodiments, before obtaining the comparison result of the standard HPV genome to be analyzed and the standard reference genome, the method further includes:
[0020] Obtaining an original HPV genome to be analyzed, filtering each genome sequence in the original HPV genome to be analyzed according to preset conditions to obtain a filtered original HPV genome to be analyzed;
[0021] The original reference genome is obtained, and the starting point of each genome in the original reference genome and the filtered original HPV genome to be analyzed is reset to obtain a standard reference genome and a standard HPV genome to be analyzed.
[0022] In some embodiments, the preset condition is: unsequenced bases / total genome length < threshold; in some embodiments, the threshold is 0.1-0.2; in some embodiments, the threshold is 0.125; in some embodiments, the starting point after the reset starting point is specifically: the starting point is reset to the first base after the L1 gene stop codon according to the annotation results.
[0023] In some embodiments, the first alignment result or the second alignment result after filtering includes base-level filtering; in some embodiments, the preset conditions of the base-level filtering include filtering out bases with a missing ratio of 90% or more in all genomes involved in the alignment.
[0024] In some embodiments, the guide tree or the phylogenetic tree is constructed based on a maximum likelihood method.
[0025] In some embodiments, based on the guide tree, performing a whole-genome multiple alignment of each genome sequence in the standard HPV genome to be analyzed and the standard reference genome to obtain a second alignment result specifically includes:
[0026] The guide tree is used as input, and based on the topology of the guide tree, a whole-genome multiple alignment is performed progressively, starting with the most closely related sequences on the guide tree and gradually expanding outward. In some embodiments, the input further includes a standard reference genome.
[0027] In some embodiments, the known HPV risk classification information includes: et al. (The New England Journal of Medicine, 2003, 348: 518-527) proposed the HPV risk level classification standard.
[0028] In some embodiments, low risk: HPV6, HPV11, HPV40, HPV42, HPV43, HPV44, HPV54, HPV61, HPV72, HPV81;
[0029] Possibly high risk: HPV26, HPV53, HPV66; and / or
[0030] High risk: HPV16, HPV18, HPV31, HPV33, HPV35, HPV39, HPV45, HPV51, HPV52, HPV56, HPV58, HPV59, HPV68, HPV82.
[0031] In some embodiments, the preset conditions of the distribution frequency specifically include:
[0032] If the reference genome is a high-risk and / or possibly high-risk HPV, then in the phylogenetic tree, the distribution frequency of the mutant genotype in low-risk HPV is greater than 60%, and the distribution frequency in high-risk and possibly high-risk HPV is less than 40%;
[0033] If the reference genome is a low-risk HPV, then in the phylogenetic tree, the distribution frequency of the mutant genotype in the low-risk HPV is less than 40%, and the distribution frequency in the high-risk and possibly high-risk HPV is greater than 60%.
[0034] In some embodiments, after the variant site risk distribution frequency analysis, at least one of the following is further included:
[0035] 1) Determine whether the variant site is in the repetitive sequence region,
[0036] If the variant site is in a repetitive sequence region, the site is discarded; if the variant site is in a non-repetitive sequence region, the site is retained;
[0037] 2) Determine whether the mutation site is in the gene coding region,
[0038] If the variant site is in the gene coding region, the variant site is included in the high-preferred set; if the variant site is in the non-gene coding region, the variant site is included in the suboptimal set;
[0039] 3) Determine whether the mutation site causes a change in the amino acid encoded by the corresponding gene,
[0040] If the variant site causes a change in the amino acid encoded by the corresponding gene, the variant site is included in the high-preferred set; if the variant site does not cause a change in the amino acid encoded by the corresponding gene, the variant site is included in the suboptimal set.
[0041] In some embodiments, the variation type of the variant site includes at least any one of: single nucleotide variant (SNV), short insertion / deletion (INDEL) variation, copy number variation (CNV) and structural variation (SV).
[0042] In one aspect, the present invention provides an apparatus for identifying variant sites that differentiate HPV risk levels, comprising: a memory and a processor;
[0043] The memory is used to store program instructions;
[0044] The processor is used to call program instructions, and when the program instructions are executed, it is used to execute the method for identifying the characteristic mutation sites and mutation genotype status of HPV high-risk and possible high-risk subtypes.
[0045] In one aspect, the present invention provides a system for identifying variant sites that distinguish HPV risk levels, comprising:
[0046] Guide tree construction module: used to align each genome sequence in the standard HPV genome to be analyzed to the standard reference genome to obtain the first alignment result; and to construct a guide tree based on the filtered first alignment result;
[0047] A phylogenetic tree construction module is used to perform a full-genome multiple alignment of each genome sequence in the standard HPV genome to be analyzed and the standard reference genome based on the guide tree to obtain a second alignment result; and to construct a phylogenetic tree based on the filtered second alignment result;
[0048] Phylogenetic tree marking module: used to mark the HPV risk level corresponding to each branch in the phylogenetic tree according to known HPV risk classification information;
[0049] Risk cutoff point determination module: used to find the differentiation nodes of the phylogenetic tree of HPVs of different risk levels based on the principle of maximum parsimony and the topological structure of the phylogenetic tree;
[0050] Variant site risk distribution frequency analysis module: used to extract the variant site information from the filtered second alignment results, and calculate the distribution frequency of the wild-type and mutant genotypes corresponding to each variant site in HPVs of different risk levels in the phylogenetic tree;
[0051] Mutation site screening module: used to screen mutation sites that meet the preset conditions of distribution frequency and extract the corresponding mutant genotype status.
[0052] In one aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for identifying variant sites that distinguish HPV risk levels.
[0053] It should be understood that the size of the serial numbers of the steps in the implementation plans and examples of the present invention does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the implementation plans and examples of the present invention.
[0054] In some embodiments, a computer device is provided, which may be a server, and its internal structure diagram may be as shown in Figure 3. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for identifying variant sites that distinguish HPV risk levels is implemented.
[0055] Those skilled in the art will appreciate that all or part of the processes in the methods of the embodiments of the present invention can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0056] Those skilled in the art will clearly understand that for the sake of convenience and conciseness of description, only the division of the aforementioned functional units and modules is used as an example. In actual applications, the aforementioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0057] In one aspect, the present invention provides an HPV risk category genotype fingerprint map, using the HPV16 reference genome (e.g., NCBI sequence access number: NC_001526.4) as a reference system,
[0058] The high-risk characteristic sites and genotypes are as follows: SNV1: g.734, C; SNV2: g.798, T; SNV3: g.985, A; SNV4: g.986, T; SNV5: g.1039, C; SNV6: g.1052, A; SNV7: g.1060, A; SNV8: g.1066, G; SNV9: g.1483, C; SNV10: g.1616, C; SNV11: g.1664, T; SNV12: g.1711, G; SNV13: g.1730, C; SNV14: g.1888, A; SNV15: g.1904, T; SNV16: g.1916, A; SNV17: g.1933, A; SNV18:g.1954,T; SNV19:g.2439,T; SNV20:g.2475,G; SNV21:g.2617,A; SN V22:g.2709,G; SNV23:g.3764,C; SNV24:g.3857,C; SNV25:g.4078,A; SNV26: g.4298,A; SNV27:g.4352,A; SNV28:g.4627,A; SNV29:g.7143,G; SNV30:g.73 53,T; SNV31:g.7357,C; SNV32:g.7407,T; SNV33:g.7436,T; SNV34:g.7832,C;
[0059] The corresponding possible high-risk characteristic sites and genotypes are as follows: SNV1:g.734,A; SNV2:g.798,T; SNV3:g.985,A; SNV4:g.986,T; SNV5:g.1039,C; SNV6:g.1052,A; SNV7:g.1060,C; SNV8:g.1066,G; SNV9:g.1483,C; SNV10:g.1616,T; SNV11:g.1664,T; SNV12:g.1711,A; SNV13:g.1730,G; SNV14:g.1888,A; SNV15:g.1904,T; SNV16:g.1916,G; SNV17:g.1933 ,A;SNV18:g.1954,A;SNV19:g.2439,C;SNV20:g.2475,G;SNV21:g.2617,A;S NV22:g.2709,G; SNV23:g.3764,C; SNV24:g.3857,C; SNV25:g.4078,A; SNV26: g.4298,A; SNV27:g.4352,A; SNV28:g.4627,G; SNV29:g.7143,C; SNV30:g.735 3,T; SNV31:g.7357,G; SNV32:g.7407,C; SNV33:g.7436,T; SNV34:g.7832,C; and
[0060] The corresponding low-risk characteristic sites and genotypes are as follows: SNV1:g.734,A; SNV2:g.798,G; SNV3:g.985,G; SNV4:g.986,C; SNV5:g.1039,G; SNV6:g.1052,C; SNV7:g.1060,C; SNV8:g.1066,A; SNV9:g.1483,G; SNV10:g.1616,T; SNV11:g.1664,C; SNV12:g.1711,A; SNV13:g.1730,A; SNV14:g.1888,G; SNV15:g.1904,G; SNV16:g.1916,G; SNV17:g.1933 ,G; SNV18:g.1954,A; SNV19:g.2439,A; SNV20:g.2475,C; SNV21:g.2617,G;S NV22:g.2709,A; SNV23:g.3764,G; SNV24:g.3857,T; SNV25:g.4078,C; SNV26: g.4298,G; SNV27:g.4352,T; SNV28:g.4627,T; SNV29:g.7143,C; SNV30:g.73 53,C; SNV31:g.7357,G; SNV32:g.7407,G; SNV33:g.7436,A; SNV34:g.7832,A.
[0061] In one aspect, the present invention provides the use of a marker detection reagent in the preparation of a reagent for assisting in the diagnosis of HPV risk subtypes, wherein the marker comprises at least any one of the following groups of sites and genotypes:
[0062] High-risk markers: SNV1: g.734,C; SNV2: g.798,T; SNV3: g.985,A; SNV4: g.986,T; SNV5: g.1039,C; SNV6: g.1052,A; SNV7: g.1060,A; SNV8: g.1066,G; SNV9: g.1483,C; SNV10: g.1616,C; SNV11: g.1664,T; SNV12: g.1711,G; SNV13: g.1730,C; SNV14: g.1888,A; SNV15: g.1904,T; SNV16: g.1916,A; SNV17: g.1933,A; SNV18: g.1954,T; SNV19: g.2439,T; SNV20: g.2475,G; SNV21: g.2617,A; SNV22: g.2709,G; SNV23: g.3764,C; SNV24: g.3857,C; SNV25: g.4078,A; SNV26: g.4298,A; SNV27: g.4352,A; SNV28: g.4627,A; SNV29: g.7143,G; SNV30: g.7353,T; SNV31: g.7357,C; SNV32: g.7407,T; SNV33: g.7436,T; SNV34: g.7832,C;
[0063] Possible high-risk markers: SNV1: g.734,A; SNV2: g.798,T; SNV3: g.985,A; SNV4: g.986,T; SNV5: g.1039,C; SNV6: g.1052,A; SNV7: g.1060,C; SNV8: g.1066,G; SNV9: g.1483,C; SNV10: g.1616,T; SNV11: g.1664,T; SNV12: g.1711,A; SNV13: g.1730,G; SNV14: g.1888,A; SNV15: g.1904,T; SNV16: g.1916,G; SNV17: g.1933,A; SNV18: g.1954,A; SNV19: g.2439,C; SNV20: g.2475,G; SNV21: g.2617,A; SNV22: g.2709,G; SNV23: g.3764,C; SNV24: g.3857,C; SNV25: g.4078,A; SNV26: g.4298,A; SNV27: g.4352,A; SNV28: g.4627,G; SNV29: g.7143,C; SNV30: g.7353,T; SNV31: g.7357,G; SNV32: g.7407,C; SNV33: g.7436,T; SNV34: g.7832,C; and
[0064] Low-risk markers: SNV1: g.734,A; SNV2: g.798,G; SNV3: g.985,G; SNV4: g.986,C; SNV5: g.1039,G; SNV6: g.1052,C; SNV7: g.1060,C; SNV8: g.1066,A; SNV9: g.1483,G; SNV10: g.1616,T; SNV11: g.1664,C; SNV12: g.1711,A; SNV13: g.1730,A; SNV14: g.1888,G; SNV15: g.1904,G; SNV16: g.1916,G; SNV17: g.1933,G; SNV18: g.1954,A; SNV19: g.2439,A; SNV20: g.2475,C; SNV21: g.2617,G; SNV22: g.2709,A; SNV23: g.3764,G; SNV24: g.3857,T; SNV25: g.4078,C; SNV26: g.4298,G; SNV27: g.4352,T; SNV28: g.4627,T; SNV29: g.7143,C; SNV30: g.7353,C; SNV31: g.7357,G; SNV32: g.7407,G; SNV33: g.7436,A; SNV34: g.7832,A。
[0065] On the one hand, the present invention provides a kit for assisting in the diagnosis of high-risk / potentially high-risk HPV subtypes, the kit comprising a detection reagent for detecting the presence of a marker in a sample from a subject, the marker comprising at least one of the following sites and genotypes: SNV1: g.734, C (high risk) and / or A (low risk); SNV2: g.798, T (high risk) and / or G (low risk); SNV3: g.985, A (high risk) and / or G (low risk); SNV4: g.986, T (high risk) and / or C (low risk); SNV5: g.1039, C (high risk) and / or G (low risk); SNV6: g.1052, A (high risk) and / or G (low risk); SNV7: g.1060, A (high risk) and / or C (low risk); SNV8: g.1066, G (high risk) and / or A (low risk); SNV9: g.1483, C (high risk) and / or G (low risk); SNV10: g.1616, C (high risk) and / or T (low risk); SNV11: g.1664, T (high risk) and / or C (low risk); SNV12: g.1711, G (high risk) and / or A (low risk); SNV13: g.1730, C (high risk) and / or A (low risk); SNV14: g.1888, A (high risk) and / or G (low risk); SNV15: g.1904, T (high risk) and / or G (low risk); SNV16: g.1916, A (high risk) and / or G (low risk); SNV17: g.1933, A (high risk) and / or G (low risk); SNV18: g.1954, T (high risk) and / or A (low risk); SNV19: g.2439, T (high risk) and / or A (low risk); SNV20: g.2475, G (high risk) and / or C (low risk); SNV21: g.2617, A (high risk) and / or G (low risk); SNV22: g.2709, G (high risk) and / or A (low risk); SNV23: g.3764, C (high risk) and / or or G (low risk); SNV24: g.3857, C (high risk) and / or T (low risk); SNV25: g.4078, A (high risk) and / or C (low risk); SNV26: g.4298, A (high risk) and / or G (low risk); SNV27: g.4352, A (high risk) and / or T (low risk); SNV28: g.4627, A (high risk) and / or T (low risk); SNV29: g.7143, G (high risk) and / or C (low risk); SNV30: g.7353, T (high risk) and / or C (low risk); SNV31: g.7357, C (high risk) and / or G (low risk); SNV32: g.7407, T (high risk) and / or G (low risk); SNV33: g.7436, T (high risk) and / or A (low risk); SNV34: g.7832, C (high risk) and / or A (low risk).
[0066] In one aspect, the present invention provides any one of the following applications:
[0067] Application of the device in diagnosing the occurrence and development of HPV infection-related lesions; application of the device in assisting in the selection of treatment options for HPV infection-related lesions; application of the device in classifying subjects or predicting the attributes of subjects;
[0068] The application of the system in diagnosing the occurrence and development of HPV infection-related lesions; the application of the system in assisting in the selection of treatment options for HPV infection-related lesions; and the application of the system in classifying subjects or predicting the attributes of subjects.
[0069] The invention provides an application of the HPV high-risk / potential high-risk mutation sites and mutant genotype status identified in the present invention in the development of HPV vaccines.
[0070] In some embodiments, the HPV infection-related lesions include: malignant tumors such as cervical cancer, oropharyngeal cancer, laryngeal cancer, anal cancer, vaginal cancer, vulvar cancer, penile cancer, and genital warts, condyloma acuminata, common warts, plantar warts, flat warts, toe warts, filiform warts, verrucous epidermodysplasia verruciformis, and Bowenoid papulosis.
[0071] In some embodiments, the HPV infection-related lesions include esophageal cancer, gastric cancer, and rectal cancer.
[0072] Existing nucleic acid fragment and probe designs for HPV molecular typing tests rely entirely on known high-risk HPV subtypes identified through epidemiological surveys. When the HPV strain being tested differs significantly from the standard sequences in its standard nucleic acid fragment / probe library due to factors such as genetic variation, the effectiveness of the test will be significantly affected. Furthermore, to cover common HPV subtypes, current detection methods require the design of standard nucleic acid fragment / probe libraries to accommodate a large number of targeted standard nucleic acid molecules, resulting in high design and preparation costs. In contrast, in some embodiments, the present invention identifies characteristic genetic variant sites associated with HPV risk, enabling the detection of all currently known low-risk and high-risk HPV subtypes using a 34-SNV HPV risk typing fingerprint consisting of 34 SNVs. Furthermore, at least 11 of these 34 SNV sites have a single-site detection accuracy exceeding 99%. Therefore, large-scale rapid screening tests based on a single site can be achieved without sacrificing accuracy, enabling rapid and precise prevention and control of high-risk / potentially high-risk HPV subtypes. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] FIG1 is a schematic flow chart of a method for identifying variant sites for distinguishing HPV risk levels provided by an embodiment of the present invention;
[0074] FIG2 is a schematic diagram of a device for identifying variant sites for distinguishing HPV risk levels, provided by an embodiment of the present invention;
[0075] FIG3 is a schematic diagram of a computer device for identifying variant sites that distinguish HPV risk levels, provided by an embodiment of the present invention;
[0076] Figure 4 shows the distribution of 34 HPV risk-associated signature SNV genotypes identified in the present invention based on an analysis of the evolutionary relationships of 2,969 HPV strains isolated worldwide. The figure shows a clear evolutionary divergence between high-risk and low-risk HPV subtypes, with the divergence points marked with red stars.
[0077] FIG5 is a flow chart illustrating the analysis and identification of HPV risk-related characteristic SNV fingerprints based on pan-genomics according to an embodiment of the present invention. DETAILED DESCRIPTION
[0078] The following is a detailed description of the technical solution of the present invention, which does not limit the scope of protection of the present invention. Non-essential modifications and adjustments made by others based on the concept of the present invention still fall within the scope of protection of the present invention.
[0079] Example 1 Download and organization of global papillomavirus genome data and identification of HPV risk-associated SNV fingerprints based on pan-genome analysis
[0080] 1) Query and download the global genomic resources of papillomaviruses before June 1, 2023 in the NCBI Genbank public database (https: / / www.ncbi.nlm.nih.gov / genbank / ) using the keywords "papillomavirus" AND "complete genome". The accession numbers of all genomic sequences that meet the keywords can be output using the NCBI web query interface, and then downloaded in batches using the NCBI's EFetch utility (v16.2) command-line tool. It should be noted that since a large amount of genomic public resources of papillomaviruses have been accumulated, here we conduct our analysis at the big data level by downloading public data to search for HPV risk-related fingerprints. Essentially, this method is also applicable to the papillomavirus genomic data obtained by researchers through sequencing experiments. However, considering that it is difficult for the papillomavirus genomic data self-tested by researchers to reach the level of the global papillomavirus genomic data in the NCBI Genbank database in terms of data volume and genetic diversity, the recommended preferred solution is to combine and analyze the data self-tested by researchers with the public data from NCBI Genbank according to the procedures described in this plan.
[0081] For the downloaded papillomavirus genomic sequences, calculate the number of unsequenced bases (base sequence 'N') and denote it as L. N And the total genome length G, only keep L N / G < alpha genomic data. It is recommended that the value range of alpha is 0.1 - 0.2. In this example, we use alpha = 0.125. After filtering, we obtained a total of 3188 papillomavirus genomes for downstream analysis, including 2969 human papillomaviruses (HPV) and 219 non-human papillomaviruses. Among these 2969 HPVs, 443 are low-risk HPVs, 76 are potentially high-risk HPVs, 1779 are high-risk HPVs, and another 671 HPVs lack clear risk annotations.
[0082] 2) Use PuMA (v1.2.2) for protein-coding gene annotation, especially for the four core genes E1, E2, L1, and L2. According to the PuMA gene annotation results, reset the starting point of each papillomavirus genomic sequence to the first base after the stop codon of the L1 gene. For the papillomavirus genomes after resetting the starting point, use windowmasker (v1.0.0) to mark (softmasking) the repetitive sequence regions and use PuMA to re-annotate the protein-coding genes.
[0083] 3) All reset-start papillomavirus genomes were aligned to the reset-start HPV16 reference genome (NCBI sequence access number NC_001526.4 for the original HPV reference genome without the reset starting point) using MAFFT (v7.310). The MAFFT genome alignment results were filtered at the base level using ClipKIT (v1.3.0) (filtering out bases with a missing ratio of 90% or more in all genomes involved in the alignment), and a preliminary evolutionary tree was constructed based on the filtered alignment using IQ-TREE (v2.0.6) (here we refer to it as a guide tree). Based on the guide tree constructed by IQ-TREE, Cactus (v2.2.0) was further used to perform a full-genome multiple alignment of all reset-start papillomavirus genome sequences (including the HPV16 reference genome). The alignment results were converted from HAL format to MAF format using the HAL package (v2.2.0) and further converted to fasta format using a python script. ClipKIT (v1.3.0) was used again to perform further base-level filtering on the results of the Cactus fine alignment (filtering out bases with a missing ratio of 90% or more in all genomes involved in the alignment), and IQ-TREE (v2.0.6) was used to implement the final phylogenetic tree construction based on the filtered alignment results. et al. (The New England Journal of Medicine, 2003, 348:518-527), with the corresponding HPV risk level for each branch of the evolutionary tree indicated below. Based on this information and the principle of maximum parsimony, we identified the cutoff point between low-risk and high-risk / possibly high-risk HPVs on the phylogenetic tree and divided HPVs into two groups: low-risk and high-risk / possibly high-risk HPVs (Figure 1). Specifically, based on the topological structure of the constructed phylogenetic tree, we calculated the distribution ratios of low-risk and high-risk / possibly high-risk HPVs on either side of each internal differentiation node. We ultimately identified a differentiation node where the difference in the distribution ratios of high-risk / possibly high-risk HPVs on the left and right sides of the differentiation node was maximized. This differentiation node was then defined as the cutoff point between low-risk and high-risk / possibly high-risk HPVs.
[0084] 4) Using the reset HPV16 reference genome as the reference coordinate system, SNP-sites (v2.5.1) was used to extract SNV variant site information based on the Cactus alignment results filtered by ClipKIT (v1.3.0). For each extracted SNV variant site, the distribution frequency of its different variant forms in low-risk, possibly high-risk, and high-risk HPVs was calculated. Genetic variant sites in which the distribution frequency of the mutant genotype (in this example, the genotype that differs from the HPV16 reference genome sequence at this site) in low-risk HPV is greater than 60%, and the distribution frequency in high-risk HPV is less than 40% were selected as HPV risk-associated sites (as is well known to those skilled in the art, in this example, since the reference genome HPV16 is a known high-risk HPV, the discriminant condition is used based on the distribution frequency; in the case of a known low-risk HPV, the discriminant condition needs to be reversed, that is, the distribution frequency of the mutant genotype in low-risk HPV is less than 40%, and the distribution frequency in high-risk and possibly high-risk HPV is greater than 60%). On this basis, further screening can be performed based on whether the SNV site is in the gene coding region and whether it causes changes in the amino acids encoded by the gene.
[0085] 5) Based on the above screening, we ultimately identified a characteristic SNV fingerprint that can distinguish low-risk from high-risk HPV subtypes (Figure 2). This fingerprint consists of 34 SNVs, all of which are non-synonymous SNVs, of which 17 occur in the E1 gene, 4 in the E2 gene, 6 in the L2 gene, 5 in the E6 gene, 1 in the E1^E4 gene, and 1 in the E7 gene. The E1 and E2 genes are core genes that regulate papillomavirus replication, and these 22 SNVs occur in key protein-coding domains of the E1 and E2 genes.
[0086] Taking the HPV16 reference genome (NCBI sequence access number: NC_001526.4) as the reference system, the occurrence positions of these 34 SNVs and the main nucleotide states of their corresponding high-risk HPVs are described as follows: SNV1: g.734, C; SNV2: g.798, T; SNV3: g.985, A; SNV4: g.986, T; SNV5: g.1039, C; S NV6:g.1052,A; SNV7:g.1060,A; SNV8:g.1066,G; SNV9:g.1483,C; SNV10:g.1616,C;S NV11:g.1664,T; SNV12:g.1711,G; SNV13:g.1730,C; SNV14:g.1888,A; SNV15:g.1904, T; SNV16:g.1916,A; SNV17:g.1933,A; SNV18:g.1954,T; SNV19:g.2439,T; SNV20:g.2 475,G; SNV21:g.2617,A; SNV22:g.2709,G; SNV23:g.3764,C; SNV24:g.3857,C; SNV25: g.4078,A; SNV26:g.4298,A; SNV27:g.4352,A; SNV28:g.4627,A; SNV29:g.7143,G; SNV 30: g.7353, T; SNV31: g.7357, C; SNV32: g.7407, T; SNV33: g.7436, T; SNV34: g.7832, C.
[0087] The corresponding major nucleotide genotypes of possible high-risk HPV are described as follows: SNV1: g.734, A; SNV2: g.798, T; SNV3: g.985, A; SNV4: g.986, T; SNV5: g.1039, C; SNV6: g.1052, A; SNV7: g.1060, C; SNV8: g.1066, G; SNV9: g.1483, C; SNV10: g.1616, T; SNV11: g.1664, T; SNV12: g.1711, A; SNV13: g.1730, G; SNV14: g.1888, A; SNV15: g.1904, T; SNV16: g.1916, G; SNV17: g. 1933,A; SNV18:g.1954,A; SNV19:g.2439,C; SNV20:g.2475,G; SNV21:g.2617, A; SNV22:g.2709,G; SNV23:g.3764,C; SNV24:g.3857,C; SNV25:g.4078,A; SNV2 6:g.4298,A; SNV27:g.4352,A; SNV28:g.4627,G; SNV29:g.7143,C; SNV30:g.7 353, T; SNV31: g.7357, G; SNV32: g.7407, C; SNV33: g.7436, T; SNV34: g.7832, C.
[0088] The corresponding major nucleotide genotypes of low-risk HPV are described as follows: SNV1:g.734,A; SNV2:g.798,G; SNV3:g.985,G; SNV4:g.986,C; SNV5:g.1039,G; SNV6:g.1052,C; SNV7:g.1060,C; SNV8:g.1066,A; SNV9:g.1483,G; SNV10:g.1616,T; SNV11:g.1664,C; SNV12:g.1711,A; SNV13:g.1730,A; SNV14:g.1888,G; SNV15:g.1904,G; SNV16:g.1916,G; SNV17:g.1 933,G; SNV18:g.1954,A; SNV19:g.2439,A; SNV20:g.2475,C; SNV21:g.2617,G ;SNV22:g.2709,A;SNV23:g.3764,G;SNV24:g.3857,T;SNV25:g.4078,C;SNV2 6:g.4298,G; SNV27:g.4352,T; SNV28:g.4627,T; SNV29:g.7143,C; SNV30:g.7 353,C; SNV31:g.7357,G; SNV32:g.7407,G; SNV33:g.7436,A; SNV34:g.7832,A.
[0089] Based on the analysis results of 2298 HPV genomes analyzed in this example (including 443 low-risk HPVs, 76 possible high-risk HPVs, and 1779 high-risk HPVs), we confirmed that the independent and combined use of these 34 SNVs can achieve accurate identification of low-risk and high-risk / possibly high-risk HPV genomes (Table 1). We named the collection of these 34 SNV sites the HPV risk typing 34-SNV fingerprint map.
[0090] Table 1 HPV risk typing 4-SNV fingerprints identified in the present invention based on the genotype distribution of 2298 HPV
[0091] Note: N in the table stands for not available, which means that the nucleotide base status of the corresponding sequence at that position is unresolved.
[0092] Example 2: Validation of HPV risk-associated SNV fingerprints based on 232 additional HPV genome data in the NCBI Genbank database
[0093] We retrieved 232 HPV genomes from the NCBI Genbank public dataset that were uploaded after or before June 1, 2023, but were not included in our pan-genomic analysis due to genome assembly integrity issues. These 232 HPV genomes include 24 low-risk HPVs, 5 potentially high-risk HPVs, and 163 high-risk HPVs. Based on the sequences of these 232 HPV genomes and the known risk subtype classifications, we independently verified the superior performance of the HPV risk-associated 34-SNV fingerprint identified in this invention in distinguishing low-risk from high-risk / potentially high-risk HPV subtypes (Table 2).
[0094] Table 2 Genotype verification results of the HPV risk typing 4-SNV fingerprints identified in the present invention based on an independent validation set of 232 HPVs
[0095] Note: N in the table stands for not available, which means that the nucleotide base status of the corresponding sequence at that position is unresolved.
[0096] Example 3: Validation of HPV risk-associated SNV fingerprints based on HPV genomic data from 45 clinical cervical cancer patients
[0097] HPV genome data isolated from tumor tissues of clinical cervical cancer patients were downloaded from the Genbase database (https: / / ngdc.cncb.ac.cn / genbase), and the corresponding Genbase access numbers are C_AA048791.1 to C_AA048835.1. These 45 patients are all cervical cancer patients admitted to the Sun Yat-sen University Cancer Center in recent years, and their HPV genome sequences are derived from de novo genome assembly based on second-generation sequencing (BGI BGIMGISEQ-2000 sequencing platform). It should be noted that these sequencing data are all from our own clinical sample collection and HPV genome sequencing, which is independent of the data set used in the big data analysis of our identification of the 34-SNV fingerprint map in Example 1.
[0098] These data were aligned to the HPV16 reference genome (NCBI Sequence Accession Number: NC_001526.4) using MAFFT (v7.310). SNV variant site information was extracted using SNP-sites (v2.5.1) based on the ClipKIT-filtered Cactus alignment results. Analysis of the genotypes of these 45 HPV genomes from cervical cancer tumor tissues based on the 34 HPV risk-associated SNV fingerprints identified in this study revealed that these genotypes clearly indicated that these HPV subtypes were high-risk (Table 3).
[0099] Table 3 Genotype validation results of the HPV risk typing 34-SNV fingerprint identified in the present invention based on an independent validation set consisting of de novo HPV genome assembly data from 45 cervical cancer patients
[0100] Note: N in the table stands for not available, which means that the nucleotide base status of the corresponding sequence at that position is unresolved.
[0101] It is also worth noting that due to the limitations of the quality of de novo genome assembly, the status of some of the 45 HPV genomes at some sites is unresolved, so they are marked as N (not available) in the table. In order to more accurately determine the precise genotypes of the corresponding HPV genomes at these sites, you can also query and download the original genome sequencing data of 45 samples from the GSA database (https: / / ngdc.cncb.ac.cn / gsa / ) through the data access number CRA013231 and the sample access number CRX839115-CRX839159. By aligning these sequencing data to the HPV16 reference genome using the bwa software, we can obtain the precise genotypes of these positions. As shown in Table 4, the HPV risk typing 34-SNV fingerprint map identified by the present invention achieved 100% high-risk typing accuracy (Table 4).
[0102] Table 4 Genotype verification results of the HPV risk typing 4-SNV fingerprints identified in the present invention based on an independent validation set consisting of HPV genome raw sequencing data from 45 cervical cancer patients
[0103] Note: N in the table stands for not available, which means that the nucleotide base status of the corresponding sequence at that position is unresolved.
[0104] Example 4: Validation of HPV risk-associated SNV fingerprints based on HPV genomic data from 45 patients with oropharyngeal cancer
[0105] HPV genome data isolated from tumor tissues of 45 clinical oropharyngeal cancer patients were downloaded from the Genbase database (https: / / ngdc.cncb.ac.cn / genbase), and the corresponding Genbase access numbers are C_AA048836.1 to C_AA048880.1. These 45 patients are all oropharyngeal cancer patients admitted to the Sun Yat-sen University Cancer Center in recent years, and their HPV genome sequences are derived from de novo genome assembly based on second-generation sequencing (BGI BGIMGISEQ-2000 sequencing platform). It should be noted that these sequencing data are all from our own clinical sample collection and HPV genome sequencing, which is independent of the data set used for our big data analysis in Example 1.
[0106] These data were aligned to the HPV16 reference genome (NCBI Sequence Accession Number: NC_001526.4) using MAFFT (v7.310). SNV variant site information was extracted using SNP-sites (v2.5.1) based on the ClipKIT-filtered Cactus alignment results. Analysis of the genotypes of the HPV genomes from these 45 oropharyngeal cancer patient tumor tissues against the 34 HPV risk-associated SNV fingerprints identified in this study revealed that these genotypes clearly indicated that these HPV subtypes were high-risk (Table 5).
[0107] Table 5: Genotyping results of the HPV risk typing 34-SNV fingerprints identified in the present invention based on an independent validation set consisting of de novo HPV genome assembly data from 45 oropharyngeal cancer patients
[0108] Note: N in the table stands for not available, which means that the nucleotide base status of the corresponding sequence at that position is unresolved.
[0109] It is also worth noting that due to the limitations of the quality of the genome de novo assembly, the status of some of the individual genomes in these 45 HPV genomes is unresolved at some sites, so they are marked as N in the table. In order to more accurately determine the precise genotypes of the corresponding HPV genomes at these sites, you can also query and download the original genome sequencing data of 45 samples from the GSA database (https: / / ngdc.cncb.ac.cn / gsa / ) through the data access number CRA013231 and the sample access number CRX839160-CRX839204. By aligning these sequencing data to the HPV16 reference genome by bwa software, we can obtain the precise genotypes of these positions. As shown in Table 6, the HPV risk typing 34-SNV fingerprint map identified by the present invention achieved 100% high-risk typing accuracy (Table 6).
[0110] Table 6: Genotype validation results of the HPV risk typing 4-SNV fingerprints identified in the present invention based on an independent validation set consisting of raw HPV genome sequencing data from 45 oropharyngeal cancer patients
[0111] Note: N in the table stands for not available, which means that the nucleotide base status of the corresponding sequence at that position is unresolved.
[0112] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A method for identifying variant sites that distinguish HPV risk levels, comprising the following steps: Constructing a guide tree: Align each genome sequence in the standard HPV genome to be analyzed with the standard reference genome to obtain the first alignment result; constructing a guide tree based on the filtered first alignment result; Constructing a phylogenetic tree: Based on the guide tree, all genomic sequences in the standard HPV genome to be analyzed and the standard reference genome are subjected to a whole-genome multiple alignment to obtain a second alignment result; and a phylogenetic tree is constructed based on the filtered second alignment result; Marking the phylogenetic tree: Based on known HPV risk classification information, mark the HPV risk level corresponding to each branch in the phylogenetic tree; Determine the risk cutoff point: Based on the principle of maximum parsimony and the topological structure of the phylogenetic tree, find the differentiation nodes of the phylogenetic tree of HPVs at different risk levels; Variant site risk distribution frequency analysis: Extract the variant site information from the filtered second alignment results and calculate the distribution frequency of the wild-type and mutant genotypes corresponding to each variant site in HPVs of different risk levels in the phylogenetic tree; Screen the variant sites that meet the preset conditions of distribution frequency and extract the corresponding mutant genotype status.
2. The method according to claim 1, wherein Before obtaining the comparison results of the standard HPV genome to be analyzed and the standard reference genome, it also includes: Obtaining an original HPV genome to be analyzed, filtering each genome sequence in the original HPV genome to be analyzed according to preset conditions to obtain a filtered original HPV genome to be analyzed; Obtaining an original reference genome, resetting the starting point of each genome in the original reference genome and the filtered original HPV genome to be analyzed, to obtain a standard reference genome and a standard HPV genome to be analyzed; Preferably, the preset condition is: unsequenced bases / total genome length < threshold; Preferably, the threshold value is 0.1-0.2; further preferably, the threshold value is 0.125; Preferably, the starting point after resetting the starting point is specifically: resetting the starting point to the first base after the stop codon of the L1 gene according to the annotation result.
3. The method according to claim 1, wherein The filtered first alignment result or the second alignment result includes filtering at the base level; Preferably, the preset conditions for filtering at the base level include filtering out bases whose missing ratio in all genomes involved in the alignment reaches 90% or more; Preferably, the guide tree or the phylogenetic tree is constructed based on the maximum likelihood method; Preferably, based on the guide tree, all genome sequences in the standard HPV genome to be analyzed and the standard reference genome are subjected to a whole-genome multiple alignment to obtain a second alignment result, specifically comprising: Taking the guide tree as input, and gradually expanding outward from the most closely related sequences on the guide tree according to the topological structure of the guide tree, a whole-genome multiple alignment is performed progressively; Preferably, the input conditions also include a standard reference genome.
4. The method according to claim 1, wherein The known HPV risk classification information includes: Low risk: HPV6, HPV11, HPV40, HPV42, HPV43, HPV44, HPV54, HPV61, HPV72, HPV81; Possibly high risk: HPV26, HPV53, HPV66; and / or High risk: HPV16, HPV18, HPV31, HPV33, HPV35, HPV39, HPV45, HPV51, HPV52, HPV56, HPV58, HPV59, HPV68, HPV82; Preferably, the variation type of the variation site includes at least any one of: single nucleotide variation, short insertion / deletion variation, copy number variation and structural variation.
5. The method according to claim 1, wherein The preset conditions of the distribution frequency specifically include: If the reference genome is a high-risk and / or possibly high-risk HPV, then in the phylogenetic tree, the distribution frequency of the mutant genotype in low-risk HPV is greater than 60%, and the distribution frequency in high-risk and possibly high-risk HPV is less than 40%; If the reference genome is a low-risk HPV, then in the phylogenetic tree, the distribution frequency of the mutant genotype in the low-risk HPV is less than 40%, and the distribution frequency in the high-risk and possibly high-risk HPV is greater than 60%.
6. The method of claim 1, wherein: After the risk distribution frequency analysis of the variant sites, at least one of the following is also included: 1) Determine whether the variant site is in the repetitive sequence region, If the variant site is in a repetitive sequence region, the site is discarded; if the variant site is in a non-repetitive sequence region, the site is retained; 2) Determine whether the mutation site is in the gene coding region, If the variant site is in the gene coding region, the variant site is included in the high-preferred set; if the variant site is in the non-gene coding region, the variant site is included in the low-preferred set; and 3) Determine whether the mutation site causes a change in the amino acid encoded by the corresponding gene, If the variant site causes a change in the amino acid encoded by the corresponding gene, the variant site is included in the high-preferred set; if the variant site does not cause a change in the amino acid encoded by the corresponding gene, the variant site is included in the suboptimal set.
7. A device for identifying variant sites that differentiate HPV risk levels, comprising: Memory and processor; The memory is used to store program instructions; and The processor is used to call program instructions, and when the program instructions are executed, is used to execute the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 6 when executed by a processor.
9. An HPV risk category genotype fingerprint, wherein: Taking the HPV16 reference genome as the standard, The high-risk genotypes are: SNV1: g.734, C; SNV2: g.798, T; SNV3: g.985, A; SNV4: g.986, T; SNV5: g.1039, C; SNV6: g.1052, A; SNV7: g.1060, A; SNV8: g.1066, G; SNV9: g.1483, C; SNV10: g.1616, C; SNV11: g.1664, T; SNV12: g.1711, G; SNV13: g.1730, C; SNV14: g.1888, A; SNV15: g.1904, T; SNV16: g.1916, A; SNV17: g.1933, A; SNV18: g.1954, T; SNV19: g.2439, T; SNV20: g.2475, G; SNV21: g.2617, A; SNV22: g.2709, G; SNV23: g.3764, C; SNV24: g.3857, C; SNV25: g.4078, A; SNV26: g.4298, A; SNV27: g.4352, A; SNV28: g.4627, A; SNV29: g.7143, G; SNV30: g.7353, T; SNV31: g.7357, C; SNV32: g.7407, T; SNV33: g.7436, T; SNV34: g.7832, C; The corresponding possible high-risk genotypes are: SNV1:g.734,A; SNV2:g.798,T; SNV3:g.985,A; SNV4:g.986,T; SNV5:g.1039,C; SNV6:g.1052,A; SNV7:g.1060,C; SNV8:g.1066,G; SNV9:g.1483,C; SNV10:g.1616,T; SNV11:g.1664,T; SNV12:g.1711,A; SNV13:g.1730,G; SNV14:g.1888,A; SNV15:g.1904,T; SNV16:g.1916,G; SNV17:g.1933,A; SNV18:g.1954,A; SNV19:g.2439,C; SNV20:g.2475,G; SNV21:g.2617,A; SNV2 2:g.2709,G; SNV23:g.3764,C; SNV24:g.3857,C; SNV25:g.4078,A; SNV26:g. 4298,A; SNV27:g.4352,A; SNV28:g.4627,G; SNV29:g.7143,C; SNV30:g.7353 ,T; SNV31:g.7357,G; SNV32:g.7407,C; SNV33:g.7436,T; SNV34:g.7832,C; and The corresponding low-risk genotypes are: SNV1:g.734,A; SNV2:g.798,G; SNV3:g.985,G; SNV4:g.986,C; SNV5:g.1039,G; SNV6:g.1052,C; SNV7:g.1060,C; SNV8:g.1066,A; SNV9:g.1483,G; SNV10:g.1616,T; SNV11:g.1664,C; SNV12:g.1711,A; SNV13:g.1730,A; SNV14:g.1888,G; SNV15:g.1904,G; SNV16:g.1916,G; SNV17:g.1933,G; SNV18:g.1954,A; SNV19:g.2439,A; SNV20:g.2475,C; SNV21:g.2617,G; SNV 22:g.2709,A; SNV23:g.3764,G; SNV24:g.3857,T; SNV25:g.4078,C; SNV26:g .4298,G; SNV27:g.4352,T; SNV28:g.4627,T; SNV29:g.7143,C; SNV30:g.735 3,C; SNV31:g.7357,G; SNV32:g.7407,G; SNV33:g.7436,A; SNV34:g.7832,A.
10. Use of a marker detection reagent in the preparation of a reagent for assisting in the diagnosis of HPV risk subtypes, wherein the marker comprises at least one of the following groups of sites and genotypes: High-risk markers: SNV1: g.734,C; SNV2: g.798,T; SNV3: g.985,A; SNV4: g.986,T; SNV5: g.1039,C; SNV6: g.1052,A; SNV7: g.1060,A; SNV8: g.1066,G; SNV9: g.1483,C; SNV10: g.1616,C; SNV11: g.1664,T; SNV12: g.1711,G; SNV13: g.1730,C; SNV14: g.1888,A; SNV15: g.1904,T; SNV16: g.1916,A; SNV17: g.1933,A; SNV18: g.1954,T; SNV19: g.2439,T; SNV20: g.2475,G; SNV21: g.2617,A; SNV22: g.2709,G; SNV23: g.3764,C; SNV24: g.3857,C; SNV25: g.4078,A; SNV26: g.4298,A; SNV27: g.4352,A; SNV28: g.4627,A; SNV29: g.7143,G; SNV30: g.7353,T; SNV31: g.7357,C; SNV32: g.7407,T; SNV33: g.7436,T; SNV34: g.7832,C; Possible high-risk markers: SNV1: g.734,A; SNV2: g.798,T; SNV3: g.985,A; SNV4: g.986,T; SNV5: g.1039,C; SNV6: g.1052,A; SNV7: g.1060,C; SNV8: g.1066,G; SNV9: g.1483,C; SNV10: g.1616,T; SNV11: g.1664,T; SNV12: g.1711,A; SNV13: g.1730,G; SNV14: g.1888,A; SNV15: g.1904,T; SNV16: g.1916,G; SNV17: g.1933,A; SNV18: g.1954,A; SNV19: g.2439,C; SNV20: g.2475,G; SNV21: g.2617,A; SNV22: g.2709,G; SNV23: g.3764,C; SNV24: g.3857,C; SNV25: g.4078,A; SNV26: g.4298,A; SNV27: g.4352,A; SNV28: g.4627,G; SNV29: g.7143,C; SNV30: g.7353,T; SNV31: g.7357,G; SNV32: g.7407,C; SNV33: g.7436,T; SNV34: g. Low-risk markers: SNV1: g.734,A; SNV2: g.798,G; SNV3: g.985,G; SNV4: g.986,C; SNV5: g.1039,G; SNV6: g.1052,C; SNV7: g.1060,C; SNV8: g.1066,A; SNV9: g.1483,G; SNV10: g.1616,T; SNV11: g.1664,C; SNV12: g.1711,A; SNV13: g.1730,A; SNV14: g.1888,G; SNV15: g.1904,G; SNV16: g.1916,G; SNV17: g.1933,G; SNV18: g.1954,A; SNV19: g.2439,A; SNV20: g.2475,C; SNV21: g.2617,G; SNV22: g.2709,A; SNV23: g.3764,G; SNV24: g.3857,T; SNV25: g.4078,C; SNV26: g.4298,G; SNV27: g.4352,T; SNV28: g.4627,T; SNV29: g.7143,C; SNV30: g.7353,C; SNV31: g.7357,G; SNV32: g.7407,G; SNV33: g.7436,A; SNV34: g.7832,A。
Citation Information
Patent Citations
Method and device for cancer risk prediction
CN116403644A
Viral genotyping method
US20080154567A1
Methods and Systems for Use in Cancer Prediction
US20200342956A1
Cited By
Measurement method of sexual reproduction animal germline mutation rate and application
CN121171342A