Method and device for recognizing a new-born chromatin loop of HPV integration
By acquiring interaction sequencing reads to identify HPV integration breakpoints, constructing fusion contigs and performing alignment corrections, clustering anchor pairs, and calculating newborn scores, the problem of low accuracy in identifying HPV integration newborn chromatin circuits was solved, achieving stable and accurate identification and characterization of regulatory relationships.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAZHONG AGRI UNIV
- Filing Date
- 2026-03-05
- Publication Date
- 2026-06-26
Smart Images

Figure CN122290694A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics, and in particular to a method and apparatus for identifying HPV-integrated neochromatin loops. Background Technology
[0002] Genomic integration of high-risk HPV is one of the key events in the development and progression of HPV-related tumors such as cervical cancer. Integration can lead to effects such as local structural variations, enhancer hijacking, and epigenetic reprogramming. However, increasing evidence suggests that its effects are not limited to a linear range near the integration site, but are closely related to three-dimensional chromatin remodeling. Therefore, accurately identifying newly formed chromatin loops induced by HPV integration at the three-dimensional genome level has become a key technical issue for understanding the oncogenic mechanism of HPV integration and supporting risk stratification and intervention research.
[0003] Existing traditional interaction localization systems (ILLs) typically align interacting sequencing reads directly to the host reference genome and identify chromatin loops based on anchor clustering relationships. However, when HPV genomes integrate and form cross-genome connections, this method struggles to correctly handle chimeric and split reads covering virus-host breakpoints. It is prone to misinterpreting or misprojecting cross-breakpoint supporting evidence into similar regions of the host genome. This results in increased false positives, missed true positives, and low reproducibility of identification results during the identification of newly formed chromatin loops, leading to low accuracy.
[0004] Therefore, there is an urgent need for a method and device for identifying HPV-integrated neochromatin loops. Summary of the Invention
[0005] This application provides a method and apparatus for identifying HPV integrated neochromatin loops, which solves the problem that traditional interaction positioning systems are prone to identification deviations and have low identification accuracy during the identification process of neochromatin loops.
[0006] The first aspect of this application provides a method for identifying HPV integration neochromatin loops. The method includes: acquiring interaction sequencing read data and determining an HPV reference sequence and a host reference sequence based on the interaction sequencing read data; the interaction sequencing read data includes sequencing read data, Hi-C data, ChIA-PET data, HiChIP data, and PLAC-seq data; identifying HPV integration breakpoints based on anomaly alignment features; the anomaly alignment features include chimeric read features, split read features, and paired end anomaly features; based on the breakpoint direction of the HPV integration breakpoint, constructing a fusion contig between the HPV reference sequence and the host reference sequence to obtain a breakpoint-aware extended reference set; and comparing the interaction sequencing read data with the breakpoint... The system performs alignment operations on an expanded reference set. Based on the alignment results and breakpoint-aware alignment constraints, it performs anchor point correction and filtering operations on the target reads to obtain high-confidence cross-genomic anchor point pairs. The target reads are those falling within a preset breakpoint-aware region, which includes a fusion region and a breakpoint neighborhood window. The breakpoint-aware alignment constraints include breakpoint neighborhood consistency constraints, directional consistency constraints, and unique alignment constraints. The high-confidence cross-genomic anchor point pairs are clustered to obtain a candidate loop set, and at least one target HPV-loop is selected from the candidate loop set based on preset screening rules. The newborn score corresponding to the target HPV-loop is calculated based on a preset newborn determination algorithm, and the chromatin loop identification result is output based on the newborn score.
[0007] Optionally, determining whether the alignment result satisfies the breakpoint neighborhood consistency constraint specifically includes: determining whether the alignment positions and alignment directions of the two segments corresponding to the cross-breakpoint read segments in the alignment result are consistent with the breakpoint connection relationship, and determining whether the two alignment positions are within the preset breakpoint neighborhood window; if the alignment positions and alignment directions of the two segments corresponding to the cross-breakpoint read segments in the alignment result are consistent with the breakpoint connection relationship, and the two alignment positions are within the preset breakpoint neighborhood window, then the alignment result is determined to satisfy the breakpoint neighborhood consistency constraint.
[0008] Optionally, clustering high-confidence cross-genome anchor pairs yields a candidate loop set. Specifically, this includes clustering high-confidence cross-genome anchor pairs and generating a candidate loop set while maintaining constraints on host-end anchor distance distribution, chromatin state stratification, and coverage stratification.
[0009] Optionally, the newborn score corresponding to the target HPV-loop is calculated based on a preset newborn determination algorithm, and the chromatin loop identification result is output according to the newborn score. Specifically, this includes: selecting high-confidence target HPV-loops from the target HPV-loops based on the number of supporting evidences for the target HPV-loop and in combination with a significance assessment calculation strategy; the significance assessment calculation strategy includes significance calculation based on a distance-dependent background model and significance calculation based on a permutation test; obtaining a set of random control loops that match the distance or chromatin state corresponding to the set of high-confidence target HPV-loops; and comparing the set of high-confidence target HPV-loops with the reference host loop library and the random control loop set, respectively, to calculate the newborn score corresponding to the target HPV-loop.
[0010] Optionally, the newborn score corresponding to the target HPV-loop is calculated based on a preset newborn determination algorithm, specifically including: calculating the HPV-loop remodeling index, 3D-haplotype consistency score and target gene loop coupling expression index respectively, and using at least one of the HPV-loop remodeling index, 3D-haplotype consistency score and target gene loop coupling expression index as the newborn score.
[0011] Optionally, the HPV-loop remodeling index is calculated, specifically by: introducing a double-control method to calculate the denovo score, and summing the denovo scores to calculate the HPV-loop remodeling index; the double-control method is the reference loop library control method and the matched random control method.
[0012] Optionally, the 3D-haplotype consistency score is calculated, specifically including: classifying the reads supporting the target HPV-loop to haplotypes based on heterozygous SNP phases and calculating loop allele bias; jointly modeling the loop allele bias with allele-specific expression in RNA-seq and outputting a set of 3D-haplotype regulatory events; and calculating the 3D-haplotype consistency score based on the set of 3D-haplotype regulatory events.
[0013] Optionally, the loop-coupled expression index of the target gene is calculated, specifically including: performing motif enrichment analysis on the HPV-terminal anchor and host-terminal anchor of the denavoHPV-loop, and screening a set of transcription factors that meet preset conditions; constructing a protein-protein interaction network based on the transcription factor set to obtain a co-transcription factor module, and outputting the loop-coupled expression index of the target gene based on the transcription factor module; the target gene includes at least one of MYC, CASC19, and FAM81A; the co-transcription factor module includes at least one of FOS, CEBPG, SP1, or CTCF.
[0014] Optionally, the chromatin loop identification results are output based on the neonatal score, specifically including: in vitro risk stratification analysis based on the chromatin loop identification results; in vitro risk stratification analysis is used to quantify and stratify the carcinogenic potential of HPV-positive samples or cell models and provide reference information; prognostic correlation analysis is performed based on the chromatin loop identification results; prognostic correlation analysis is used to analyze HPV-loop coupling activation events corresponding to different target genes and use HPV-loop coupling activation events as candidate prognostic indicators; intervention screening is performed based on the chromatin loop identification results; intervention screening is used to screen candidate interventions that inhibit HPV-loop formation or stabilization; at least one output result corresponding to in vitro risk stratification analysis, prognostic correlation analysis, and intervention screening is used as the chromatin loop identification result.
[0015] A second aspect of this application provides an HPV integrated neochromatin loop identification device, the device comprising an acquisition module and a processing module, wherein, The acquisition module is used to acquire interaction sequencing read data and determine the HPV reference sequence and host reference sequence based on the interaction sequencing read data. The interaction sequencing read data includes sequencing read data, Hi-C data, ChIA-PET data, HiChIP data, and PLAC-seq data. HPV integration breakpoints are identified based on anomaly alignment features, including chimeric read features, split read features, and paired end anomaly features. Based on the breakpoint direction of the HPV integration breakpoint, a breakpoint-aware extended reference set is obtained by constructing fusion contigs between the HPV reference sequence and the host reference sequence. The interaction sequencing read data is compared with the breakpoint-aware extended reference set. Based on the alignment results and breakpoint-aware alignment constraints, anchor point correction and filtering operations are performed on the target reads to obtain high-confidence cross-genome anchor point pairs. The target reads are reads falling within a preset breakpoint-aware region, which includes fusion regions and breakpoint neighborhood windows. Breakpoint-aware alignment constraints include breakpoint neighborhood consistency constraints, direction consistency constraints, and unique alignment constraints.
[0016] The processing module is used to cluster high-confidence cross-genomic anchor pairs to obtain a candidate loop set, and select at least one target HPV-loop from the candidate loop set based on preset screening rules; calculate the newborn score corresponding to the target HPV-loop based on a preset newborn determination algorithm, and output the chromatin loop identification result according to the newborn score.
[0017] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described above.
[0018] A fourth aspect of this application provides a non-transitory computer-readable storage medium storing a computer program, the computer program being executed by a processor using any of the methods described above.
[0019] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. Acquire interaction sequencing read data and determine HPV and host reference sequences; identify HPV integration breakpoints based on chimeric read characteristics, split read characteristics, and abnormal pairing end characteristics; construct fusion contigs based on breakpoint directions to obtain a breakpoint-aware extended reference set; align interaction sequencing read data with the breakpoint-aware extended reference set, and obtain high-confidence cross-genome anchor pairs based on breakpoint-aware alignment constraints; cluster anchor pairs and screen target HPV-loops; calculate the nascent score of target HPV-loops and output chromatin loop identification results, thereby achieving stable and accurate identification of nascent chromatin loops in the context of HPV integration, reducing interference from cross-genome interaction mispairing identification results, and improving the accuracy and stability of nascent chromatin loop identification.
[0020] 2. Based on heterozygous SNP phasing, reads supporting the target HPV-loop are assigned to haplotypes, and loop allelic bias is calculated. Loop allelic bias and allele-specific expression from RNA-seq are jointly modeled, and a 3D-haplotype regulatory event set is output. Based on the 3D-haplotype regulatory event set, a 3D-haplotype consistency score is calculated to characterize the synergistic consistency between the spatial interaction structure and allele-specific transcriptional regulation of the target HPV-loop, so that the identification results of the newly formed chromatin loop can further reflect its potential allele regulatory effect.
[0021] 3. Calculate the target gene loop coupling expression index, specifically including: performing motif enrichment analysis on the HPV-terminal anchor and host-terminal anchor of the denovoHPV-loop, and screening a set of transcription factors that meet the preset conditions; constructing a protein-protein interaction network based on the transcription factor set to obtain a co-transcription factor module, and outputting the target gene loop coupling expression index based on the transcription factor module, thereby characterizing the coupling relationship between the denovoHPV-loop and the expression regulation of target genes at the structural level, so that the identification results of the newly formed chromatin loop can further reflect its potential transcriptional regulatory effect and biological function. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating a method for identifying HPV integrated neochromatin loops provided in an embodiment of this application. Figure 2 This is a schematic diagram of a module of an HPV integrated neochromatin loop identification device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0023] Explanation of reference numerals in the attached figures: 21. Acquisition module; 22. Processing module; 301. Processor; 302. Communication bus; 303. User interface; 304. Network interface; 305. Memory. Detailed Implementation
[0024] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0025] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.
[0026] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0027] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0028] Please refer to Figure 1 The flowchart illustrates a method for identifying HPV integrated neochromatin loops provided in an embodiment of this application. The flowchart mainly includes the following steps: S101 to S106.
[0029] Step S101: Obtain interaction sequencing read data and determine the HPV reference sequence and host reference sequence based on the interaction sequencing read data.
[0030] Specifically, interaction sequencing reads are first acquired, including but not limited to ChIA-PET, Hi-C, HiChIP, and PLAC-seq data. Preprocessing is performed on the acquired raw interaction sequencing data (FASTQ), including adapter sequence removal and quality control. For ChIA-PET, HiChIP, or PLAC-seq data, linker or bridge adapter sequences are identified and removed. Tags at both ends of each interaction event are extracted and labeled tag1 and tag2, respectively. Tag1 and tag2 are then treated as independent reads and aligned separately. For Hi-C data, read pairs are screened based on restriction enzyme sites and valid pair rules to retain valid linker fragments that meet the spatial interaction criteria.
[0031] After preprocessing the interaction sequencing read data, HPV and host reference sequences were determined based on the data. The steps for determining the HPV reference sequence were as follows: Based on the analysis object and detection target of the interaction sequencing read data, an HPV genome reference sequence (HPVreferencegenome) corresponding to the interaction sequencing read data was selected as the viral-side reference sequence for identifying viral-origin reads, chimeric reads, and cross-genome connectivity. The steps for determining the host reference sequence were as follows: Based on the biological sample source corresponding to the interaction sequencing read data, a host genome reference sequence (hostreferencegenome) consistent with the species of the biological sample was selected as the host-side reference sequence for host-side alignment, integration breakpoint localization, and subsequent chromatin loop analysis. The HPV reference sequence and the host reference sequence together constitute the basic reference sequence system for subsequent anomaly identification and breakpoint-aware extended reference construction.
[0032] Step S102: Identify HPV integration breakpoints based on anomaly comparison features; the anomaly comparison features include chimeric read features, split read features, and paired end anomaly features.
[0033] Specifically, joint analysis of alignment results of interacting sequencing reads is performed using whole-genome sequencing (WGS) or exome sequencing (WES) data, or by directly collecting and locating anomalous alignment signals based on cross-genome chimeric tags detected in the interacting sequencing read data, to identify candidate HPV integration breakpoints. During the identification process, the genomic coordinates and corresponding alignment direction of the integration breakpoint in the host reference sequence are determined by analyzing the alignment distribution of anomalous alignment reads on the host and HPV genomes. Simultaneously, the genomic coordinates and corresponding alignment direction of the integration breakpoint in the HPV reference sequence are also determined. This forms breakpoint location and orientation information characterizing the connection between HPV and the host genome, serving as the basis for subsequent breakpoint-aware extended reference construction.
[0034] Step S103: Based on the breakpoint direction of the HPV integrated breakpoint, a breakpoint-aware extended reference set is obtained by constructing a fusion contig between the HPV reference sequence and the host reference sequence.
[0035] Specifically, based on the breakpoint information in step S102, a host-HPV fusion contig is constructed according to the breakpoint direction based on the HPV reference sequence and the host reference sequence. The steps are as follows: Flanking sequences of length Lh on the left and right sides of the host breakpoint and flanking sequences of length Lv on the left and right sides of the HPV breakpoint are extracted respectively, for example, Lh and Lv are 200-5000bp; and the flanking sequences are spliced after forward / reverse complementarity according to the breakpoint directions of the host and HPV ends to generate at least one fusion connection sequence; in one embodiment, candidate connection sequences covering different connection directions can be generated simultaneously, for example... ⊕ , ⊕ And its reverse complementary form. Then, an index is built for all candidate connection sequences to form a breakpoint-aware extended reference set, and the interaction reads are aligned to the breakpoint-aware extended reference set.
[0036] Step S104: The interaction sequencing read data is compared with the breakpoint-aware extended reference set. Based on the comparison results and the breakpoint-aware alignment constraints, the target read is subjected to anchor point correction and filtering operations to obtain high-confidence cross-genome anchor point pairs.
[0037] Specifically, after aligning the interacting sequencing reads to the breakpoint-aware extended reference set, anchor point correction and filtering operations are performed on the target reads based on the alignment results and breakpoint-aware alignment constraints. The target reads are those falling within a preset breakpoint-aware region, which includes a fusion region and a breakpoint neighborhood window. Breakpoint-aware alignment constraints include breakpoint neighborhood consistency constraints, directional consistency constraints, and unique alignment constraints.
[0038] The breakpoint neighborhood consistency constraint is used to limit the spatial location consistency and connectivity consistency of the alignment results within the breakpoint neighborhood. This means that the alignment boundary point of the cross-breakpoint reads on the fusion contig should fall within a breakpoint neighborhood window centered on the breakpoint location with a preset threshold δ in width, and the alignment directions of the two corresponding cross-breakpoint reads should be consistent with the construction direction of the fusion connection sequence. δ is a preset window parameter, with a value range of, for example, 10–50 bp, used to limit the consistency determination range of the breakpoint neighborhood.
[0039] Directional consistency constraints are used to limit the consistency of paired or split reads in alignment direction. This means that in the same candidate breakpoint event, the alignment direction combination of different supporting reads should be consistent with the breakpoint direction and the connection direction of the fusion contig, so as to exclude non-genome connections caused by directional conflicts.
[0040] The unique alignment constraint is used to reduce the risk of mismatch caused by repeated sequences or low-complexity sequences. It means that when there are multiple alignment results for a read segment, the read segment is retained only if the difference in alignment score between the best and second-best alignment results is greater than a preset threshold. When the difference in alignment score is less than the preset threshold, the corresponding read segment is removed or its support is reduced according to the preset weight allocation strategy.
[0041] Taking the breakpoint neighborhood consistency constraint as an example, the process of determining whether the alignment result meets the breakpoint neighborhood consistency constraint includes: determining whether the alignment positions and alignment directions of the two segments corresponding to the cross-breakpoint reads in the alignment result are consistent with the breakpoint connection relationship, and determining whether the two alignment positions are within the preset breakpoint neighborhood window; when the alignment boundary point of the cross-breakpoint reads on the fusion contig falls within the neighborhood window of the breakpoint ±δ, and the alignment directions of the two segments are consistent with the fusion connection direction, it is further confirmed that one end of the reads is labeled on the HPV side and the other end is labeled on the host side, or one end is labeled on the HPV side of the fusion contig and the other end is labeled on the host reference sequence.
[0042] Step S105: Cluster the high-confidence cross-genomic anchor pairs to obtain a candidate loop set, and select at least one target HPV-loop from the candidate loop set based on preset screening rules.
[0043] Specifically, the high-confidence cross-genomic anchor pairs obtained from step S104 are mapped to a unified anchor window system. Each anchor point is windowed and merged within the coordinate space of the HPV reference sequence, fusion contig, and host reference sequence using a preset window size, thereby converting the scattered anchor coordinates into clusterable anchor window identifiers. Subsequently, high-confidence PETs are clustered using anchor windows as clustering units. PETs with the same or adjacent anchor window combinations are grouped into the same interaction cluster, and the anchor window pair corresponding to each interaction cluster is treated as a candidate loop entry, thus forming a candidate loop set. Each candidate loop entry is associated with at least one HPV-side anchor window and one host-side anchor window, and is also associated with the PET set and its statistical characteristics supporting the candidate loop.
[0044] After forming a candidate loop set, HPV-loops are selected from the candidate loop set based on preset screening rules, and the selected HPV-loops are used as candidates for target HPV-loops. The preset screening rules include at least structural location rules and evidence strength rules. Structural location rules are used to determine whether a candidate loop possesses the structural characteristics of a cross-genome loop. Specifically, at least one anchor point of the candidate loop must be located in the HPV reference sequence, the HPV side fragment of a fusion contig, or a breakpoint neighborhood window, and the other anchor point must be located in a host genome region, thus ensuring that the candidate loop has a cross-genome connection morphology of "viral end anchor point - host end anchor point". Evidence strength rules are used to determine whether a candidate loop has sufficient supporting evidence. Specifically, the number of valid PETs supporting the candidate loop must not be less than a preset threshold k, to avoid low-support false positives caused by accidental alignments or local noise. k can be set according to sequencing depth, sample noise level, or preset detection sensitivity requirements, so that the screening rules can both suppress false positives and ensure the detection capability of true HPV-loops.
[0045] Step S106: Calculate the newborn score corresponding to the target HPV-loop based on the preset newborn determination algorithm, and output the chromatin loop identification result according to the newborn score.
[0046] Specifically, the HPV-loop remodeling index, 3D-haplotype consistency score, and target gene loop coupling expression index were calculated respectively, and at least one of the HPV-loop remodeling index, 3D-haplotype consistency score, and target gene loop coupling expression index was used as the newborn score.
[0047] In one possible implementation, step S106 further includes: selecting high-confidence target HPV-loops from the target HPV-loops based on the number of supporting evidences for the target HPV-loop and in conjunction with a significance assessment calculation strategy; the significance assessment calculation strategy includes significance calculation based on a distance-dependent background model and significance calculation based on a permutation test; obtaining a set of randomized control loops that match the distance or chromatin state corresponding to the set of high-confidence target HPV-loops; and comparing the set of high-confidence target HPV-loops with the reference host loop library and the randomized control loop set, respectively, to calculate the neonatal score corresponding to the target HPV-loop.
[0048] Specifically, after identifying the target HPV-loops, to obtain a set of high-confidence HPV-loops suitable for neoprediction, support, evidence consistency score, and significance p-value are further calculated for each HPV-loop. Multiple test error rate control is then performed to achieve convergence from "structural candidate" to "statistically high confidence." Support is the number of valid PETs supporting the HPV-loop, used to characterize the strength of direct evidence for that loop. The statistical range of support can be limited to valid PETs that meet the breakpoint-aware alignment constraint and pass the deduplication process. Evidence consistency scoring characterizes whether supporting evidence for the same HPV-loop exhibits stable consistency across key consistency dimensions, specifically including breakpoint neighborhood consistency, PET direction consistency, fragment length distribution consistency, and repeatability. Breakpoint neighborhood consistency measures whether HPV-side anchors stably cluster in the breakpoint neighborhood or whether HPV-side fragments of the fusion contig are consistent with the breakpoint connection. PET direction consistency measures whether the directional combination of supporting evidence aligns with the breakpoint direction and the connection direction of the fusion contig. Fragment length distribution consistency measures whether the length of the inserted fragments corresponding to the supporting evidence conforms to the statistical distribution within the same library or candidate event. Repeatability measures whether there is excessive repetition of supporting evidence or non-independent repetition due to low-complexity sequences. Evidence consistency scores can be used as filters or weighting factors in subsequent significance assessments, allowing HPV-loops with higher consistency to have higher credibility when support levels are similar.
[0049] A set of randomized control loops matching the high-confidence target HPV-loop in terms of distance or chromatin state is obtained to construct a background control system comparable to the high-confidence target HPV-loop in key structural statistical features. The randomized control loop set is generated while maintaining constraints on the distance distribution between host-end anchor points and the stratification of chromatin state in the anchor point regions. This ensures that the randomized control loops are consistent with the high-confidence target HPV-loop in terms of spatial scale and chromatin environment, thus characterizing the distribution features of background loops that meet the same structural conditions without introducing HPV integration effects.
[0050] Based on this, the set of high-confidence target HPV-loops was compared with a reference host loop library and a randomized control loop set. The reference host loop library was used to determine whether the host-end anchor points of the high-confidence target HPV-loops could be explained by existing host 3D chromatin loop structures, thus excluding pre-existing host background loop events. The randomized control loop set was used to estimate the background null distribution based on permutation statistics, thereby assessing the probability level of observing high-confidence target HPV-loops under conditions of distance and chromatin state matching. Through this dual-comparison analysis, event-level determination results or baseline scores were obtained to characterize the nascent attributes of target HPV-loops. These event-level determination results serve as one of the basic inputs for subsequent calculations of the HPV-loop remodeling index, 3D-haplotype consistency score, and target gene loop coupling expression index.
[0051] In one possible implementation, step S106 further includes: calculating the HPV-loop remodeling index, the 3D-haplotype consistency score, and the target gene loop coupling expression index, respectively, and using at least one of these scores as the newborn score. Specifically, calculating the HPV-loop remodeling index includes: using a dual-control method to calculate the denovo score and summing the denovo scores to calculate the HPV-loop remodeling index; the dual-control method is a reference loop library control method and a matched random control method. Calculating the 3D-haplotype consistency score specifically includes: assigning reads supporting the target HPV-loop to haplotypes based on heterozygous SNP phasing and calculating loop allelic bias; jointly modeling the loop allelic bias with allele-specific expression from RNA-seq and outputting a set of 3D-haplotype regulatory events; and calculating the 3D-haplotype consistency score based on the set of 3D-haplotype regulatory events. The calculation of target gene loop coupling expression indices specifically includes: performing motif enrichment analysis on the HPV-terminal and host-terminal anchors of the denavoHPV-loop, and screening a set of transcription factors that meet preset conditions; constructing a protein-protein interaction network based on the transcription factor set to obtain a co-transcription factor module, and outputting the target gene loop coupling expression indices based on the transcription factor module; the target gene includes at least one of MYC, CASC19, and FAM81A; the co-transcription factor module includes at least one of FOS, CEBPG, SP1, or CTCF.
[0052] Specifically, the HPV-loop remodeling index, 3D-haplotype consistency score, and target gene loop coupling expression index were calculated based on the overall remodeling degree at the structural level, the three-dimensional regulatory consistency at the allele level, and the target gene expression coupling effect at the functional level. At least one of the above scoring indicators was used as the newborn score to quantitatively characterize the newborn chromatin loops induced by HPV integration.
[0053] In calculating the HPV-loop remodeling index, a dual-control method is introduced to quantify the event-level regeneration of each target HPV-loop. The dual-control method includes two complementary control pathways: a reference loop library comparison and a matched randomized controlled trial. The reference loop library comparison compares the target HPV-loop with a pre-constructed reference host loop library to determine whether the host-end anchor points and their spatial interactions can be explained by the existing three-dimensional chromatin structure of the host, thereby identifying possible host background loops. The matched randomized controlled trial, while maintaining consistency in the genomic distance distribution between anchor points and the chromatin state stratification of the anchor points, constructs a set of randomized controlled loops. Based on this set, the null distribution is estimated to assess the statistical probability of observing the target HPV-loop under the same structural conditions. Combining the results of the two types of controls mentioned above, a denovo score was calculated for each target HPV-loop to reflect the degree of neon formation relative to the host background structure. After obtaining the denovo scores for each target HPV-loop, the denovo scores were aggregated at the sample or event level to form the HPV-loop remodeling index, thereby comprehensively characterizing the intensity and degree of remodeling and neon formation introduced by HPV integration events on the three-dimensional chromatin loop structure in a given sample. The HPV-loop remodeling index can serve as a form of neon scoring, and together with the 3D-haplotype consistency score and target gene loop coupling expression indicators, it constitutes a multi-dimensional and scalable neon scoring system to support subsequent sample stratification analysis, mechanism analysis, and in vitro intervention screening applications. The formula for calculating the HPV-loop remodeling index is as follows:
[0054] in, This indicates that the host-end anchor point of the candidate HPV-loop can be matched with the same or neighboring host loop in the reference host loop library; This indicates that no matching or neighboring host loops were found; This indicates the observed support for the candidate HPV-loop; Indicates the first Support of a randomized control loop; Indicates the size of the set of randomized control loops; This indicates the number of elements in the set that satisfy the condition; the smaller the empirical p-value, the more difficult it is for the observation support to be interpreted by the matching zero distribution. This represents the event-level neonatal score of the candidate HPV-loop; Used to suppress candidate loops that can be interpreted by the reference host loop library; The empirical p-values of randomized controlled trials are converted into a monotonically increasing significance level. The values range from 0 to 1 and are obtained through random comparison and statistics. In actual implementation, this is usually achieved by increasing... To improve resolution; Indicates the sample-level HPV-loop remodeling index; This indicates the number of HPV-loops included in the aggregation, which can typically be the number of elements in the high-confidence HPV-loop set or the number of elements judged to be new. Indicates the first The event-level newborn score of each HPV-loop; Indicates the first The aggregate weight of each HPV-loop is determined by the implementation strategy. If no weighting is introduced, it can be set as follows: If weighting is introduced, then... Monotonic correlation with support or evidence consistency scores of HPV-loop is used to reflect differences in the strength of evidence, and the weights can be further normalized to maintain comparability between different samples.
[0055] In calculating the 3D-haplotype consistency score, spatial interaction evidence of the target HPV-loop is linked to allele-level genetic background and expression bias, enabling new structural events to be further characterized as allele-specific three-dimensional regulatory events. This calculation process revolves around the logic of "read haplotype attribution—loop allele bias quantification—joint modeling with expression ASE—consistency score output," ensuring that each intermediate result can be traced back to specific supporting evidence of the target HPV-loop and can be quantified comparably across samples.
[0056] When assigning reads supporting a target HPV-loop to haplotypes based on heterozygous SNP phasing and calculating loop allelic bias, the phasing results corresponding to the target sample are first obtained, ensuring that each heterozygous SNP has a clear haplotype assignment relationship. Subsequently, in the supporting evidence set for the target HPV-loop, interacting sequencing reads or read tags containing heterozygous SNPs are screened, and based on the allelic information of the heterozygous SNPs they cover, the corresponding reads are assigned to either the first or second haplotype. To avoid instability in assignment due to insufficient coverage by a single SNP, for the same target HPV-loop, multiple available heterozygous SNP assignment evidence can be aggregated within the host-terminal anchor neighborhood and the spatially interacting fragment range. Consistency checks are performed on the read assignment results, ensuring that the supportive read counts assigned to different haplotypes form loop allelic support. Based on loop allelic support, loop allelic bias is constructed to characterize the difference in spatial interaction strength between the target HPV-loop on two haplotypes. The value of loop allelic bias can be obtained by the ratio of the support of the two haplotypes, and a smoothing term can be introduced when the support is low to suppress extreme ratios and maintain comparability under different sequencing depth conditions.
[0057] When jointly modeling loop allelic bias and allele-specific expression from RNA-seq to output a 3D-haplotype regulatory event set, the allele-specific expression results of target genes potentially regulated by the target HPV-loop are first calculated based on RNA-seq data, establishing the ASE bias direction and strength based on the expression support of each target gene across two haplotypes. Subsequently, the loop allelic bias of the target HPV-loop is correlated with the ASE bias of the target genes, establishing a testable consistency relationship between loop enhancement and expression enhancement within the same haplotype. In the joint modeling, loop allelic bias is used as a structural evidence variable, and ASE bias as a transcriptional evidence variable. Information such as the amount of supporting evidence for the target HPV-loop, anchor confidence, or neologism determination basis is used as confidence weights to output the regulatory consistency determination result for each "target HPV-loop—target gene" pairing. Pairs meeting preset consistency criteria are converged into 3D-haplotype regulation events, and these events are correlated with their corresponding haplotype orientation, loop bias strength, and expression bias strength, thus forming a set of 3D-haplotype regulation events. When calculating the 3D-haplotype consistency score based on this set, the consistency judgment results in the event set are aggregated and quantified at the sample level or the target HPV-loop level, making them a form of output for the emerging score. Specifically, for each target HPV-loop, the number of associated 3D-haplotype regulation events, the event consistency strength, and the stability of the consistency orientation are statistically analyzed. Weights related to the amount of supporting evidence for the target HPV-loop are introduced to weight and aggregate the event-level consistency contributions, thereby obtaining the 3D-haplotype consistency score for the target HPV-loop. The formula for calculating the 3D-haplotype consistency score is as follows:
[0058] in, This is due to ASE isotropic bias. This is due to the equipotential bias of the ring road. This indicates the count of allele-specific reads in RNA-seq that fall on the target gene or its equivalent transcription unit associated with the target HPV-loop, and is assigned to h1 according to the same set of phase results; This represents the corresponding count belonging to h2; This represents a smoothing constant used to avoid the logarithm becoming undefinable due to a count of zero. Its value can be a very small positive number and can be fixed as a constant to ensure cross-sample comparability. Similarly, the logarithmic ratio characterizes the direction and strength of allelic bias on the expression side, making it consistent with the bias on the loop side. In terms of numerical form, they can be directly modeled together. This represents the number of valid PETs that support the target HPV-loop after being phased by heterozygous SNPs and assigned to haplotype-1 (h1). The counting method is to count the PETs that satisfy the deduplication and alignment constraints within the anchor point windows at both ends of the target HPV-loop. This indicates the corresponding valid PET count belonging to haplotype-2 (h2). As a directional consistency factor, For a sign function, when Time to take This indicates that the direction of the loop bias is consistent with the direction of the expressed bias. Time to take Indicates opposite directions, when Time to take This indicates that at least one side has no discernible bias; Indicates the 3D-haplotype consistency score. This represents the Sigmoid mapping, used to compress the joint quantization result to... arrive The scoring range is used to threshold the output; The amplification factor representing the directional consistency and bias strength on the score can be obtained by fitting to the training samples or by setting empirically. This represents the directional consistency factor, which increases the score when directions are consistent and decreases the score when directions are opposite. The consistency strength term is taken as the smaller of the bias strength between the loop side and the expression side, which is used to reflect the joint constraint that a high consistency event is formed only when both sides are significantly biased and in the same direction. This represents the uncertainty penalty coefficient; This represents an uncertainty term, used to reflect instability caused by undercounting or insufficient evidence of phase separation. Its value can be constructed from the count size, for example, by letting... Follow and The decrease in coverage leads to an increase in coverage, which in turn causes the scores of low-coverage events to shrink to the neutral range.
[0059] When calculating the target gene loop coupling expression index, during motif enrichment analysis of the HPV-end anchors and host-end anchors of the de novo HPV-loop, the anchor regions at both ends of each loop are first located in the obtained de novo HPV-loop set. These anchor regions are then uniformly mapped to the corresponding reference sequence coordinate system, so that the HPV-end anchors correspond to the coordinates of the HPV reference sequence or the HPV side fragment of the fusion contig, and the host-end anchors correspond to the coordinates of the host reference sequence. To ensure comparability between different samples and different loops, the anchor region is typically expanded around the anchor center by a pre-defined neighborhood window, and the window length is uniformly set, allowing subsequent motif scanning and statistical enrichment to be performed on a consistent sequence background. Subsequently, the HPV-terminal anchor region set and the host-terminal anchor region set were used as the sequence sets to be analyzed. Motif patterns of candidate transcription factor binding sites were scanned, and the frequency and enrichment degree of each motif in the sequence sets were statistically analyzed. Simultaneously, a matching background sequence set was introduced to eliminate biases caused by differences in GC content, comparability, and sequence complexity, thus obtaining the HPV-terminal enriched motif set and the host-terminal enriched motif set. Furthermore, the enrichment results from both ends were jointly screened, prioritizing the retention of transcription factors corresponding to motifs showing significant enrichment trends in both the HPV and host ends. This strengthens the consistency of the interpretation of "common binding potential across both ends of the genome." The screening results were then constrained to be consistent with sample expression information or publicly available binding information to avoid introducing low-expression or biologically unfeasible candidate factors into subsequent network modeling.
[0060] When screening a set of transcription factors that meet predefined criteria, these criteria are used to converge enrichment signals at the motif level into an interpretable and verifiable set of candidate transcription factors. The predefined criteria consider at least the consistency of enrichment at both ends, the significance of enrichment, and the strength of evidence related to the stability of the denovoHPV-loop. Consistency of enrichment at both ends ensures that candidate factors have binding site enrichment characteristics at both the HPV-terminal and host-terminal anchor sites, making it reasonable for them to have the capacity for cross-terminal cooperative mediation of loop formation or stability. Significance of enrichment avoids misjudging occasional fluctuations in motif counts as genuine signals, thereby improving the robustness of the candidate set. The strength of evidence links the screening of candidate factors with the amount of supporting evidence for the denovoHPV-loop, anchor consistency, or the basis for newborn determination, prioritizing more consistent and stronger motif signals in high-confidence loops for inclusion in the candidate set. After meeting the above conditions, a set of transcription factors is obtained, and this set is associated with the anchor region of each denovoHPV-loop and stored to form a traceable correspondence of "loop - anchor at both ends - candidate transcription factor", providing a clear source of node set for subsequent protein interaction network modeling.
[0061] When constructing a protein-protein interaction network based on a set of transcription factors to obtain co-transmission factor modules, each transcription factor in the set is treated as a network node. Interaction edges are established between nodes based on a pre-defined protein-protein interaction relation library or known interaction evidence, thus forming a protein-protein interaction network. To make the network structure more closely resemble the formation mechanism of the denovo HPV-loop, interaction edges can be further weighted by combining sample expression levels, consistency of motif enrichment intensity at both ends, and the strength of loop support evidence, so that interaction relationships more likely to be co-transmission factors in specific sample or loop contexts receive higher weights. Subsequently, module identification processing is performed on the protein-protein interaction network, causing the network to converge topologically into several closely interacting and functionally related co-transmission factor subnetworks. Subnetworks that meet pre-defined size, connectivity, or centrality conditions are identified as co-transmission factor modules. When a co-transcription factor module contains at least one of FOS, CEBPG, SP1, or CTCF, its motif enrichment distribution characteristics at the HPV and host ends, as well as its interaction topological position, can be further combined to interpret and constrain the potential role of the module in loop stability. This allows the module to not only have statistical clustering significance but also to provide mechanistic guidance on "why this denovo HPV-loop can be formed and exist stably".
[0062] When outputting target gene loop-coupled expression indices based on the co-transcription factor module, the host-end anchor of the denavoHPV-loop is associated with gene annotations or regulatory element annotations to determine the set of target genes that may be coupled to each loop, including at least one candidate target gene from MYC, CASC19, and FAM81A. Subsequently, the expression evidence of the co-transcription factor module and the target genes is coupled and modeled so that the co-synergistic strength at the module level can be mapped to a quantifiable index of the target gene expression response. Specifically, the motif enrichment intensity of the co-transcription factor module in the neighborhood of the HPV-end and host-end anchors, the network centrality or connectivity of key nodes within the module, and the strength of supporting evidence for the loop are used as joint structural and mechanistic evidence. Changes in the expression of the corresponding target gene in RNA-seq or other transcriptome data, allelic bias, or consistency characteristics with the loop are used as transcriptional evidence, thus forming the target gene loop-coupled expression index. The calculation formula for the target gene loop-coupled expression index is as follows: [Formula omitted for brevity].
[0063] in, Indicates the first Transcription factors on the denovoHPV-loop The intensity of dual-end co-enrichment; This indicates transcription factors within the HPV end-anchor window. The motif enrichment score can be obtained by comparing the motif scan count of the anchor sequence with the background sequence count, and can be further combined with factors such as sequence GC for background correction. Indicates transcription factors within the host-terminal anchor window The motif enrichment score; the principle of using geometric mean requires simultaneous enrichment at both ends. Only when the enrichment is high will the co-enrichment intensity be significantly reduced, as non-enrichment at either end will significantly suppress the co-enrichment intensity. Indicates transcription factor Expression constraint factors; Indicates transcription factor The expression intensity in the sample can be derived from the normalized expression level of RNA-seq; This indicates a high expression screening threshold; Indicates the scale parameter; The Sigmoid mapping smoothly transforms the expression intensity into weights of 0 to 1 near the threshold. Indicates transcription factor Normalized centrality weights in the co-transcription factor module; This indicates that in a set of transcription factors Transcription factors in the constructed protein-protein interaction network The centrality score can be any of degree centrality, betweenness centrality, or eigenvector centrality, as long as it reflects the contribution of key nodes in the collaborative network. Used to normalize centrality, making weights comparable across different samples or module sizes; Indicates the first A denovoHPV-loop targets genes The intensity of the readout expression; Indicates target gene The change in expression can be obtained from the difference between the treatment group and the control group, or the difference before and after the intervention, or the change of the sample relative to the baseline; Indicates the first A denovoHPV-loop and target gene Loop coupling expression index; This represents a set of transcription factors that satisfy the condition of being enriched at both ends and highly expressed. Indicates transcription factor The centrality weight in the collaborative module is used to reflect the contribution of the collaborative network structure; Indicates transcription factor The synergistic enrichment intensity at both ends of the HPV-loop anchor point is used to reflect the binding potential on the anchor point side; This indicates transcription factor expression constraint factors, used to reflect whether the transcription factor in the sample has actual regulatory capacity; It represents the readout intensity of target gene expression and is used to couple structural events and transcriptional outcomes into the same dimensional score.
[0064] In one possible implementation, step S106 further includes: performing in vitro risk stratification analysis based on chromatin loop identification results; the in vitro risk stratification analysis is used to quantify and stratify the carcinogenic potential of HPV-positive samples or cell models and provide reference information; performing prognostic correlation analysis based on chromatin loop identification results; the prognostic correlation analysis is used to analyze HPV-loop coupling activation events corresponding to different target genes and use HPV-loop coupling activation events as candidate prognostic indicators; performing intervention screening operations based on chromatin loop identification results; the intervention screening operations are used to screen candidate intervention methods to inhibit HPV-loop formation or stabilization; and using at least one output result corresponding to the in vitro risk stratification analysis, prognostic correlation analysis, and intervention screening operations as the chromatin loop identification result.
[0065] Specifically, when performing in vitro risk stratification analysis based on chromatin loop identification results, the set of newly formed chromatin loop events and their scoring information corresponding to the sample or cell model are used as the source of stratification features. This allows in vitro stratification to reflect the degree of three-dimensional chromatin structural remodeling and the strength of regulatory coupling induced by HPV integration. In in vitro risk stratification analysis, at least one of the following indicators is used as the stratification basis: the newborn score or its constituent indicators, including at least one of the HPV-loop remodeling index, 3D-haplotype consistency score, and target gene loop coupling expression indicators. The number of denovo HPV-loops, spatial distribution density, and the strength of loop supporting evidence can be combined to form sample-level structural load characteristics. Based on the above stratification features, HPV-positive samples or cell models are quantitatively stratified to obtain stratification results for different risk levels or carcinogenic potential levels. High-risk stratification results correspond to higher structural remodeling intensity, more obvious allelic regulatory consistency, or stronger target gene loop coupling expression effects, thus providing reference information for in vitro experimental grouping, model selection, or subsequent intervention validation.
[0066] When performing prognostic correlation analysis based on chromatin loop identification results, a candidate indicator system for prognostic assessment is constructed around HPV-loop coupling activation events corresponding to different target genes. Specifically, the host-terminal anchor of the de novo HPV-loop is associated with gene annotation and regulatory element annotation to identify HPV-loops spatially coupled to target gene regions, and HPV-loop coupling activation events corresponding to target genes are extracted as event-level candidate prognostic indicators. HPV-loop coupling activation events can be defined by the consistency between the structural and transcriptional sides, so that they simultaneously reflect loop evidence and expression effects. Structural evidence includes at least the quantity of loop supporting evidence, significance control results, and basic results for newborn determination. Transcriptional evidence includes at least changes in target gene expression, allele-specific expression bias, or mechanistic evidence consistent with co-transcription factor modules. The above event-level candidate prognostic indicators are statistically compared among samples or correlated with follow-up outcomes to screen out HPV-loop coupling activation events that are significantly associated with prognosis, and output as candidate prognostic indicators to support the prognostic interpretation and risk indication of loop events related to specific target genes.
[0067] When conducting intervention screening based on chromatin loop identification results, these results are used as a quantitative basis for intervention readout, screening candidate interventions to inhibit or stabilize HPV-loops. Specifically, based on the mechanistic analysis results of the denovoHPV-loop and its anchor region, key candidate factors that may be involved in loop formation or stabilization are identified, and suppression, knockdown, or editing strategies targeting these factors are used as candidate interventions. Simultaneously, based on the structural characteristics of the HPV-loop, an intervention evaluation index system is constructed, centered on the HPV-loop remodeling index, neonatal score, or target gene loop coupling expression indicators, allowing for comparison of different candidate interventions under a unified readout index. Interaction sequencing reads are re-acquired from samples or cell models after each candidate intervention, and the chromatin loop identification process is executed again. Changes in the denovoHPV-loop set, neonatal score, and target gene coupling activation events before and after intervention are compared, thereby screening candidate interventions that can significantly reduce the probability of HPV-loop formation, weaken loop stability, or inhibit the expression effect of key target genes. The screening results are used as the output of the intervention screening operation.
[0068] When using at least one output result from in vitro risk stratification analysis, prognostic correlation analysis, and intervention screening operations as chromatin circuit identification results, the chromatin circuit identification results further include application layer output content in addition to the basic identification output. The application layer output content includes at least one of the following: stratification label or stratification score of in vitro risk stratification, candidate prognostic indicators obtained from prognostic correlation analysis, and candidate intervention methods obtained from intervention screening operations. This allows the output results of this application to form a closed loop that directly addresses application needs such as in vitro stratification, prognostic research, and intervention verification.
[0069] Please refer to Figure 2 This illustration shows a schematic diagram of a module for identifying HPV integrated neochromatin loops according to an embodiment of this application. The device includes an acquisition module 21 and a processing module 22, wherein... The acquisition module 21 is used to acquire interaction sequencing read data and determine the HPV reference sequence and host reference sequence based on the interaction sequencing read data. The interaction sequencing read data includes sequencing read data, Hi-C data, ChIA-PET data, HiChIP data, and PLAC-seq data. HPV integration breakpoints are identified based on abnormal alignment features. The abnormal alignment features include chimeric read features, split read features, and paired end abnormal features. Based on the breakpoint direction of the HPV integration breakpoint, a breakpoint-aware extended reference set is obtained by constructing a fusion contig between the HPV reference sequence and the host reference sequence. The interaction sequencing read data is compared with the breakpoint-aware extended reference set. Based on the comparison results and breakpoint-aware alignment constraints, anchor point correction and filtering operations are performed on the target reads to obtain high-confidence cross-genome anchor point pairs. The target reads are reads that fall within a preset breakpoint-aware region. The preset breakpoint-aware region includes a fusion region and a breakpoint neighborhood window. The breakpoint-aware alignment constraints include breakpoint neighborhood consistency constraints, direction consistency constraints, and unique alignment constraints.
[0070] The processing module 22 is used to cluster high-confidence cross-genomic anchor pairs to obtain a candidate loop set, and select at least one target HPV-loop from the candidate loop set based on a preset screening rule; calculate the newborn score corresponding to the target HPV-loop based on a preset newborn determination algorithm, and output the chromatin loop identification result based on the newborn score.
[0071] In one possible implementation, the acquisition module 21 is used to determine whether the comparison result satisfies the breakpoint neighborhood consistency constraint. Specifically, it includes: determining whether the two comparison positions and comparison directions corresponding to the cross-breakpoint reading segment in the comparison result are consistent with the breakpoint connection relationship, and determining whether the two comparison positions are within the preset breakpoint neighborhood window range; if the two comparison positions and comparison directions corresponding to the cross-breakpoint reading segment in the comparison result are consistent with the breakpoint connection relationship, and the two comparison positions are within the preset breakpoint neighborhood window range, then the comparison result is determined to satisfy the breakpoint neighborhood consistency constraint.
[0072] In one possible implementation, the processing module 22 is used to cluster high-confidence cross-genome anchor pairs to obtain a candidate loop set, specifically including: clustering high-confidence cross-genome anchor pairs, and generating a candidate loop set while maintaining host-end anchor distance distribution constraints, chromatin state stratification constraints, and coverage stratification constraints.
[0073] In one possible implementation, the processing module 22 is used to calculate the newborn score corresponding to the target HPV-loop based on a preset newborn determination algorithm, and output the chromatin loop identification result according to the newborn score. Specifically, it includes: selecting high-confidence target HPV-loops from the target HPV-loops based on the number of supporting evidences for the target HPV-loop and in combination with a significance assessment calculation strategy; the significance assessment calculation strategy includes significance calculation based on a distance-dependent background model and significance calculation based on a permutation test; obtaining a set of random control loops that match the distance or chromatin state corresponding to the set of high-confidence target HPV-loops; comparing the set of high-confidence target HPV-loops with the reference host loop library and the random control loop set respectively to calculate the newborn score corresponding to the target HPV-loop.
[0074] In one possible implementation, the processing module 22 is used to calculate the newborn score corresponding to the target HPV-loop based on a preset newborn determination algorithm, specifically including: calculating the HPV-loop remodeling index, 3D-haplotype consistency score and target gene loop coupling expression index respectively, and using at least one of the HPV-loop remodeling index, 3D-haplotype consistency score and target gene loop coupling expression index as the newborn score.
[0075] In one possible implementation, the processing module 22 is used to calculate the HPV-loop remodeling index, specifically including: introducing a double-control method to calculate the denovo score, and summing the denovo scores to calculate the HPV-loop remodeling index; the double-control method is a reference loop library control method and a matched random control method.
[0076] In one possible implementation, the processing module 22 is used to calculate the 3D-haplotype consistency score, specifically including: classifying the reads supporting the target HPV-loop to haplotypes based on heterozygous SNP phasing, and calculating loop allelic bias; jointly modeling the loop allelic bias with allele-specific expression in RNA-seq, and outputting a set of 3D-haplotype regulatory events; and calculating the 3D-haplotype consistency score based on the set of 3D-haplotype regulatory events.
[0077] In one possible implementation, the processing module 22 is used to calculate the target gene loop coupling expression index, specifically including: performing motif enrichment analysis on the HPV end anchor and host end anchor of the denavoHPV-loop, and screening a set of transcription factors that meet preset conditions; constructing a protein interaction network based on the transcription factor set to obtain a co-transcription factor module, and outputting the target gene loop coupling expression index based on the transcription factor module; the target gene includes at least one of MYC, CASC19, and FAM81A; the co-transcription factor module includes at least one of FOS, CEBPG, SP1, or CTCF.
[0078] In one possible implementation, the processing module 22 is used to output chromatin loop identification results based on the neonatal score, specifically including: performing in vitro risk stratification analysis based on the chromatin loop identification results; the in vitro risk stratification analysis is used to quantify and stratify the carcinogenic potential of HPV-positive samples or cell models and provide reference information; performing prognostic correlation analysis based on the chromatin loop identification results; the prognostic correlation analysis is used to analyze HPV-loop coupling activation events corresponding to different target genes and use HPV-loop coupling activation events as candidate prognostic indicators; performing intervention screening operations based on the chromatin loop identification results; the intervention screening operations are used to screen candidate intervention methods to inhibit HPV-loop formation or stabilization; and using at least one output result corresponding to the in vitro risk stratification analysis, prognostic correlation analysis, and intervention screening operations as the chromatin loop identification results.
[0079] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0080] This application also provides an electronic device. (See reference...) Figure 3 , Figure 3This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: at least one processor 301, at least one communication bus 302, a user interface 303, at least one network interface 304, and a memory 305.
[0081] The communication bus 302 is used to enable communication between these components.
[0082] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0083] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0084] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.
[0085] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. (Refer to...) Figure 3 The memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an HPV integrated neochromatin loop identification application.
[0086] exist Figure 3 In the illustrated electronic device, the user interface 303 is primarily used to provide an input interface for the user and acquire user input data; while the processor 301 can be used to call the HPV integrated neochromatin loop identification application stored in the memory 305. When executed by one or more processors 301, the electronic device performs one or more of the methods described in the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0087] This application also provides a non-transitory computer-readable storage medium storing instructions. When executed by one or more processors, these instructions cause an electronic device to perform one or more of the methods described in the above embodiments.
[0088] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0089] In the various embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.
[0090] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0091] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0092] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0093] The above description is merely an exemplary embodiment disclosed in this application and should not be construed as limiting the scope of this application. Any equivalent changes and modifications made in accordance with the teachings of this application shall still fall within the scope of this application.
[0094] This application is intended to cover any variations, uses, or adaptations disclosed herein that follow the general principles disclosed herein and include common knowledge or customary technical means in the art that are not described in this application.
Claims
1. A method for identifying HPV-integrated neochromatin loops, characterized in that, The method includes: The interaction sequencing read data is obtained, and the HPV reference sequence and the host reference sequence are determined based on the interaction sequencing read data; the interaction sequencing read data includes sequencing read data, Hi-C data, ChIA-PET data, HiChIP data, and PLAC-seq data; HPV integration breakpoints are identified based on anomaly comparison features; the anomaly comparison features include chimeric read features, split read features, and paired end anomaly features. Based on the breakpoint direction of the HPV integrated breakpoint, a breakpoint-aware extended reference set is obtained by constructing a fusion contig between the HPV reference sequence and the host reference sequence. The interactive sequencing reads are compared with the breakpoint-aware extended reference set. Based on the comparison results and breakpoint-aware alignment constraints, anchor point correction and filtering operations are performed on the target reads to obtain high-confidence cross-genome anchor point pairs. The target reads are reads that fall within a preset breakpoint-aware region, which includes a fusion region and a breakpoint neighborhood window. The breakpoint-aware alignment constraints include breakpoint neighborhood consistency constraints, directional consistency constraints, and unique alignment constraints. Clustering of the high-confidence cross-genomic anchor pairs yields a candidate loop set, and at least one target HPV-loop is selected from the candidate loop set based on preset screening rules; The newborn score corresponding to the target HPV-loop is calculated based on the preset newborn determination algorithm, and the chromatin loop identification result is output according to the newborn score.
2. The method according to claim 1, characterized in that, Determining whether the comparison result satisfies the breakpoint neighborhood consistency constraint specifically includes: Determine whether the comparison positions and directions of the two segments corresponding to the cross-breakpoint reading segment in the comparison result are consistent with the connection relationship of the breakpoint, and determine whether the two comparison positions are within the preset breakpoint neighborhood window. If the comparison results show that the two comparison positions and the comparison direction of the cross-breakpoint reading segment are consistent with the breakpoint connection relationship, and the two comparison positions are within the preset breakpoint neighborhood window, then the comparison results are determined to satisfy the breakpoint neighborhood consistency constraint.
3. The method according to claim 1, characterized in that, The clustering of the high-confidence cross-genomic anchor pairs to obtain a candidate loop set specifically includes: Clustering is performed on the high-confidence cross-genome anchor pairs, and the candidate loop set is generated while maintaining constraints on host-end anchor distance distribution, chromatin state stratification, and coverage stratification.
4. The method according to claim 1, characterized in that, The step of calculating the neonatal score corresponding to the target HPV-loop based on a preset neonatal determination algorithm, and outputting the chromatin loop identification result based on the neonatal score, specifically includes: Based on the number of supporting evidences for the target HPV-loop, and in conjunction with a significance assessment calculation strategy, a high-confidence target HPV-loop is selected from the target HPV-loops; the significance assessment calculation strategy includes significance calculation based on a distance-dependent background model and significance calculation based on a permutation test; Obtain a set of random control loops that match the distance or chromatin state of the set of high-confidence target HPV-loops; The set of high-confidence target HPV-loops is compared with the reference host loop library and the set of random control loops to calculate the neonatal score corresponding to the target HPV-loop.
5. The method according to claim 1, characterized in that, The calculation of the newborn score corresponding to the target HPV-loop based on the preset newborn determination algorithm specifically includes: The HPV-loop remodeling index, 3D-haplotype consistency score, and target gene loop coupling expression index were calculated separately, and at least one of the HPV-loop remodeling index, the 3D-haplotype consistency score, and the target gene loop coupling expression index was used as the newborn score.
6. The method according to claim 5, characterized in that, The calculation of the HPV-loop remodeling index specifically includes: A dual-control method is introduced to calculate the denovo score, and the denovo scores are summarized to calculate the HPV-loop remodeling index; the dual-control method is the reference loop library control method and the matched random control method.
7. The method according to claim 5, characterized in that, The calculation of the 3D-haplotype consistency score specifically includes: Based on the heterozygous SNP phase separation, the reads supporting the target HPV-loop are assigned to haplotypes, and loop equipotential bias is calculated; The loop allelic bias and allelic-specific expression from RNA-seq are jointly modeled, and a 3D-haplotype regulatory event set is output. The 3D-haplotype consistency score is calculated based on the 3D-haplotype control event set.
8. The method according to claim 5, characterized in that, The calculation of the target gene loop coupling expression index specifically includes: Motif enrichment analysis was performed on the HPV end anchor and host end anchor of the denavoHPV-loop, and a set of transcription factors that met the preset conditions was screened. Based on the set of transcription factors, a protein-protein interaction network is constructed to obtain a co-transcription factor module, and a target gene loop coupling expression index is output based on the transcription factor module; the target gene includes at least one of MYC, CASC19, and FAM81A; the co-transcription factor module includes at least one of FOS, CEBPG, SP1, or CTCF.
9. The method according to claim 1, characterized in that, The step of outputting chromatin loop identification results based on the neonatal score specifically includes: In vitro risk stratification analysis is performed based on the chromatin circuit identification results; the in vitro risk stratification analysis is used to quantify and stratify the carcinogenic potential of HPV-positive samples or cell models and provide reference information. Prognostic correlation analysis is performed based on the chromatin loop identification results; the prognostic correlation analysis is used to analyze HPV-loop coupling activation events corresponding to different target genes and to use the HPV-loop coupling activation events as candidate prognostic indicators. An intervention screening operation is performed based on the chromatin loop identification results; the intervention screening operation is used to screen candidate intervention methods to inhibit or stabilize HPV-loop formation. At least one output result corresponding to the in vitro risk stratification analysis, the prognostic correlation analysis, and the intervention screening operation is used as the chromatin circuit identification result.
10. A device for identifying HPV-integrated neochromatin loops, characterized in that, The device includes an acquisition module and a processing module, wherein, The acquisition module is used to acquire interaction sequencing read data and determine the HPV reference sequence and host reference sequence based on the interaction sequencing read data; the interaction sequencing read data includes sequencing read data, Hi-C data, ChIA-PET data, HiChIP data, and PLAC-seq data; HPV integration breakpoints are identified based on anomaly alignment features; the anomaly alignment features include chimeric read features, split read features, and paired end anomaly features; based on the breakpoint direction of the HPV integration breakpoint, the HPV reference sequence and the host reference sequence are identified; A fusion contig is constructed between reference sequences to obtain a breakpoint-aware extended reference set. The interacting sequencing reads are aligned with the breakpoint-aware extended reference set. Based on the alignment results and breakpoint-aware alignment constraints, anchor point correction and filtering operations are performed on the target reads to obtain high-confidence cross-genome anchor point pairs. The target reads are reads falling within a preset breakpoint-aware region, which includes a fusion region and a breakpoint neighborhood window. The breakpoint-aware alignment constraints include breakpoint neighborhood consistency constraints, directional consistency constraints, and unique alignment constraints. The processing module is used to cluster the high-confidence cross-genomic anchor pairs to obtain a candidate loop set, and select at least one target HPV-loop from the candidate loop set based on a preset screening rule; calculate the newborn score corresponding to the target HPV-loop based on a preset newborn determination algorithm, and output the chromatin loop identification result according to the newborn score.