A method and device for super edge evidence analysis of multi-site chromatin interaction
By decomposing and normalizing multi-site interaction data, generating hyperedge signatures and performing incremental mining, a saliency discrimination model is constructed, which solves the problem of inaccurate saliency determination in multi-site interaction analysis and achieves efficient identification and verification of multi-site interaction structures.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAZHONG AGRI UNIV
- Filing Date
- 2026-02-26
- Publication Date
- 2026-06-26
AI Technical Summary
Existing multi-site interaction analysis systems suffer from inaccurate significance determination results when evaluating multi-site interactions. They are unable to uniformly model the co-occurrence relationships of multiple anchor points within the same analysis unit, leading to information loss and significance bias.
The target read is decomposed into multiple genomic fragments and mapped to a set of anchor points. Normalization is performed to generate a superedge signature, and evidence pointers are saved to form a superedge evidence transaction. Incremental candidate mining is carried out through biological and statistical constraints, consistency statistics are calculated, a significance discrimination null model is constructed, and the corrected significance probability value is output through multiple tests.
It improves the accuracy and interpretability of multi-site interaction analysis, ensures the reproducibility and consistency of analysis results, improves storage and retrieval efficiency through coding, and ensures the consistency of identification and processing of the same multi-site interaction structure in different reads.
Smart Images

Figure CN122290712A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of high-order graph data processing, and in particular to a method and apparatus for analyzing hyperedge evidence of multi-site chromatin interactions. Background Technology
[0002] Multi-site interaction data, such as Pore-C, are generated from multiple connected reads, each naturally corresponding to multiple anchor points and forming higher-order hyperedge evidence. Compared to traditional pairwise interactions, this data exhibits differences in coverage, distance distribution bias, duplicate / near-duplicate evidence, and the need for verifiability that requires backtracking to reads. Projecting multi-site interactions as pairwise interactions can easily lead to information loss and significance bias. Therefore, how to verify the significance of multi-site interactions while ensuring data integrity and accuracy has become an important research direction.
[0003] Existing multi-site interaction analysis systems typically validate multi-site interaction data using simple statistical significance methods. For example, they calculate the support and anchor hit rate for each interaction and determine whether the interaction is significant based on a fixed threshold. However, these systems only examine the support of each read segment for a single anchor or paired anchors during the statistical process, failing to uniformly model the co-occurrence relationships of multiple anchors within the same analysis unit. Therefore, existing multi-site interaction analysis systems struggle to accurately reflect the true high-order interaction structure when assessing information loss and significance bias introduced by read segment splitting or pairing during multi-site interaction evaluation, easily leading to inaccurate determinations of the significance of multi-site interactions.
[0004] Therefore, there is an urgent need for a method and apparatus for analyzing hyperboundary evidence of multi-site chromatin interactions. Summary of the Invention
[0005] This application provides a method and apparatus for super-edge evidence analysis of multi-site chromatin interactions, which solves the problem that existing multi-site interaction analysis systems have inaccurate significance determination results when evaluating multi-site interactions.
[0006] The first aspect of this application provides a method for analyzing superedge evidence of multi-site chromatin interactions. The method includes: decomposing a target read into multiple genomic fragments and mapping them to an anchor set; the target read being any read in multi-connected read interaction data; normalizing the anchor set matched by the target read to obtain a superedge signature, and storing at least one evidence pointer for each superedge signature to form a superedge evidence transaction; incrementally mining candidates for the superedge evidence transaction under preset biological and statistical constraints, and obtaining a candidate superedge set using preset pruning rules; calculating the target candidate... The consistency statistic between the selected superedge and the reading evidence is used; the target candidate superedge is any one of the candidate superedges in the candidate superedge set; the reading evidence is obtained by backtracking through the evidence pointer associated with the target candidate superedge; a significance null model is constructed based on the consistency statistic, and the significance probability value corresponding to the target candidate superedge is output through the significance null model; multiple tests are performed on the significance probability value of each target candidate superedge, and the corrected significance probability value is output; the superedge evidence analysis results and verifiable data packets are output based on the corrected significance probability values; the verifiable data packets are used to backtrack the evidence pointer.
[0007] Optionally, the set of anchor points hit by the target read segment is normalized to obtain the superedge signature. Specifically, this includes: performing deduplication and sorting processing on the set of anchor points hit by the target read segment to obtain the superedge signature; and further representing the superedge signature as an encoding form for counting or querying, wherein the encoding form includes bitmap, hash value or compressed integer sequence.
[0008] Optionally, after normalizing the set of anchor points that the target read segment hits to obtain a super-edge signature, and saving at least one evidence pointer for each super-edge signature to form a super-edge evidence transaction, the method further includes: constructing a set of read segment evidence pointers for each target read segment; the evidence pointer includes a read segment identifier, original file offset information, and a parsed index key; and constructing an inverted mapping from anchor points to the set of read segment evidence pointers as the anchor point set.
[0009] Optionally, the consistency statistic includes at least observational support and evidence consistency coefficient. Calculating the consistency statistic between the target candidate superedge and the read segment evidence specifically includes: calculating observational support by counting the number of read segment evidence pieces that satisfy a first preset consistency condition; the first preset consistency condition is that the read segment evidence can cover the set of anchor points corresponding to the target candidate superedge; calculating the evidence consistency coefficient by counting the degree of overlap between read segment evidence pieces that satisfy a second preset consistency condition; the second preset consistency condition is that the read segment evidence can appear simultaneously on each anchor point corresponding to the target candidate superedge; and calculating the consistency statistic between the target candidate superedge and the read segment evidence based on the observational support and the evidence consistency coefficient.
[0010] Optionally, a significance-discriminating null model is constructed based on consistency statistics, and the significance probability value corresponding to the target candidate hyperedge is output through the significance-discriminating null model. Specifically, this includes: constructing a significance-discriminating null model based on consistency statistics that maintains at least one structural layering; the layering includes at least two of the following: chromosome layering, hyperedge order layering, span or distance layering, and layering based on anchor point heat; calculating the significance probability value based on the significance-discriminating null model through a first random control method or a second random control method; the first random control method involves translating and permuting the anchor point identifier corresponding to the target candidate hyperedge within the contig or chromosome; the second random control method involves performing degree-preserving rearrangement sampling on the anchor point-read segment association relationship while maintaining the anchor point chromosome count unchanged.
[0011] Optionally, after performing multiple tests on the significance probability value of each target candidate hyperedge and outputting the corrected significance probability value, the method further includes: performing a difference or reconnection statistical test on the target candidate hyperedge between at least two sets of conditions or groups; the at least two sets of conditions or groups include at least one of different samples, different treatment conditions, different cell clusters, and different haplotypes or alleles; the difference or reconnection statistical test includes at least one or more of Fisher's exact test, chi-square test, negative binomial or generalized linear model test, β-binomial overdispersion correction test, and hierarchical label permutation test.
[0012] Optionally, a verifiable data package is output, specifically including: a verifiable data package based on the corrected significance probability value, wherein the verifiable data package includes an input data summary, parameter and software version information, a list of result files, checksum or hash check information, and a minimal backtracking script or interface description for extraction or recalculation.
[0013] Optionally, after outputting the hyperedge evidence analysis results and verifiable data packets based on the corrected significance probability values, the method further includes: analyzing the conditions of the structural variation breakpoint neighborhood based on the hyperedge evidence analysis results and verifiable data packets, specifically including: obtaining a set of structural variation breakpoints based on the hyperedge evidence analysis results and verifiable data packets, and defining a breakpoint window of a preset length for each structural variation breakpoint; determining whether the genomic region corresponding to any anchor point of a valid candidate hyperedge overlaps with the breakpoint window; valid candidate hyperedges are target candidate hyperedges that pass the significance discrimination; if the genomic region corresponding to any anchor point of a valid candidate hyperedge overlaps with the breakpoint window, then the valid candidate hyperedge is confirmed to have hit the breakpoint window; calculating the event-level statistics of valid candidate hyperedges that hit the breakpoint window, and calculating the event-level difference based on the event-level statistics; the event-level statistics include the number of hit candidate hyperedges, the total support of hit candidate hyperedges, and the average order of hit candidate hyperedges; outputting the result table corresponding to the conditions of the structural variation breakpoint neighborhood based on the event-level difference, and grouping the result table according to haplotype or allele.
[0014] A second aspect of this application provides a hyperborder evidence analysis device for multi-site chromatin interactions. The device includes a data preprocessing module, a candidate mining module, a significance evaluation module, a testing module, and an output module, wherein... The data preprocessing module is used to decompose the target read into multiple genomic fragments and map them to a set of anchor points; the target read is any read data in the multi-connected read interaction data; the set of anchor points hit by the target read is normalized to obtain a superedge signature, and at least one evidence pointer is saved for each superedge signature to form a superedge evidence transaction.
[0015] The candidate mining module is used to perform incremental candidate mining of superedge evidence transactions under preset biological and statistical constraints, and to obtain a set of candidate superedges using preset pruning rules.
[0016] The saliency assessment module is used to calculate the consistency statistic between the target candidate hyperedge and the read segment evidence. The target candidate hyperedge is any one of the candidate hyperedges in the candidate hyperedge set. The read segment evidence is obtained by backtracking through the evidence pointer associated with the target candidate hyperedge. A saliency discrimination null model is constructed based on the consistency statistic, and the saliency probability value corresponding to the target candidate hyperedge is output through the saliency discrimination null model.
[0017] The testing module performs multiple tests on the significance probability value of each target candidate hyperedge and outputs the corrected significance probability value.
[0018] The output module is used to output the hyperedge evidence analysis results and verifiable data packets based on the corrected significance probability values; the verifiable data packets are used to backtrack the evidence pointers.
[0019] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described above.
[0020] A fourth aspect of this application provides a non-transitory computer-readable storage medium storing a computer program, the computer program being executed by a processor using any of the methods described above.
[0021] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. Decompose and map the target read segment in the multi-connected read segment interaction data to the anchor point set. Normalize the hit anchor points to generate hyperedge signatures and save the evidence pointers to form hyperedge evidence transactions. Incrementally mine the hyperedge evidence transactions under biological and statistical constraints to obtain a candidate hyperedge set. Backtrack the read segment evidence based on the evidence pointers and calculate the consistency statistics of the target candidate hyperedges. Construct a significance discrimination null model based on the consistency statistics, output the significance probability value, and obtain the corrected significance probability value through multiple tests. Output the hyperedge evidence analysis results and verifiable data packets for backtracking the evidence pointers based on the corrected significance probability value, thereby realizing the closed-loop association between the results and evidence and improving the interpretability, reproducibility, and reproducibility consistency of the analysis results.
[0022] 2. Perform deduplication and sorting on the set of anchor points hit in the target read segment to obtain the hyperedge signature; further represent the hyperedge signature as an encoding form for counting or querying, including bitmap, hash value or compressed integer sequence, thereby improving the storage efficiency and retrieval efficiency of the hyperedge signature in the candidate mining, support statistics and consistency analysis process while maintaining the semantics of the hyperedge structure, and ensuring that the same multi-site interaction structure in different read segments can be consistently identified and merged.
[0023] 3. Construct a significance null model that maintains at least one structural hierarchy based on consistency statistics; Based on the significance null model, calculate the significance probability value by translating and permuting the anchor point identifiers corresponding to the target candidate hyperedge within the contig or chromosome, or by performing degree-preserving rearrangement sampling on the anchor point-read segment association relationship while keeping the anchor point chromosome count unchanged. This allows the estimation of the occurrence probability of the consistency statistics of the target candidate hyperedge in a random background that satisfies structural constraints, and the corresponding significance probability value is obtained. Attached Figure Description
[0024] Figure 1 This is a flowchart illustrating a method for analyzing hyperedge evidence of multi-site chromatin interactions provided in an embodiment of this application. Figure 2 This is a schematic diagram of a module of a hyperedge evidence analysis device for multi-site chromatin interaction provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0025] Explanation of reference numerals in the attached figures: 21. Data preprocessing module; 22. Candidate mining module; 23. Significance evaluation module; 24. Validation module; 25. Output module; 301. Processor; 302. Communication bus; 303. User interface; 304. Network interface; 305. Memory. Detailed Implementation
[0026] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0027] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.
[0028] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0029] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0030] Please refer to Figure 1 The flowchart illustrates a method for analyzing super-edge evidence of multi-site chromatin interaction provided in this application embodiment. The flowchart mainly includes the following steps: S101 to S107.
[0031] Step S101: Decompose the target read into multiple genomic fragments and map them to an anchor set.
[0032] Specifically, in multi-site chromatin interaction sequencing experiments, a large number of interaction records consisting of multiple linked reads are generated. Each read often covers multiple scattered genomic locations at the same time, and the frequency and stability of different location combinations appearing in the data vary significantly. Among them, the target read is any read data in the multi-linked read interaction data.
[0033] Existing multi-site chromatin interaction analysis systems typically focus on candidate structure generation or saliency assessment separately, emphasizing algorithm implementation for each single processing stage. They lack a comprehensive system capable of collaboratively addressing multi-stage constraints within a unified technical framework. Therefore, they struggle to simultaneously meet the following technical requirements: First, while ensuring traceability of evidence, effectively constraining the high-dimensional combinatorial space induced by multi-connection reads to suppress combinatorial explosion and generate a controllable set of structurally consistent candidate hyperedges; Second, without compromising key structural hierarchical features such as chromosome affiliation, hyperedge order, and anchor distance distribution. Under these conditions, a permutation or resampling null model is constructed to robustly calibrate the significance of the consistency statistics of the hyperedge. Third, under different sample sources, different treatment conditions, different cell clusters, and different haplotypes or allelic backgrounds, and even within the neighborhood of structural variation breakpoints, differential or reconnection analyses are uniformly performed on the same hyperedge structure, and event-level effect sizes with clear statistical significance are output. Fourth, at the result output stage, an evidence package that is recalculated, verifiable, and can be accurately traced back to the original read evidence is provided, so that the analysis conclusions are not only statistically significant, but also have complete data sources and reproduction paths.
[0034] This application addresses this type of data and aims to systematically organize the multi-site co-occurrence relationships reflected in the reads. Through statistical and constraint analysis, it distinguishes between multi-site interaction relationships driven by biological spatial structure that occur stably and repeatedly in the data, and co-occurrence combinations that occur only occasionally in a small number of reads and lack stable support. Furthermore, it can clarify the original read support source corresponding to each type of multi-site interaction relationship that is determined to be valid.
[0035] Furthermore, this application not only focuses on whether multi-site interaction relationships objectively exist, but also on how these interactions change under different acquisition conditions, different data sources, or different genetic backgrounds. Specifically, it aims to determine whether existing multi-site interactions are enhanced, weakened, or reconnected, thereby characterizing the reconstruction behavior of multi-site spatial structures under different conditions. Therefore, this application's technical objectives cover both the determination of the validity of evidence for multi-site interactions and the analysis of the changing characteristics of multi-site interaction relationships.
[0036] Based on the aforementioned technical objectives, this application transforms the original multi-connected read evidence into hyperedge evidence transactions carrying evidence pointers. Under preset biological constraints and pruning rules, incremental candidate mining is performed on possible multi-site combinations to control the candidate size and reduce combination complexity. Using read evidence consistency as the core statistical basis, and combining a permutation null model or a matching sampling null model that maintains structural features at the contig or chromosome level, the significance of candidate hyperedges is evaluated, and the corrected significance probability value is obtained through multiple tests. On this basis, difference or reconnection statistical tests are performed on the candidate hyperedges after consistency analysis between different multi-connected read interaction data and at the haplotype or allelic level, and null model settings such as hierarchical label permutation and connection number-preserving random rearrangement are supported. Finally, a verifiable evidence package containing a parameter list, verification information, and traceable read evidence pointers is output, realizing the correspondence between the analysis results and the original evidence. This analytical framework can also be extended to the neighborhood of structural variation breakpoints. By statistically comparing the support and order of hyperedges within the breakpoint window, it quantifies the conditional or isotropic reconnection effect and outputs the corresponding event-level change indicators and their representative evidence.
[0037] Based on the above overall technical solution, in step S101, this application first processes a single multi-linked read by splitting the read into multiple locatable genomic fragments and mapping each fragment to a predefined set of anchor points. Specifically, when normalizing the set of anchor points matched by the read, a symbolic representation is used to formally describe the processing result. The hyperedge signature can be represented as:
[0038] in, This represents the super-edge signature generated for the currently processed single read segment. This represents the set of all anchor identifiers that the read segment has matched after the genome alignment and mapping to the anchor set. This indicates that deduplication is performed on the set of anchor point identifiers to eliminate duplicate anchor points introduced by segment overlap or repeated comparison within the same read segment; This means that the deduplicated anchor point identifiers are sorted according to a preset global order rule, resulting in anchor point combinations that are unique in representation and independent of the anchor point order. Through the above deduplication and sorting operations, as long as different read segments hit the same set of anchor points, a consistent hyperedge signature can be generated to represent the same multi-site chromatin interaction structure.
[0039] While generating a superedge signature S(read), an evidence pointer associated with that read segment is stored in the corresponding superedge signature record to establish a one-to-one or many-to-one correspondence between the superedge signature and the original read segment. The evidence pointer is used in subsequent analysis stages to back-locate the original read segment supporting the multi-site interaction structure based on the superedge signature. Its specific implementation can include the unique identifier of the read segment in the original data, its position offset information in the data file, or a resolvable index key. By retaining the evidence pointer at the superedge signature level, any superedge analysis result can be traced back to the corresponding original read segment evidence set, thereby ensuring the verifiability of the analysis results.
[0040] For multiple read segments with the same S(read) value, it is assumed that they reflect the same type of multi-site interaction relationship at the anchor point level. During the processing, the superedge signature records corresponding to these read segments are merged and their support is accumulated to form a statistical support strength representation of the superedge signature.
[0041] Step S102: Normalize the set of anchor points hit by the target read segment to obtain the super-edge signature, and save at least one evidence pointer for each super-edge signature to form a super-edge evidence transaction.
[0042] Specifically, the set of anchor points hit by the target read segment is deduplicated and sorted to generate a superedge signature that can uniquely represent the multi-point interaction structure of the read segment. At least one evidence pointer for tracing back to the original read segment is associated and stored in the superedge signature, thereby standardizing the interaction information at the read segment level into a traceable superedge evidence transaction, providing a unified input for subsequent statistical analysis.
[0043] In one possible implementation, step S102 further includes: performing deduplication and sorting processing on the set of anchor points hit by the target read segment to obtain a superedge signature; and further representing the superedge signature as an encoding form for counting or querying, the encoding form including bitmap, hash value or compressed integer sequence.
[0044] Specifically, after completing the deduplication and sorting of the anchor point set matched by the target read segment and obtaining the superedge signature, in order to improve the efficiency of subsequent statistical analysis, candidate mining, and consistency calculation, this embodiment further expresses the superedge signature using an encoding representation suitable for computer processing. The purpose of setting the encoding representation is to enable the superedge signature to participate in counting, comparison, and retrieval operations in a more compact or efficient manner without changing the anchor point combination semantics expressed by the superedge signature.
[0045] In one possible implementation, step S102 further includes: constructing a set of read segment evidence pointers for each target read segment; the evidence pointer includes a read segment identifier, original file offset information, and a parsed index key; and constructing an inverted mapping from the anchor point to the set of read segment evidence pointers as an anchor point set.
[0046] Specifically, in one implementation, the hyperedge signature can be represented as a bitmap. The bitmap format uses a predefined set of anchor points as the index space, assigning a fixed position to each anchor point. When a hyperedge signature contains an anchor point, it is marked as valid at the corresponding position; otherwise, it is marked as invalid. Using the bitmap format, the inclusion or intersection / merger relationships between different hyperedge signatures at the anchor point level can be quickly determined, making it suitable for scenarios requiring frequent set operations or consistency statistics.
[0047] In another implementation, the superedge signature can be represented as a hash value. The hash value form involves taking the deduplicated and sorted sequence of anchor identifiers as input and generating a fixed-length digest identifier using a deterministic hash function. This digest identifier is used for fast comparison and indexing of superedge signatures without explicitly storing the complete anchor sequence, making it suitable for scenarios involving large-scale candidate superedge deduplication, counting, or rapid location. Because the superedge signature has already undergone normalization during generation, identical anchor combinations will stably map to the same hash result, thus ensuring semantic consistency.
[0048] In another implementation, the superedge signature can be represented as a compressed integer sequence. A compressed integer sequence refers to compactly representing the sorted anchor identifiers according to predetermined integer encoding rules, such as differential encoding, variable-length encoding, or other compression strategies, to reduce storage space usage. This representation effectively reduces memory and storage overhead while maintaining anchor sequence information, making it suitable for scenarios that require both preserving the anchor sequence structure and performing large-scale statistical analysis.
[0049] Through any of the above encoding representations, the hyperedge signature logically still uniquely corresponds to a definite set of anchor points, without changing its technical meaning as an identifier of a multi-site interaction structure. Different encoding representations only affect the storage method and computational efficiency of the hyperedge signature at the implementation level, without affecting the implementation logic of subsequent candidate hyperedge mining, read segment evidence consistency statistics, saliency discrimination, and difference or reconnection analysis based on the hyperedge signature. By introducing the above encoding representations in step S102, the hyperedge evidence transaction maintains both semantic clarity and traceability, and possesses computational performance suitable for large-scale data processing.
[0050] Step S103: Under preset biological and statistical constraints, incremental candidate mining is performed on the superedge evidence transactions, and a set of candidate superedges is obtained using preset pruning rules.
[0051] Specifically, based on the obtained superedge evidence transactions, and under the combined constraints of pre-set biological and statistical constraints, incremental candidate mining is performed on the superedge evidence transactions to generate a candidate superedge set and avoid runaway candidate size. A superedge evidence transaction is a transactional record centered on a superedge signature and associated with at least one evidence pointer. The superedge signature corresponds to the anchor point combination hit by the read segment, and the evidence pointer is used to backtrack to the original read segment supporting that anchor point combination. Since a single multi-connected read segment may hit multiple anchor points, if all anchor point combinations are enumerated without constraints, the number of candidates will expand rapidly with the number of anchor points. Therefore, in this step, constraint and pruning mechanisms are used to control the mining space.
[0052] Statistical constraints are used to define the minimum reliability requirements that candidate hyperedges should meet at the data level. Statistical constraints must include at least a lower bound on the order of the hyperedge. And minimum support threshold. The hyperedge order is used to represent the number of different anchor points in the hyperedge signature. When the order is low, it easily degenerates into pairwise interactions or low-order co-occurrence, making it difficult to reflect the structural features of multi-site interactions. Therefore, by setting... This ensures that candidate hyperedges only enter the candidate set when their order reaches a preset lower limit, thus focusing the analysis on multi-site interaction structures. A minimum support threshold is used to limit the number of reads that must support a candidate hyperedge. Support can be characterized by the number of reads that match the hyperedge signature or cover the anchor combination. This threshold filters out anchor combinations that are only incidentally supported by a very small number of reads, reducing the proportion of noisy candidates.
[0053] Biological constraints are used to ensure that the generation of candidate hyperedges conforms to genomic structure and experimental characteristics, thereby avoiding the generation of candidates that are biologically unreasonable or have low interpretative value. Biological constraints can include span or distance constraints and annotation combination constraints. Span or distance constraints limit the genomic location span or relative distance between anchor points in candidate hyperedges to a predetermined range, preventing the indiscriminate inclusion of anchor point combinations that are too far apart and lack reasonable biological explanation, while also reducing statistical bias caused by differences in distance distribution. Annotation combination constraints filter candidates based on the functional annotations or regional attributes associated with the anchor points. For example, they limit the annotation type combinations of anchor points in candidate hyperedges to meet predetermined rules, thereby making candidate hyperedges more relevant to the target biological question and reducing the number of irrelevant combinations.
[0054] In terms of candidate generation strategy, an incremental candidate mining approach is adopted to gradually expand anchor point combinations, rather than enumerating all combinations at once. The basic idea of incremental candidate mining is to start with lower-order anchor point combinations and decide whether to continue expanding to higher-order combinations based on the support and constraint satisfaction of existing combinations, thereby continuously filtering out combinations that do not meet the conditions during the expansion process. This strategy can be implemented using the mining idea of frequent itemsets, where the hyperedge signature corresponding to each read segment is regarded as a transaction, and the anchor point identifier is regarded as an element in the transaction. Candidate hyperedges are obtained by mining anchor point combinations that occur repeatedly in a large number of transactions. In this way, the generation of the candidate hyperedge set is directly related to the true co-occurrence pattern in the hyperedge evidence transaction, which can prioritize the discovery of anchor point combinations that occur repeatedly in the read segment evidence.
[0055] In terms of pruning rules, invalid expansions are avoided by terminating the candidate expansion process early. The core of the pruning rule is to utilize the monotonicity or upper bound estimation of support. When the upper bound of the support of a certain set of anchor points is already lower than the minimum support threshold, any higher-order expansion that includes this set cannot meet the minimum support requirement, thus directly terminating further expansion of that branch. Here, the set of anchor points can be understood as intermediate combinations in the candidate expansion process, and the upper bound of support can be obtained based on the currently observed number of support reads or transaction counts. Through this pruning mechanism, a large number of combinations that cannot meet the threshold can be eliminated before the candidate expands to a higher order, suppressing the exponential growth of the number of candidates and controlling the consumption of computational resources.
[0056] Step S104: Calculate the consistency statistic between the target candidate superedge and the read segment evidence; the target candidate superedge is any candidate superedge in the candidate superedge set.
[0057] Specifically, by statistically analyzing the number of read segments supporting the target candidate superedge, we obtain the observational support, reflecting the strength of the superedge's presence in the data. Simultaneously, based on the correspondence between anchor points and read segment evidence, we assess the overlap of read segment evidence corresponding to different anchor points, obtaining the evidence consistency coefficient to characterize the degree of evidence consistency. By combining these two types of statistical results, we form the consistency statistic for the target candidate superedge, reflecting whether the candidate superedge is stably and collaboratively supported by multiple read segments, thus providing a basis for subsequent significance determination.
[0058] In one possible implementation, step S104 further includes: calculating the observation support by counting the number of read segment evidences that satisfy the first preset consistency condition; the first preset consistency condition is that the read segment evidence can cover the set of anchor points corresponding to the target candidate hyperedge; calculating the evidence consistency coefficient by counting the degree of overlap between read segment evidences that satisfy the second preset consistency condition; the second preset consistency condition is that the read segment evidence can appear simultaneously on each anchor point corresponding to the target candidate hyperedge; and calculating the consistency statistic between the target candidate hyperedge and the read segment evidence based on the observation support and the evidence consistency coefficient.
[0059] Specifically, this embodiment quantifies the existence of a stable and cooperative support relationship between the target candidate hyperedge and the read evidence. The observed support reflects the sufficiency of the number of read evidences satisfying the coverage relationship, while the evidence consistency coefficient reflects the high overlap of read evidence corresponding to each anchor point, thus avoiding situations where there are many reads at individual anchor points but a lack of common evidence overall. The target candidate hyperedge is any one of the candidate hyperedges in the candidate hyperedge set, which can be represented by a set of anchor points. Read evidence can be backtracked and statistically analyzed using evidence pointers.
[0060] Anchor points are used to uniformly express the location results of read segments on the reference genome. Anchor points can be represented as triples containing chromosome identifiers, start coordinates, and end coordinates, or as index identifiers corresponding one-to-one with those triples. Read segment This represents a multi-join read segment or a set of its parsed segments. Indicates reading segment The set of anchor points that were hit. Deduplication is performed during the formation process to avoid duplicate hits of the same anchor point by the same read segment, which would affect subsequent statistics. Candidate superedges The corresponding set of anchor points is denoted as , where k is the order of the candidate superedge. Different anchor point identifiers. This represents the set of read segments participating in the consistency check, and its selection is used to ensure that the read segments have the basic conditions to support the target order hyperedge. The number of segments contained in the parsed segment is not less than k, and the mapping quality of the segment meets the preset filtering conditions, thereby avoiding interference from low-quality or insufficient segments on consistency statistics.
[0061] The observation support corresponds to the first pre-set consistency condition, which requires that the evidence for the read segment must cover the entire set of anchor points corresponding to the target candidate hyperedge, that is, the read segment hit set tokens(r) contains all anchor points in A(e). The formula for calculating the observation support is:
[0062]
[0063] in, Indicates candidate superedges The observation support is the number of reads that satisfy the coverage relationship. The larger the value, the more reads can hit all the anchors of the candidate superedge at the same time, provided that the read quality and number of segments are satisfied. The set of read segments participating in the statistics is determined by the threshold for the number of read segments and the mapping quality filter. Let k be the set of anchor points corresponding to the candidate superedges, and its size is determined by the order k. For reading paragraph The set of anchor points hit. (Symbol "") "Used to indicate a set containment relationship, here it means Each anchor point in the text is contained within In the middle. This formula is derived from the fact that... The reading segments are checked one by one for coverage relationships and counted, so that the observation support directly describes the amount of evidence that "simultaneously covers all anchor points", thus corresponding to the first preset consistency condition.
[0064] The evidence consistency coefficient corresponds to the second pre-set consistency condition. This condition requires that the evidence for the read segment must appear simultaneously at each anchor point corresponding to the target candidate hyperedge. Essentially, it utilizes the inverted mapping from anchor points to the read segment evidence to check the degree of overlap between the evidence set that "hit all anchor points" and the evidence set that "hit any anchor point." Evidence pointer Used to indicate the location or access path of the original read segment in super-edge evidence transactions. It can include segment identifiers. At least one of the following is used: original file offset or resolvable index key, to ensure that the original read segment can be traced back when outputting verifiable data packets later. An inverted index mapping R(a) from anchor points to the read segment evidence pointer set is constructed based on the evidence pointers, so that any anchor point a can quickly locate the read segment evidence set that hits that anchor point. Its definition is:
[0065]
[0066] in, This represents the set of evidence pointers corresponding to anchor point a. For any anchor point identifier, For reading paragraph The corresponding evidence pointers. This inverted index map aggregates all hit anchor points. The read segment evidence pointer allows subsequent intersection and union operations on multiple anchor point combinations to be completed at the evidence pointer level, thereby simultaneously satisfying the requirements for consistent statistics and evidence traceability.
[0067] For candidate hyperedges Its anchor point set is Define the set of evidence that matches all anchor points. and the union of evidence that hits any anchor point They are respectively:
[0068]
[0069] in, This indicates that it can simultaneously hit candidate superedges. The set of read segment evidence pointers for all anchor points, whose values are determined by the... Each anchor point corresponds to The intersection of the sets is obtained by taking the intersection of the sets. The more elements in the intersection, the more read segments there are that simultaneously support all anchor points in the same read segment evidence. Indicates that the candidate superedge is hit. The set of read segment evidence pointers for any anchor point in the middle, whose values are determined by the... Each anchor point corresponds to The union of these sets is used to determine the overall evidence coverage related to the candidate hyperedge. Based on the intersection and union of these sets, the coefficient of evidence consistency is calculated. The calculation formula is:
[0070]
[0071] in, The coefficient of consistency of evidence is represented by the numerator, which ranges from 0 to 1. The denominator represents the number of evidence pointers that simultaneously support all anchor points. This represents the number of evidence points supporting any given anchor. This coefficient measures consistency by the proportion of all relevant evidence that genuinely supports all anchors simultaneously. A larger value indicates that the supporting reads from different anchor points are more concentrated on the same batch of read evidence, thus better meeting the second presupposed consistency condition. Because The elements in the data are all evidence pointers. These evidence pointers can be directly used to backtrack the original set of read segments that support the candidate superedge in the verifiable data packet, thereby ensuring that the consistency statistics are verifiable.
[0072] Consistency statistics are a comprehensive quantitative result based on observational support and the coefficient of evidence consistency. Observational support emphasizes the threshold of the amount of evidence, while the coefficient of evidence consistency emphasizes the quality of evidence overlap. Combining the two can avoid misjudgments caused by relying on a single indicator. For example, when Larger but When the value is low, it indicates that although many read segments support some anchor points individually, the proportion of read segments that collectively support all anchor points is low. Such candidate superedges are more likely to reflect unstable combinations. Smaller but A high consistency rate indicates a high concentration of common evidence but insufficient evidence size. Subsequent significance assessments can further utilize a null model to determine its reliability. Based on the aforementioned consistency statistics, to further quantify the reliability of evidence for target candidate hyperedges, this embodiment can also introduce a confidence calculation method based on empirical Bayesian principles. This method establishes a comparative relationship between observed evidence and background evidence. The posterior mean of the consistency ratio can be expressed as... The relevant algorithm is as follows:
[0073]
[0074]
[0075]
[0076]
[0077] in, By applying the observed count x together with the prior parameters a and b to the numerator and denominator, the counting of small samples is smoothed, avoiding extreme probabilities when x is 0 or x is n. By comparing the differences between the foreground and background in the logarithmic dominance space, the characterization of the intensity of the difference is enhanced; By using the sigmoid function, the intensity of the difference is mapped to a confidence level between 0 and 1, making it easier to use in conjunction with other statistics. Multiplying the confidence level by the evidence consistency coefficient increases the overall score when the candidate hyperedge shows a higher support advantage under the foreground condition relative to the background condition, and the corresponding reading evidence shows a higher degree of concentration and overlap among the anchor points. This ensures that the score is consistent with the joint evaluation logic of evidence consistency and support advantage in this step.
[0078] Step S105: Construct a significance null model based on the consistency statistic, and output the significance probability value corresponding to the target candidate hyperedge through the significance null model.
[0079] Specifically, based on the consistency statistics of the target candidate hyperedge, a null model for significance discrimination is constructed, and a random control distribution is generated while keeping the preset structural constraints unchanged. By comparing the consistency statistics of the target candidate hyperedge in the real data with the random control distribution, the significance probability value corresponding to the target candidate hyperedge is calculated, which is used to quantify the probability of the consistency level occurring under random conditions, thereby distinguishing between multi-site interaction relationships supported by stable read evidence and combinations caused by random co-occurrence.
[0080] In one possible implementation, step S105 further includes: constructing a significance-discriminating null model that maintains at least one structural stratification based on consistency statistics; the stratification includes at least two of chromosome stratification, hyperedge order stratification, span or distance stratification, and anchor point heat stratification; calculating the significance probability value based on the significance-discriminating null model using a first random control method or a second random control method; the first random control method involves translating and permuting the anchor point identifier corresponding to the target candidate hyperedge within the contig or chromosome; the second random control method involves performing degree-preserving rearrangement sampling on the anchor point-read segment association relationship while maintaining the anchor point chromosome count.
[0081] Specifically, firstly, the anchor-read segment association is a data structure used to characterize the hit correspondence between anchor points and read segment evidence. It is established by associating each anchor point with read segment evidence that hits that anchor point. The associated object corresponding to the anchor point can be expressed using a set of evidence pointers, which can consist of at least one of read segment identifiers, original file offsets, or resolvable index keys. This allows any anchor point to be traced back to the set of read segment evidence that hits that anchor point through the anchor-read segment association, and supports performing intersection or union operations on the evidence sets of multiple anchor points at the candidate hyperedge level to form the evidence set required for consistency statistics. A significance-discriminating null model is constructed based on consistency statistics, maintaining at least one structural layering. Structural layering is used to constrain randomization operations, ensuring that the random control distribution remains comparable to real data in terms of key structural attributes. Structural layering includes at least two of the following: chromosome layering, hyperedge order layering, span or distance layering, and layering based on anchor point heat. Chromosome stratification is used to ensure that anchor points always originate from the same chromosome or contig during randomization, preventing cross-chromosomal permutations from disrupting the spatial background structure; hyperedge order stratification is used to ensure that the hyperedges of randomized controls and target candidates are of similar order. Maintaining consistency at the hierarchical level ensures that the comparison of consistency statistics is unaffected by differences in the number of anchor points; span or distance stratification constrains the distribution range of randomly generated anchor point combinations on the genomic coordinate scale, avoiding systematic shifts introduced by differences in genomic distance; stratification based on anchor point popularity maintains the overall usage ratio of high-frequency and low-frequency anchor points in randomized controls, reducing statistical bias caused by uneven anchor point hit frequencies.
[0082] After setting the structural hierarchical constraints as described above, based on the saliency discrimination null model, the saliency probability values corresponding to the target candidate hyperedges are calculated using a random comparison method. The random comparison method includes at least a first random comparison method and a second random comparison method, and in one possible implementation, it can also be combined with matched sampling to jointly construct the background distribution, thereby enhancing the robustness of the null model. In the matched sampling method, the hyperedge order is used to calculate the saliency probability. Read segments are binned to ensure consistency in the number of anchor points within each bin. Further grouping by chromosome ensures comparability between background reads and target candidate hyperedges in terms of chromosome affiliation. Subsequently, the number of background reads is randomly sampled from the set of reads belonging to the same bin as the target candidate hyperedge. And calculate background support based on these background reads. And the corresponding number of trials. .in, This indicates the number of background read segments that satisfy the same coverage or consistency conditions as the target candidate superedge. This represents the total number of reads participating in the background statistics. The background support distribution obtained in this way is used to characterize the randomness of the read support consistency statistic under the same order and chromosome structure conditions. In the first randomized controlled method, the anchor markers corresponding to the target candidate superedges within the contig or chromosome are translated and permuted. Specifically, for each read, an overall offset is generated within its contig or chromosome range, and the anchor markers hit by the read are translated as a whole in the anchor list sorted by genomic location, allowing wrapping around the list boundaries. This translation and permutation operation preserves the relative density structure and read complexity characteristics of the anchors within the chromosome, but disrupts the precise co-occurrence relationship between the original anchor combinations. The above translation and permutation process is repeated. Second-rate, This represents the number of times the translation permutation process is repeated in the significance-discriminating null model, resulting in a set of background consistency statistics. Consistency statistics of the target candidate hyperedge observed in real data. The significance probability of the target candidate hyperedge is estimated by comparing it with the set of background consistency statistics and using the following formula:
[0083]
[0084] The numerator is used to count the number of times the background consistency statistic is not less than the true observation under random control conditions, and the denominator is used to smooth and normalize the number of permutations so that the significance probability value can stably reflect the degree to which the target candidate hyperedge deviates from random co-occurrence while maintaining the chromosome structure.
[0085] In the second randomized controlled trial, anchor markers are resampled to maintain their degree of uniformity while keeping the anchor chromosome count constant. Specifically, firstly, the number of anchor markers matched by each read on each chromosome or contig is counted and used as a constraint in the randomization process. Then, from the complete set of anchor markers on the corresponding chromosome or contig, anchor markers are randomly sampled from the set of anchor markers without replacement, ensuring that the number of anchor markers for each read on each chromosome remains unchanged, but the specific anchor marker positions are randomly rearranged. The background read set generated in this way is consistent with the real data in terms of connectivity distribution and chromosome distribution, but the specific co-occurrence relationships between the original anchor markers are broken, thus obtaining a background consistency statistic distribution used for significance determination.
[0086] Step S106: Perform multiple tests on the significance probability value of each target candidate hyperedge and output the corrected significance probability value.
[0087] Specifically, for all target candidate hyperedges, the p-values, i.e. the significance probability values, are corrected using the BH method with FDR, and the q-values, i.e. the corrected significance probability values, are output.
[0088] In one possible implementation, step S106 further includes: performing a difference or reconnection statistical test on the target candidate hyperedge between at least two sets of conditions or groups; the at least two sets of conditions or groups include at least one of different samples, different treatment conditions, different cell clusters, and different haplotypes or alleles; the difference or reconnection statistical test includes at least one or more of Fisher's exact test, chi-square test, negative binomial or generalized linear model test, β-binomial overdispersion correction test, and hierarchical label permutation test.
[0089] Specifically, statistical tests are performed on the changes in the strength of evidence for the same target candidate superedge under different acquisition conditions or genetic backgrounds to identify whether multi-site interaction relationships exhibit increased, decreased, or reconnected characteristics. Here, "difference" refers to a statistically significant change in the support or consistency of the same target candidate superedge under different conditions. "Reconnected" refers to a target candidate superedge gaining stronger common evidence support in one set of conditions while its support significantly weakens in another set, or exhibiting a reconstructed trend in connectivity relationships within the breakpoint neighborhood, haplotype, or allelic background. By further comparing the statistical differences between different conditions based on the randomized controlled significance assessment completed in step S105, candidate superedges that are significant only under a single condition can be distinguished from candidate superedges that undergo systematic changes when conditions are switched, providing a basis for subsequent event interpretation and verification.
[0090] At least two sets of conditions or groupings are used to define the comparison objects for the statistical test of differences or reconnection. These can be different samples corresponding to different sources of multi-linked read interaction data, different treatment conditions resulting from different experimental or treatment methods, different cell cluster groups formed after dividing read evidence by cell clusters, or different haplotype or allelic groups formed after attributing read evidence based on haplotype or allelic information. In this step, different samples are used to reflect differences at the data source level, different treatment conditions are used to reflect differences brought about by external intervention or changes in experimental conditions, different cell clusters are used to reflect population differences in interaction relationships under conditions of cellular heterogeneity, and different haplotypes or alleles are used to reflect allelic-specific changes in interaction relationships under conditions of genetic background differences. The above grouping methods can be used individually or in combination while maintaining the established structural stratification, so as to exclude irrelevant structural biases when comparing differences.
[0091] When performing statistical tests for differences or reconnections, it is first necessary to summarize comparable evidence statistics for the same target candidate superedge under different conditions or groups. Evidence statistics can be derived from the observed support, evidence consistency coefficient, or consistency statistic in step S104, or from the count results used for significance assessment in step S105, as long as they use consistent statistical methods across different groups. To ensure that the statistical test is traceable and does not introduce additional ambiguity, the summary process can trace back to the read evidence set within each group using the evidence pointer corresponding to the target candidate superedge. Within each group, the number of read evidence pieces that satisfy the conditions for covering the anchor point set is counted, and the union size of the read evidence pieces related to each anchor point is also counted, thus forming count pairs or ratio pairs for testing. Based on these group-specific statistics, a contingency structure or regression structure for difference testing can be constructed, and the statistical determination of differences or reconnections can be completed accordingly.
[0092] When using Fisher's exact test or chi-square test, a contingency table is typically constructed by combining the support and non-support counts of the target candidate hyperedge under the two sets of conditions. The support count represents the number of reads within each group that satisfy the requirement of covering the anchor point set of the target candidate hyperedge, while the non-support count represents the difference between the total number of reads included in the statistics and the support count within each group. Fisher's exact test is suitable for cases with small or sparse counts, assessing the significance of the difference by evaluating the exact distribution of the contingency table under a fixed margin. Chi-square test is suitable for cases with large counts and where the expected frequency meets the requirement, assessing the significance of the difference by comparing the deviation between the observed frequency and the expected frequency. These tests can determine whether there is a significant difference in the support ratio of the target candidate hyperedge under the two sets of conditions, thus completing the basic determination of discrepancies or reconnections.
[0093] When using negative binomial or generalized linear models for testing, the support count of target candidate hyperedges within each condition or group is considered as the response variable, while condition identifiers, group identifiers, and necessary structural stratification factors are used as explanatory variables. The significance of the conditional effect is assessed through a regression model. Negative binomial models are suitable for cases where count data exhibit excessive dispersion with variance greater than the mean, and can more robustly characterize the fluctuations in read support counts. Generalized linear models can adapt different statistical measures of evidence after selecting appropriate link functions and error distributions, such as using support counts as the response variable, or using the total number of reads or anchor point heat as biases or covariates, thereby testing the significance of conditional differences while controlling for confounding factors. These models are particularly suitable for unified modeling and joint testing when multiple conditions, multiple cell clusters, or multi-level groupings exist, avoiding the information loss caused by relying solely on pairwise contingency tables.
[0094] When using the β-binomial overdispersion correction test, the support ratio of the target candidate hyperedge within each group is treated as a random variable with inter-individual random fluctuations. The population difference in the ratio parameter is characterized by a β distribution, and overdispersion correction is introduced on the basis of the binomial observation model, so that robust significance assessment can still be obtained when there is heterogeneity between different samples or different cell clusters. The core of this test is to consider both intra-group count fluctuations and inter-group ratio fluctuations, thereby reducing the risk of false alarms due to extreme samples or local noise. It is suitable for situations where the support ratio of reads is unevenly distributed among different haplotypes or different cell clusters.
[0095] When using the stratified label permutation test, while maintaining the structural stratification constraints set in step S105, the conditional labels or group labels are randomly permuted to construct a control distribution of the difference statistics under the null hypothesis. During the permutation process, at least two of the following can be kept unchanged: chromosome stratification, hyperedge order stratification, span or distance stratification, and anchor point heat stratification, ensuring that the permuted data is comparable to the true data in key structural attributes. For each permutation, the between-group difference statistics of the target candidate hyperedge are recalculated, and the observed difference statistics are compared with the permutation distribution to assess its significance under stratification constraints. This test is suitable for scenarios insensitive to model distribution assumptions and can robustly validate differences or reconnections under complex stratification constraints, maintaining methodological consistency with the aforementioned significance discrimination null model.
[0096] Step S106 further includes performing a difference or reconnection statistical test on the target candidate hyperedge between at least two sets of conditions or groups. The at least two sets of conditions or groups may include at least one of different samples, different treatment conditions, different cell clusters, and different haplotypes or alleles, used to characterize the difference in evidence strength of the same multi-site interaction structure under changes in conditions or genetic background. In specific implementation, firstly, using the hyperedge signature as a unified key, the valid candidate hyperedges from different conditions or groups are aligned using a union. Then, within each condition or group, the support corresponding to the target candidate hyperedge is statistically analyzed and denoted as:
[0097]
[0098] in, Indicates conditions or group identifiers. This represents the target candidate superedge; a minimum order constraint can be further applied during the statistical process. And a minimum support threshold is used to exclude low-order or insufficiently supported hyperedges, thereby ensuring the consistency of the comparison objects at the structural level between different conditions.
[0099] After summarizing support across conditions or groups, the difference or reconnection effect size is calculated for the target candidate hyperedge to characterize the direction and magnitude of changes in evidence strength between different conditions. In the case of comparing two groups of conditions, the difference effect size can be expressed as:
[0100]
[0101] in, and These represent the support of the target candidate hyperedge under the two sets of conditions, respectively. This is a smoothing constant used to avoid computational instability caused by zero values; in another representation, the support difference can also be directly used as the effect size. :
[0102] This method is used to characterize the absolute change in support strength of candidate hyperedges before and after condition switching. Based on the effect size calculation, difference or reconnection statistical tests are performed on the target candidate hyperedges to determine whether the observed support changes significantly deviate from random fluctuations. Statistical tests can include Fisher's exact test or chi-square test. Fisher's test can be used to test whether there are significant differences in support ratios under different conditions by constructing a contingency structure consisting of support and non-support counts. When support counts in a reading exhibit excessive dispersion or when multi-condition joint modeling is required, a negative binomial model or a generalized linear model can be used, with support counts as the response variable and condition identifiers and necessary structural stratification factors as explanatory variables, thereby assessing the conditional effect while controlling for confounding factors. When there is intra-group heterogeneity in support ratios, a β-binomial excessive dispersion correction test can be used to improve the robustness of the difference test by simultaneously characterizing the fluctuation of the ratio parameter and the uncertainty of the observed counts. In addition, randomization-based null models can be used for testing, including shuffling condition labels while keeping chromosome grouping unchanged, performing degree preservation redistribution while maintaining the overall support structure, and performing hierarchical label permutation tests under multiple structural constraints, to obtain the results of differential significance assessment in complex structural contexts.
[0103] When using the hierarchical label permutation test, a set of interference features for hierarchical classification can be constructed, including chromosome identifier, hyperedge order, anchor span or distance, and anchor heat. Then, the order, distance, and heat are binned using quantiles to obtain... Then, it combines with chromosome markers to form layered bonds. In each Internally, permutations are performed on conditions or group labels to construct a randomized control distribution of the difference statistics while keeping the key structural distributions unchanged, and the significance of the true effect size is then assessed accordingly.
[0104] After obtaining the significance probability value corresponding to each target candidate hyperedge, to control the risk of false discovery caused by multiple comparisons, multiple tests are performed to correct the significance probability values of all target candidate hyperedges. Specifically, let the total number of target candidate hyperedges participating in the difference or reconnection test be... The significance probability value corresponding to each candidate hyperedge is denoted as . And sort them in ascending order to get: Let the preset error detection rate control level be: And calculate the corresponding threshold sequence. :
[0105]
[0106] Determine the largest index in an ordered sequence. Make Based on this, the target candidate hyperedge set controlled by the false discovery rate is determined. To obtain continuously comparable correction results, an initial correction amount is calculated for each significance probability value after sorting:
[0107]
[0108] The final corrected significance probability value is obtained through monotonicity correction:
[0109] Then The results are backfilled to the corresponding target candidate hyperedges and output as the corrected significance results of the difference or reconnection statistical test. Through the above steps, step S106 can not only determine whether there are significant differences or reconnections between multi-site interaction evidence under different conditions or groups, but also stably control the false discovery rate in the case of multiple comparisons, providing a reliable basis for subsequent event-level interpretation and biological validation.
[0110] Step S107: Output the hyperedge evidence analysis results and verifiable data packets based on the corrected significance probability values; the verifiable data packets are used to backtrack the evidence pointers.
[0111] Specifically, based on the corrected significance probability values, the target candidate hyperedges that have passed the preset significance threshold are compiled and summarized to form hyperedge evidence analysis results, and verifiable data packages for backtracking evidence pointers are generated simultaneously. The hyperedge evidence analysis results are a set of structured analysis conclusions obtained by summarizing the target candidate hyperedges after significance discrimination and multiple testing correction. These conclusions describe the anchor point composition, order, support, consistency statistic, significance probability value, and differences or reconnection effects under different conditions or groupings for each target candidate hyperedge. The verifiable data package includes an input data summary, parameter and software version information, a list of result files, checksum or hash check information, and a minimal backtracking script or interface description for extraction or recalculation.
[0112] In the output process, a set of result files describing the overall analysis conclusions of multi-site interactions is first generated. The candidate hyperedge set is then sorted according to the hyperedge signature or its corresponding hash identifier, generating a hyperedge result table. This result table includes at least the hyperedge identifier or signature information used to uniquely identify the hyperedge, the corresponding anchor list, the hyperedge order, and the observation support. This includes span or distance statistics used to characterize the spatial distribution of anchor points, and may also include stratification key information for subsequent stratification or difference analysis. Based on this, and combining the significance discrimination null model and multiple test correction process constructed in step S105, a significance result table is generated. This result table must contain at least the significance probability value corresponding to each target candidate hyperedge. and the corrected significance probability value Simultaneously record the consistency coefficient of the evidence. The system also includes information on the null model type and key parameters used to reproduce the significance assessment process, such as the number of permutations or samplings, the random seed, and the structural stratification scheme used. If group comparisons were performed in step S106, a difference or reconnection result table is further generated to record the support statistics and corresponding effect sizes of the same target candidate hyperedge under different conditions or groups. Alternatively, the support difference, the statistical test method used, and the corresponding p-value and q-value. If the analysis involves the neighborhood of structural variation breakpoints, an event-level reconnection result table can also be generated to describe the support of hyperedges within the breakpoint window. This result table should at least include the event identifier, breakpoint window size, statistics under different conditions or isotropic conditions, and event-level effect size.
[0113] After generating the above result files, the next step is to output evidence pointers and index information for evidence backtracking. In the preceding steps, a corresponding evidence pointer has been generated and saved for each original read segment, and its supporting read segment's evidence pointer set is stored in the hyperedge signature record. Based on this association, an evidence pointer lookup table is generated to establish the mapping relationship between the hyperedge identifier and its supporting read segment evidence pointer set. The evidence pointer lookup table can be stored in blocks or compressed, and supports the extraction of representative evidence pointers as needed, thus ensuring that each analysis result can be traced back to the corresponding original read segment evidence without significantly increasing the storage burden.
[0114] In addition to the results and evidence pointers, a parameter list and verification information for reproduction and validation are output synchronously. The parameter list should include at least the input data summary information and the lower bound of the hyperedge order. Key information such as support and significance thresholds, zero-model configuration, random seed, software version number, and runtime timestamp are used to fully record the configuration environment for this analysis. Simultaneously, checksums are calculated for all output files, for example, using a hash algorithm to generate checksum values, and a unified manifest file is created to record the name, size, checksum, and generation order of each output file, thereby verifying the integrity and consistency of the result files during subsequent transmission or storage.
[0115] After the aforementioned result file, evidence pointer lookup table, and parameter and verification information are generated, they are encapsulated into a verifiable data package. The verifiable data package also includes a minimal backtracking script or interface description, used to guide the extraction of the corresponding read segment from the original multi-connection read segment data based on the evidence pointer when needed, and to recalculate... , The results include difference or reconnection statistics, and the recalculated results are compared with the results recorded in the verifiable data package to achieve verifiable recalculation of the hyperedge evidence analysis results. Through the above method, step S107 achieves a closed-loop output between the analysis conclusion, statistical basis and original evidence, so that each multi-site interaction result judged to be significant or showing difference reconnection has a clear data source and a reproducible path.
[0116] In one possible implementation, step S107 further includes: analyzing the conditions of the structural variation breakpoint neighborhood based on the hyperedge evidence analysis results and verifiable data packets, specifically including: obtaining a set of structural variation breakpoints based on the hyperedge evidence analysis results and verifiable data packets, and defining a breakpoint window of a preset length for each structural variation breakpoint; determining whether the genomic region corresponding to any anchor point of a valid candidate hyperedge overlaps with the breakpoint window; a valid candidate hyperedge is a target candidate hyperedge that has passed the significance discrimination; if the genomic region corresponding to any anchor point of a valid candidate hyperedge overlaps with the breakpoint window, then the valid candidate hyperedge is confirmed to have hit the breakpoint window; calculating the event-level statistics of the valid candidate hyperedges that hit the breakpoint window, and calculating the event-level difference based on the event-level statistics; the event-level statistics include the number of hit candidate hyperedges, the total support of hit candidate hyperedges, and the average order of hit candidate hyperedges; outputting a result table corresponding to the conditions of the structural variation breakpoint neighborhood based on the event-level difference, and grouping the result table according to haplotype or allele.
[0117] Specifically, based on the hyperedge evidence analysis results and verifiable data packets, a set of structural variation breakpoints is read. This set can be represented using BEDPE or VCF style fields, where each structural variation event must contain at least an event identifier. and the chromosome and genome coordinates where the breakpoint is located And, in the case of a double breakpoint description, further include Based on the breakpoint coordinates, define a breakpoint window of preset length for each structural variation breakpoint. The breakpoint window is represented on the genome coordinate system as follows: Used to define genomic regions that are spatially adjacent to breakpoints.
[0118] After the breakpoint window is defined, target candidate superedges that pass the saliency test are considered valid candidate superedges for hit determination. If the genomic region corresponding to any anchor point in the valid candidate superedge is on the same chromosome as the breakpoint window... If interval overlap occurs, the valid candidate superedge hits the breakpoint window; for structural mutation events containing the second breakpoint, respectively in The above hit determination is executed within the corresponding breakpoint window, and the hit results are merged and statistically analyzed.
[0119] After the hit determination is completed, event-level statistics are performed on the valid candidate hyperedges of the hit breakpoint window under different conditions or groups. The event-level statistics include at least the number of hyperedges of the hit breakpoint window. Used to characterize the number of significant multi-site interaction structures within the neighborhood of a breakpoint; total support of hit candidate hyperedges. Used to characterize the overall strength of evidence of multi-site interactions within the neighborhood of a breakpoint; and the average order of the hit candidate hyperedges. Used to characterize the complexity of the interaction structure of multiple sites within the neighborhood of a breakpoint.
[0120] Between different conditions or groupings, calculate the event-level differences based on the above event-level statistics. Among these, the event-level support difference... It can be represented as:
[0121]
[0122] in, The sum of support for all valid candidate hyperedges that hit the same structural variation breakpoint window under condition A or group A is represented by the sum of the observed support for each valid candidate hyperedge under condition A. This sum of support is used to characterize the overall strength of evidence of multi-site interaction within the breakpoint neighborhood under condition A. This represents the sum of support for all valid candidate hyperedges hitting the same structural variation breakpoint window under condition B or group B, and its statistical caliber is the same as... To maintain consistency, this is used to characterize the overall strength of multi-site interaction evidence within the neighborhood of the breakpoint under condition B, thus providing a comparable basic statistic for calculating event-level variance or logarithmic change. Furthermore, to facilitate comparisons across event scales, the event-level logarithmic change can be further calculated. The relevant expression is:
[0123] Among them, subscript These represent the two sets of conditions or groups being compared.
[0124] If a valid candidate superedge contains a haplotype or equivalence identifier, further [the following can be done] based on [the following]. The field groups the valid candidate superedges of the breakpoint window, and repeatedly performs breakpoint hit determination, event-level statistics summary and event-level difference calculation within each haplotype or isoplethysmographic group, thereby obtaining isoplethysmographic or haplotype-specific reconnection features in the neighborhood of the structural variation breakpoint.
[0125] Finally, based on event-level statistics and event-level differences, a result table corresponding to the neighborhood of structural variation breakpoints is output. The result table contains at least the following: , , , , , as well as And can be in accordance with Alternatively, output can be grouped using equivalence markers. The results table is associated with the evidence pointers, parameter lists, and backtracking scripts in the verifiable data package, enabling the analysis results of each structural variation breakpoint neighborhood to be traced back to the corresponding hyperedge evidence and original read data, thus achieving verifiable recalculation of event-level results.
[0126] Please refer to Figure 2 This document illustrates a schematic diagram of a hyperedge evidence analysis device for multi-site chromatin interactions provided in an embodiment of this application. The device includes a data preprocessing module 21, a candidate mining module 22, a significance evaluation module 23, a testing module 24, and an output module 25. Data preprocessing module 21 is used to decompose the target read into multiple genomic fragments and map them to an anchor set; the target read is any read data in the multi-connected read interaction data; the anchor set hit by the target read is normalized to obtain a superedge signature, and at least one evidence pointer is saved for each superedge signature to form a superedge evidence transaction; The candidate mining module 22 is used to perform incremental candidate mining on hyperedge evidence transactions under preset biological and statistical constraints, and to obtain a set of candidate hyperedges using preset pruning rules. The significance assessment module 23 is used to calculate the consistency statistic between the target candidate hyperedge and the reading segment evidence; the target candidate hyperedge is any candidate hyperedge in the candidate hyperedge set; the reading segment evidence is obtained by backtracking through the evidence pointer associated with the target candidate hyperedge; a significance discrimination null model is constructed based on the consistency statistic, and the significance probability value corresponding to the target candidate hyperedge is output through the significance discrimination null model; The testing module 24 is used to perform multiple tests on the significance probability value of each target candidate hyperedge and output the corrected significance probability value.
[0127] Output module 25 is used to output the hyperedge evidence analysis results and verifiable data packets based on the corrected significance probability values; the verifiable data packets are used to backtrack the evidence pointer.
[0128] In one possible implementation, the data preprocessing module 21 is used to normalize the set of anchor points hit by the target read segment to obtain a superedge signature. Specifically, it includes: performing deduplication and sorting processing on the set of anchor points hit by the target read segment to obtain a superedge signature; and further representing the superedge signature as an encoding form for counting or querying, wherein the encoding form includes a bitmap, a hash value, or a compressed integer sequence.
[0129] In one possible implementation, the data preprocessing module 21 is used to normalize the set of anchor points that the target read segment hits to obtain a super-edge signature, and save at least one evidence pointer for each super-edge signature to form a super-edge evidence transaction. The method further includes: constructing a set of read segment evidence pointers for each target read segment; the evidence pointer includes a read segment identifier, original file offset information and parsed index key; and constructing an inverted index mapping from anchor points to the set of read segment evidence pointers as the anchor point set.
[0130] In one possible implementation, the consistency statistic includes at least observational support and evidence consistency coefficient. The significance assessment module 23 is used to calculate the consistency statistic between the target candidate superedge and the read segment evidence, specifically including: calculating the observational support by counting the number of read segment evidences that meet a first preset consistency condition; the first preset consistency condition is that the read segment evidence can cover the set of anchor points corresponding to the target candidate superedge; calculating the evidence consistency coefficient by counting the degree of overlap between read segment evidences that meet a second preset consistency condition; the second preset consistency condition is that the read segment evidence can appear simultaneously on each anchor point corresponding to the target candidate superedge; and calculating the consistency statistic between the target candidate superedge and the read segment evidence based on the observational support and the evidence consistency coefficient.
[0131] In one possible implementation, the significance assessment module 23 is used to construct a significance null model based on consistency statistics and output the significance probability value corresponding to the target candidate hyperedge through the significance null model. Specifically, it includes: constructing a significance null model based on consistency statistics that maintains at least one structural layering; the layering includes at least two of chromosome layering, hyperedge order layering, span or distance layering, and layering based on anchor point heat; calculating the significance probability value based on the significance null model through a first random control method or a second random control method; the first random control method is to perform translation and permutation of the anchor point identifier corresponding to the target candidate hyperedge within the contig or chromosome; the second random control method is to perform degree-preserving rearrangement sampling of the anchor point-read segment association relationship while keeping the anchor point chromosome count unchanged.
[0132] In one possible implementation, the testing module 24 is configured to, after performing multiple tests on the significance probability value of each target candidate hyperedge and outputting the corrected significance probability value: perform a difference or reconnection statistical test on the target candidate hyperedge between at least two sets of conditions or groups; the at least two sets of conditions or groups include at least one of different samples, different treatment conditions, different cell clusters, and different haplotypes or alleles; the difference or reconnection statistical test includes at least one or more of Fisher's exact test, chi-square test, negative binomial or generalized linear model test, β-binomial overdispersion correction test, and hierarchical label permutation test.
[0133] In one possible implementation, the output module 25 is used to output a verifiable data packet, specifically including: outputting a verifiable data packet based on the corrected significance probability value, wherein the verifiable data packet includes an input data summary, parameter and software version information, a list of result files, checksum or hash check information, and a minimal backtracking script or interface description for extraction or recalculation.
[0134] In one possible implementation, the output module 25 is used to: after outputting the hyperedge evidence analysis results and verifiable data packets based on the corrected significance probability values, analyze the conditions of the structural variation breakpoint neighborhood based on the hyperedge evidence analysis results and verifiable data packets, specifically including: obtaining a set of structural variation breakpoints based on the hyperedge evidence analysis results and verifiable data packets, and defining a breakpoint window of a preset length for each structural variation breakpoint; determining whether the genomic region corresponding to any anchor point of a valid candidate hyperedge overlaps with the breakpoint window; a valid candidate hyperedge is a target candidate hyperedge that passes the significance discrimination; if the genomic region corresponding to any anchor point of a valid candidate hyperedge overlaps with the breakpoint window, then confirming that the valid candidate hyperedge hits the breakpoint window; calculating the event-level statistics of the valid candidate hyperedges that hit the breakpoint window, and calculating the event-level difference based on the event-level statistics; the event-level statistics include the number of hit candidate hyperedges, the total support of hit candidate hyperedges, and the average order of hit candidate hyperedges; outputting a result table corresponding to the conditions of the structural variation breakpoint neighborhood based on the event-level difference, and grouping the result table according to haplotype or allele.
[0135] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0136] This application also provides an electronic device. (See reference...) Figure 3 , Figure 3This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: at least one processor 301, at least one communication bus 302, a user interface 303, at least one network interface 304, and a memory 305.
[0137] The communication bus 302 is used to enable communication between these components.
[0138] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0139] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0140] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.
[0141] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include a non-transitory computer-readable storage medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. (Refer to...) Figure 3 The memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a hyperborder evidence analysis application for multi-site chromatin interactions.
[0142] exist Figure 3 In the illustrated electronic device, the user interface 303 is primarily used to provide an input interface for the user and acquire user input data; while the processor 301 can be used to call the hyperedge evidence analysis application storing multi-site chromatin interaction data in the memory 305. When executed by one or more processors 301, the electronic device performs one or more of the methods described in the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0143] This application also provides a non-transitory computer-readable storage medium storing instructions. When executed by one or more processors, these instructions cause an electronic device to perform one or more of the methods described in the above embodiments.
[0144] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0145] In the various embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.
[0146] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0147] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0148] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0149] The above description is merely an exemplary embodiment disclosed in this application and should not be construed as limiting the scope of this application. Any equivalent changes and modifications made in accordance with the teachings of this application shall still fall within the scope of this application.
[0150] This application is intended to cover any variations, uses, or adaptations disclosed herein that follow the general principles disclosed herein and include common knowledge or customary technical means in the art that are not described in this application.
Claims
1. A method for analyzing hyperborder evidence of multi-site chromatin interactions, characterized in that, The method includes: The target read is decomposed into multiple genomic fragments and mapped to a set of anchor points; the target read is any read data in the multi-linked read interaction data. The set of anchor points hit by the target read segment is normalized to obtain a super-edge signature, and at least one evidence pointer is saved for each super-edge signature to form a super-edge evidence transaction; Under preset biological and statistical constraints, incremental candidate mining is performed on the superedge evidence transactions, and a set of candidate superedges is obtained using preset pruning rules. Calculate the consistency statistic between the target candidate superedge and the read segment evidence; the target candidate superedge is any candidate superedge in the set of candidate superedges; the read segment evidence is obtained by backtracking through the evidence pointer associated with the target candidate superedge; A significance null model is constructed based on the consistency statistic, and the significance probability value corresponding to the target candidate hyperedge is output through the significance null model. Perform multiple tests on the significance probability value of each of the target candidate hyperedges, and output the corrected significance probability value; Based on the corrected significance probability value, output the hyperedge evidence analysis results and verifiable data packets; the verifiable data packets are used to backtrack the evidence pointer.
2. The method according to claim 1, characterized in that, The process of normalizing the set of anchor points matched by the target read segment to obtain the superedge signature specifically includes: The set of anchor points hit by the target read segment is deduplicated and sorted to obtain the super-edge signature; The superedge signature is further represented as an encoded form for counting or querying, the encoded form including bitmaps, hash values, or compressed integer sequences.
3. The method according to claim 1, characterized in that, After normalizing the set of anchor points in the target read segment to obtain a superedge signature, and storing at least one evidence pointer for each superedge signature to form a superedge evidence transaction, the method further includes: Construct a set of read segment evidence pointers corresponding to each target read segment; the evidence pointer includes a read segment identifier, original file offset information, and a parsing index key; Construct an inverted index mapping from the anchor point to the set of read segment evidence pointers as the anchor point set.
4. The method according to claim 1, characterized in that, The consistency statistics include at least the observation support and the evidence consistency coefficient. Specifically, the calculation of the consistency statistics between the target candidate hyperedge and the read segment evidence includes: The observation support is calculated by counting the number of read segment evidences that satisfy the first preset consistency condition; the first preset consistency condition is that the read segment evidences can cover the set of anchor points corresponding to the target candidate hyperedge. The evidence consistency coefficient is calculated by statistically analyzing the degree of overlap between the read segment evidences that satisfy the second preset consistency condition; the second preset consistency condition is that the read segment evidences can appear simultaneously at each anchor point corresponding to the target candidate hyperedge. The consistency statistic between the target candidate hyperedge and the read segment evidence is calculated based on the observation support and the evidence consistency coefficient.
5. The method according to claim 1, characterized in that, The step of constructing a significance null model based on the consistency statistic and outputting the significance probability value corresponding to the target candidate hyperedge through the significance null model specifically includes: The significance-discriminating null model is constructed based on the consistency statistic, which maintains at least one structural layering; the layering includes at least two of the following: chromosome layering, hyperedge order layering, span or distance layering, and layering based on anchor point heat. Based on the saliency discrimination null model, the saliency probability value is calculated by a first random control method or a second random control method; the first random control method is to perform translation and permutation on the anchor point identifier corresponding to the target candidate superedge within the contig or chromosome; the second random control method is to perform degree-preserving rearrangement sampling on the anchor point-read segment association relationship while keeping the anchor point chromosome count unchanged.
6. The method according to claim 1, characterized in that, After performing multiple tests on the significance probability value of each of the target candidate hyperedges and outputting the corrected significance probability value, the method further includes: Perform a difference or reconnection statistical test on the target candidate hyperedge between at least two sets of conditions or groups; the at least two sets of conditions or groups include at least one of different samples, different treatment conditions, different cell clusters, and different haplotypes or alleles; the difference or reconnection statistical test includes at least one or more of Fisher's exact test, chi-square test, negative binomial or generalized linear model test, β-binomial overdispersion correction test, and hierarchical label permutation test.
7. The method according to claim 1, characterized in that, Output verifiable data packets, specifically including: The verifiable data package is output based on the corrected significance probability value. The verifiable data package includes an input data summary, parameter and software version information, a list of result files, checksum or hash check information, and a minimal backtracking script or interface description for extraction or recalculation.
8. The method according to claim 1, characterized in that, After outputting the hyperedge evidence analysis results and verifiable data packets based on the corrected significance probability values, the method further includes: Based on the results of the hyperedge evidence analysis and the conditions of the neighborhood of the structural variation breakpoint in the verifiable data packet analysis, the specific conditions include: Based on the hyperedge evidence analysis results and the verifiable data packet, a set of structural variation breakpoints is obtained, and a breakpoint window of a preset length is defined for each structural variation breakpoint. Determine whether the genomic region corresponding to any anchor point of the valid candidate superedge overlaps with the breakpoint window; the valid candidate superedge is the target candidate superedge that has been identified through saliency determination; If the genomic region corresponding to any anchor point of the valid candidate superedge overlaps with the breakpoint window, then the valid candidate superedge is confirmed to have hit the breakpoint window. The event-level statistics of the valid candidate superedges that hit the breakpoint window are statistically analyzed, and the event-level difference is calculated based on the event-level statistics; the event-level statistics include the number of hit candidate superedges, the total support of hit candidate superedges, and the average order of hit candidate superedges. Based on the event-level difference, output the result table corresponding to the conditions of the structural variation breakpoint neighborhood, and output the result table in groups according to haplotype or isotype.
9. A hyperborder evidence analysis device for multi-site chromatin interactions, characterized in that, The device includes a data preprocessing module, a candidate mining module, a significance evaluation module, a testing module, and an output module, wherein... The data preprocessing module is used to decompose the target read into multiple genomic fragments and map them to an anchor set; the target read is any read in the multi-connected read interaction data; the anchor set hit by the target read is normalized to obtain a superedge signature, and at least one evidence pointer is saved for each superedge signature to form a superedge evidence transaction; The candidate mining module is used to perform incremental candidate mining on the hyperedge evidence transactions under preset biological and statistical constraints, and to obtain a set of candidate hyperedges using preset pruning rules. The significance evaluation module is used to calculate the consistency statistic between the target candidate hyperedge and the read segment evidence; the target candidate hyperedge is any candidate hyperedge in the set of candidate hyperedges; the read segment evidence is obtained by backtracking through the evidence pointer associated with the target candidate hyperedge; a significance discrimination null model is constructed based on the consistency statistic, and the significance probability value corresponding to the target candidate hyperedge is output through the significance discrimination null model; The testing module is used to perform multiple tests on the significance probability value of each target candidate hyperedge and output the corrected significance probability value. The output module is used to output the hyperedge evidence analysis results and verifiable data packets based on the corrected significance probability value; the verifiable data packets are used to backtrack the evidence pointer.
10. An electronic device, characterized in that, The device includes a processor, a communication bus, a user interface, a network interface, and a memory. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The communication bus is used to enable communication between the components within the electronic device. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-8.