Method, device, terminal and medium for determining adjacent connection signals generated between different genomes
By obtaining the target genome set and relative abundance of multi-genome mixed DNA samples, and estimating the number of adjacent connections for alignment, the problem of false positive connection interference was solved, achieving high accuracy and low false positive neighbor relationship judgment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
- Filing Date
- 2026-03-06
- Publication Date
- 2026-05-15
AI Technical Summary
In existing technologies, false positive connections significantly interfere with the determination of proximity relationships between different genomes in a sample, resulting in severe noise signal interference.
By acquiring the target genome set of a multi-genome mixed DNA sample, the relative abundance and total number of neighboring connections of each genome are determined. Based on the relative abundance, the expected number of neighboring connections of any genome pair under completely random spatial collision conditions is estimated. The authenticity of the neighboring connection signal is determined by comparing the actual number of neighboring connections with the expected number of neighboring connections.
It effectively filters out noisy connections, preserves true spatial proximity connections, reduces the interference of false positive connections, and improves the accuracy and stability of proximity relationship judgment.
Smart Images

Figure CN121811976B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics analysis technology, and in particular to methods, devices, terminals and media for determining proximity connection signals generated between different genomes. Background Technology
[0002] Sample DNA (Deoxyribonucleic acid) spatial proximity joining technology obtains information on the spatial proximity relationships between different DNA fragments by cross-linking, cleaving, and joining DNA molecules in situ under certain conditions, followed by high-throughput sequencing of the ligation products, without depending on culture conditions. In recent years, this technology has been introduced into the fields of microbiome and environmental sample research and has been used to resolve the spatial associations between different microorganisms, viruses, and various mobile genetic elements at the mixed genome scale. For example, in virus-host relationship studies, sample DNA spatial proximity joining technology is widely used to infer the correspondence between extrachromosomal genetic elements such as viruses or plasmids and their potential microbial hosts. Its basic principle is that if a viral genome and a host genome have an in situ infection or spatial proximity relationship in a sample, a connection signal can be generated between them.
[0003] However, during experiments, the presence of abundant extracellular DNA in samples—such as DNA released from dead cells, DNA released from ruptured cells during sample handling, and free viral DNA molecules—can lead to the formation of numerous false-positive connections between different genomes in subsequent high-throughput sequencing data. These false-positive connections significantly interfere with the assessment of proximity relationships between different genomes within the sample.
[0004] Therefore, existing technologies have shortcomings and need to be improved and developed. Summary of the Invention
[0005] The technical problem to be solved by this application is to provide a method, device, terminal and medium for determining the proximity connection signal generated between different genomes, in order to address the above-mentioned defects of the prior art. The aim is to solve the problem that false positive connections in the prior art can significantly interfere with the determination of the proximity relationship between different genomes in a sample.
[0006] The technical solution adopted by this application to solve the technical problem is as follows:
[0007] A method for determining proximity connectivity signals between different genomes, wherein the method includes:
[0008] Obtain the target genome set corresponding to the multi-genome mixed DNA sample, and determine the relative abundance of each genome in the target genome set and the total number of neighbor connections in the target genome set;
[0009] The expected number of neighbor connections for any pair of genomes under completely random spatial collision conditions is estimated based on the relative abundance of each genome and the total number of neighbor connections.
[0010] The actual number of neighbor connections for each genome pair is obtained. Based on the actual number of neighbor connections and the expected number of neighbor connections for each genome pair, the determination result of the neighbor connection signal generated between each genome pair is determined.
[0011] In one embodiment of this application, estimating the expected number of neighbor connections for any pair of genomes under completely random spatial collision conditions based on the relative abundance of each genome and the total number of neighbor connections includes:
[0012] The probability of random spatial collision between any two genomes is calculated based on the relative abundance of each genome.
[0013] The product of the random spatial collision probability and the total number of neighbor connections is calculated to obtain the expected number of neighbor connections for any genome pair under completely random spatial collision conditions.
[0014] In one embodiment of this application, calculating the probability of random spatial collision between any two genomes based on the relative abundance of each genome includes:
[0015] Based on the relative abundance of each genome, the product of the relative abundance of any pair of genomes in the target genome set is calculated to obtain the product of the relative abundance of each pair of genomes in the target genome set.
[0016] Summing the products of the relative abundance of all genome pairs yields the summation result;
[0017] The ratio of the product of the relative abundances to the summation result is used as the probability of random spatial collision between the two genomes.
[0018] In one embodiment of this application, calculating the probability of random spatial collision between any two genomes based on the relative abundance of each genome includes:
[0019] The weight correction factor for each genome is determined based on the preset weight correction information;
[0020] The relative abundance weighted value of each genome is calculated based on the relative abundance of each genome and the weight correction factor.
[0021] Calculate the product of the relative abundance weights of any pair of genomes in the target genome set to obtain the product of the relative abundance weights of each pair of genomes in the target genome set;
[0022] Summing the products of the relative abundance weights of all genome pairs yields the summation result;
[0023] The ratio of the product of the relative abundance weighted values to the summation result is used as the probability of random spatial collision between the two genomes;
[0024] The weight correction information includes at least one of the following: GC content, restriction site density, and adjacent linker fragment recovery efficiency.
[0025] In one embodiment of this application, the actual number of neighboring connections for each genome pair is obtained, and based on the actual number of neighboring connections and the expected number of neighboring connections for each genome pair, the determination result of the neighboring connection signal generated between each genome pair is determined, including:
[0026] Count the number of actual neighbor connections observed for each genome pair;
[0027] Based on the actual number of neighbor connections and the expected number of neighbor connections for each genome pair, a predetermined index statistic is obtained;
[0028] Obtain a pre-defined threshold, and compare the predetermined indicator statistics of each genome pair with the threshold to obtain the comparison result;
[0029] Based on the comparison results, the determination of the proximity connection signal generated between the two genomes in each genome pair is obtained.
[0030] In one embodiment of this application, the determination result of the proximity connection signal generated between two genomes in each genome pair based on the alignment result includes:
[0031] If the alignment result is a ratio less than the threshold, then the proximity connection signal generated between the two genomes in the genome pair is an extracellular random connection signal;
[0032] If the alignment result is a ratio greater than or equal to the threshold, then the proximity connection signal generated between the two genomes in the genome pair is the intracellular real spatial proximity signal;
[0033] The threshold is obtained by evaluating the distribution characteristics of the ratio of actual neighbor connections to expected neighbor connections based on known real genomes in several constructed simulated community datasets. The threshold includes at least one of a fixed empirical threshold, a quantile threshold, and a sample adaptive threshold.
[0034] In one embodiment of this application, the actual number of neighboring connections for each genome pair is obtained, and based on the actual number of neighboring connections and the expected number of neighboring connections for each genome pair, the determination result of the neighboring connection signal generated between each genome pair is determined, including:
[0035] Count the number of actual neighbor connections observed for each genome pair;
[0036] Based on the actual number of neighbor connections and the expected number of neighbor connections for each genome pair, a predetermined index statistic is obtained;
[0037] A pre-constructed statistical model is obtained, and the predetermined index statistics are input into the statistical model to obtain the determination result of the proximity connection signal generated between each genome pair.
[0038] This application also provides a device for determining proximity connectivity signals generated between different genomes, the device comprising:
[0039] The acquisition module is used to acquire the target genome set corresponding to the multi-genome mixed DNA sample, and to determine the relative abundance of each genome in the target genome set and the total number of neighboring connections of the target genome set;
[0040] An estimation module is used to estimate the expected number of neighbor connections for any pair of genomes under completely random spatial collision conditions based on the relative abundance of each genome and the total number of neighbor connections.
[0041] The determination module is used to obtain the actual number of neighbor connections for each genome pair, and based on the actual number of neighbor connections and the expected number of neighbor connections for each genome pair, to determine the determination result of the neighbor connection signal generated between each genome pair.
[0042] This application also provides a terminal, including: a memory, a processor, and a proximity connection signal determination program between different genomes stored in the memory and executable on the processor, wherein when the proximity connection signal determination program between different genomes is executed by the processor, it implements the steps of the proximity connection signal determination method between different genomes as described above.
[0043] This application also provides a computer-readable storage medium storing a computer program that can be executed to implement the steps of the method for determining proximity connection signals generated between different genomes as described above.
[0044] Technical effects of this application:
[0045] This application obtains the target genome set corresponding to a multi-genome mixed DNA sample and determines the relative abundance of each genome in the target genome set and the total number of neighboring connections in the target genome set. Based on the relative abundance of each genome and the total number of neighboring connections, it estimates the expected number of neighboring connections between any pair of genomes under completely random spatial collision conditions. It then obtains the actual number of neighboring connections for each genome pair and, based on the actual and expected number of neighboring connections for each genome pair, determines the determination result of the neighboring connection signal generated between each genome pair. This application derives the expected number of neighboring connections under random collision backgrounds through the relative abundance of each genome, filters out noisy connections by comparing the expected and actual number of neighboring connections, and retains true spatial neighboring connections, thereby avoiding the problem of false positive connections significantly interfering with the determination of the proximity relationship between different genomes in the sample. Attached Figure Description
[0046] Figure 1 This is a flowchart of a preferred embodiment of the method for determining proximity connection signals between different genomes in this application.
[0047] Figure 2 This is a functional principle block diagram of a preferred embodiment of the proximity connection signal determination device between different genomes in this application.
[0048] Figure 3 This is a functional principle block diagram of a preferred embodiment of the terminal in this application. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this application clearer and more explicit, the following detailed description of this application is provided with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0050] In existing technologies, effectively filtering noise signals generated by extracellular DNA linkages in sample DNA spatial proximity linkage data while preserving intracellular proximity linkage signals between different genomes to the greatest extent has become a core technical requirement.
[0051] Existing technologies for filtering false-positive linkage signals caused by extracellular DNA in sample DNA spatial proximity linkage data mainly include the following categories:
[0052] First, the minimum evidence thresholds method.
[0053] These methods subjectively infer that Hi-C links with small read counts may not be genuine connections between different genomes. Therefore, they filter out possible false positives by setting a minimum Hi-C read count threshold (e.g., ≥2, ≥5, or ≥10 links). Hi-C is a high-throughput chromosome conformation capture technology specifically designed to study the spatial relationships of chromatin DNA across the entire genome. By capturing spatially adjacent DNA fragment connection data, it reveals the three-dimensional structure of chromatin and its regulatory mechanisms.
[0054] This method is simple to implement and can remove single, accidental connections to some extent, but its threshold is only based on empirical estimation, and its reliability is affected by factors such as sample heterogeneity, differences in sample processing operations, and variations in sequencing depth. Its accuracy lacks theoretical support and experimental verification.
[0055] Second, the relative connectivity or host preference method.
[0056] This type of method does not rely on an absolute read count threshold. Instead, it compares the relative strength (read count) of Hi-C linkages between a viral genome and multiple possible host genomes, and assigns them to the host with the highest linkage ratio or significant enrichment. A representative approach first introduces the virus-host copy number ratio (VPH) to correct for Hi-C linkage abundance.
[0057]
[0058] Based on this, we further constructed a relative connectivity index for virus-host pairs. Taking into account the virus-host connection density, self-connectivity density, and abundance correction term:
[0059]
[0060] In the formula, Indicates the abundance of viral genomes (or vOTUs); Indicates the abundance of the host genome (or MAG); This indicates the number of Hi-C connections between the viral genome and all possible hosts; This represents the sum of Hi-C connections between the viral genome and the entire host genome; This represents the sum of all virus-host Hi-C connections in the sample; Indicates virus With the host Hi-C connectivity density normalized to unit sequence length; Represents the host genome Internal (self-connected) unit sequence length normalized Hi-C connectivity density; VPH represents the estimated average viral copies per host cell. This is a score indicating the relative connectivity or host preference of a virus-host pair.
[0061] Filter and sort virus-host pairings; if Then it is considered a trustworthy connection.
[0062] This strategy can mitigate the impact of sequencing depth differences to some extent and reduce multi-host misjudgment problems, but it is highly sensitive to differences in genome size, abundance, and coverage, and is prone to systematic bias in real multi-host infection scenarios.
[0063] Third, methods based on statistical measures or noise models (model-based filtering).
[0064] This type of method assesses the significance of cross-genomic proximity linkage signals by constructing empirical noise models or statistical background distributions. Commonly used indicators include three categories: standard score (Z-score), noise-to-signal ratio, and empirical distribution threshold.
[0065] These methods can remove non-specific connections to some extent, but their model assumptions are often out of touch with the physical mechanisms of actual extracellular random connections, and they are highly sensitive to sequencing depth, making it difficult to generalize their application across different datasets.
[0066] Although the above methods have alleviated the noise problem in DNA proximity-ligation sequencing data to some extent, they still generally have the following shortcomings:
[0067] First, it does not explicitly model the probability of random spatial collisions and connections between different genomes based on the generation mechanism of extracellular nonspecific connections;
[0068] Second, it mainly relies on the absolute number of read segments or the relative connection ratio as filtering criteria, lacking a unified, interpretable and cross-dataset comparable discrimination scale;
[0069] Third, the impact of differences in genomic abundance on random collision probability and the cumulative effect of noise connectivity was not systematically considered;
[0070] Fourth, the lack of an accurate signal-noise differentiation mechanism leads to a large number of false negative results while removing a large number of false positive results.
[0071] This application proposes a method for filtering noisy connections between different genomes in sample DNA spatial proximity linkage data, used to determine intracellular or extracellular connection signals between different genomes. This method constructs a random collision background model based on statistical results from DNA proximity linkage sequencing data, and systematically screens cross-genome connections accordingly.
[0072] The core idea of this application is that, given DNA-protein spatial cross-linking during sample processing and the possibility of incompletely inactivated cross-linking agents remaining in the reaction system after cell disruption, the probability of a connection between the extracellular DNA of any two genomes should be proportional to their relative abundance in the sample, consistent with a random spatial collision model. Therefore, the expected number of connections under random collision conditions can be derived based on the relative abundance of each genome, and the ratio of the observed number of connections to this expected value can be used as a unified criterion to filter out noisy connections while preserving true spatially adjacent connections.
[0073] The following description, with reference to the accompanying drawings, illustrates a method, apparatus, terminal, and medium for determining proximity links between different genomes according to embodiments of this application. Addressing the problem mentioned in the background art where false positive links significantly interfere with the determination of proximity relationships between different genomes in a sample, this application provides a method for determining proximity links between different genomes. In this method, a target genome set corresponding to a multi-genome mixed DNA sample is obtained, and the relative abundance of each genome in the target genome set and the total number of proximity links in the target genome set are determined. Based on the relative abundance of each genome and the total number of proximity links, the expected number of proximity links between any pair of genomes under completely random spatial collision conditions is estimated. The actual number of proximity links for each genome pair is obtained, and based on the actual number of proximity links and the expected number of proximity links for each genome pair, the determination result of the proximity link signal generated between each genome pair is determined. This application derives the expected number of proximity links under random collision conditions by using the relative abundance of each genome, filters out noisy links by comparing the expected number of proximity links with the actual number of proximity links, and retains true spatial proximity links, thereby avoiding the problem of false positive links significantly interfering with the determination of proximity relationships between different genomes in a sample.
[0074] Please see Figure 1 , Figure 1 This is a flowchart illustrating the method for determining proximity connectivity signals between different genomes in this application. Figure 1 As shown in the embodiments of this application, the method for determining proximity connectivity signals between different genomes includes:
[0075] Step S100: Obtain the target genome set corresponding to the multi-genome mixed DNA sample, and determine the relative abundance of each genome in the target genome set and the total number of neighboring connections in the target genome set.
[0076] In this context, a multi-genome mixed DNA sample refers to the entire object sampled and processed in the experimental workflow (e.g., Hi-C experimental materials that have undergone cross-linking, library construction, and sequencing). The target genome set corresponding to the multi-genome mixed DNA sample refers to the set of target genomes that actually participated in the statistical analysis of proximity joins and the construction of the contact matrix after alignment and quality filtering of sequencing data obtained using DNA spatial proximity joining technology. The random collision model is established within this genome set because only these genomes contribute to the observed spatial proximity joins and can be used for a reasonable estimation of the expected number of proximity joins. DNA spatial proximity joining technologies include, but are not limited to, Hi-C, Micro-C, Pore-C, SPRITE, GAM, and other experimental or sequencing technologies capable of characterizing the spatial proximity relationships of DNA fragments. Micro-C is a chromatin conformation capture technology based on micrococcal nuclease (MNase) digestion that achieves nucleosome-level resolution; Pore-C combines chromatin conformation capture with nanopore long-read sequencing to achieve direct detection of multiple chromatin spatial interactions; SPRITE (Split-Pool Recognition of Interactions by Tag Extension) uses a barcode co-labeling strategy to resolve multiple chromatin interactions without the need for ligation reactions; GAM (Genome Architecture Mapping) constructs a three-dimensional structure association map at the whole genome scale based on ultrathin nuclear slices and DNA co-occurrence statistics.
[0077] Multi-genome mixed DNA samples include, but are not limited to, metagenomic samples, co-culture samples, multi-strain mixed samples, or other samples containing DNA from multiple different genomic sources.
[0078] The relative abundance of a genome refers to the relative proportion of a particular genome within a target genome set. This relative abundance is typically obtained by quantitatively analyzing the sequencing signals of different genomes in a mixed sample, thus characterizing the relative abundance level of that genome within the mixed sample.
[0079] The total number of neighboring connections in the target genome set refers to the sum of all valid DNA neighboring connection events originating from the target genome set in DNA neighboring connection sequencing technology.
[0080] like Figure 1 As shown in the embodiments of this application, the method for determining proximity connectivity signals between different genomes further includes:
[0081] Step S200: Estimate the expected number of neighbor connections for any pair of genomes under completely random spatial collision conditions based on the relative abundance of each genome and the total number of neighbor connections.
[0082] In multi-genome mixed DNA samples, the probability of spatial proximity connections between DNA fragments from different genome sources differs fundamentally between real intracellular spatial proximity and random extracellular collision conditions. In this application, the estimation of the expected number of proximity connections in the target genome set under completely random spatial collision conditions is achieved based on a random collision model of relative genome abundance.
[0083] In one embodiment of this application, step S200 specifically includes:
[0084] Step S210: Calculate the probability of random spatial collision between any two genomes based on the relative abundance of each genome;
[0085] Step S220: Calculate the product of the random spatial collision probability and the total number of neighbor connections to obtain the expected number of neighbor connections for any genome pair under completely random spatial collision conditions.
[0086] This application derives the random space collision probability based on the theory that the probability of a proximity connection between any two genomes in the target genome set is proportional to the product of their relative abundance in the target genome set, and realizes random space collision probability modeling based on abundance.
[0087] In the first embodiment of this application, step S210 specifically includes:
[0088] Step S210a: Based on the relative abundance of each genome, calculate the product of the relative abundance of any pair of genomes in the target genome set to obtain the product of the relative abundance of each pair of genomes in the target genome set.
[0089] Step S220a: Sum the products of the relative abundance of all genome pairs to obtain the summation result;
[0090] Step S230a: The ratio of the product of the relative abundances to the summation result is used as the probability of random spatial collision between the two genomes.
[0091] The formula for the random collision model in this embodiment is:
[0092] ;
[0093] in, Indicates the expected number of neighbor connections. Indicates the relative abundance of genome i. Indicates the relative abundance of genome j. This represents the total number of neighbor connections. This represents all possible genome pairs in the target genome set. Indicates the relative abundance of genome k. Indicates the relative abundance of genome l. This represents the sum of the products of the relative abundance of all possible genome pairs in the target genome set, used for normalization so that the sum of all expected connectivity numbers equals the observed total connectivity number. .
[0094] Based on the construction probability mechanism of different genome connections during sample processing, this application uses DNA proximity connection sequencing data to accurately determine the intracellular or extracellular connection relationship between two genomes, achieving high-precision determination with low false positive and low false negative rates.
[0095] In the second embodiment of this application, step S210 specifically includes:
[0096] Step S210b: Determine the weight correction factor for each genome based on the preset weight correction information;
[0097] Step S220b: Calculate the relative abundance weighted value of each genome based on the relative abundance and weight correction factor of each genome;
[0098] Step S230b: Calculate the product of the relative abundance weights of any pair of genomes in the target genome set to obtain the product of the relative abundance weights of each pair of genomes in the target genome set;
[0099] Step S240b: Sum the products of the relative abundance weights of all genome pairs to obtain the summation result;
[0100] Step S250b: The ratio of the product of the relative abundance weighted values to the summation result is used as the probability of random spatial collision between the two genomes.
[0101] The formula for the random collision model in this embodiment is:
[0102] ;
[0103] in, Indicates the expected number of neighbor connections. Indicates the relative abundance of genome i. This represents the weight correction factor for genome i. Indicates the relative abundance of genome j. This represents the weight correction factor for genome j. This represents the total number of neighbor connections. This represents all possible genome pairs in the target genome set. Indicates the relative abundance of genome k. This represents the weighting correction factor for genome k. Indicates the relative abundance of genome l. The weighting correction factor for genome l. This represents the sum of the relative abundance weights of all possible genome pairs in the target genome set, used for normalization so that the sum of all expected connections equals the observed total connections. .
[0104] The weighting correction information includes at least one of GC content, restriction enzyme site density, and adjacent linker fragment recovery efficiency. That is, one, a combination of two, or a combination of three of GC content, restriction enzyme site density, and adjacent linker fragment recovery efficiency can be introduced into the relative abundance term.
[0105] Specifically, GC content and restriction enzyme site density are inherent properties of the genome sequence itself; each genome has corresponding GC content and restriction enzyme site distribution characteristics. Neighbor-linked fragment recovery efficiency, however, is not an inherent property of the genome but rather a statistical or empirical parameter related to sample processing, experimental procedures, and sequencing processes. The neighbor-linked fragment recovery efficiency in this application can be estimated separately for different genomes, or a uniform parameter can be used when differentiation is not possible.
[0106] This application characterizes the extracellular nonspecific join mechanism based on a random collision model of relative genomic abundance. In existing technologies, the mainstream methods for filtering noisy joins generated in DNA proximity joining techniques in multi-genome mixed samples mainly include: methods based on a minimum inter-genome join filtering threshold, methods based on viral connectivity and preference for hosts, and methods based on statistical or empirical noise models. These methods do not start from the physical generation mechanism of extracellular nonspecific joins to model the probability of random spatial collisions and joins between different genomes, but mainly rely on absolute read count thresholds, relative connectivity ratios, or data-driven statistical significance criteria. This application, based on the real mechanism of extracellular joins generated in multi-microbial mixed samples, including microbiomes and environmental samples, during DNA proximity joining processing, introduces for the first time a random collision model based on relative genomic abundance. It assumes that in the absence of real biological spatial proximity or infection relationships, the probability of proximity joins between different genome pairs is proportional to the product of their relative abundance in the sample, thus constructing the desired number of proximity joins for each genome pair. The random collision model in this application characterizes the source of extracellular random connections at the mechanistic level, so that noisy connections are no longer judged by "more or less" but by "whether they significantly exceed the expectation of random collisions", providing an interpretable physical basis for subsequent fine-grained filtering.
[0107] Based on the construction probability mechanism of different genome connections during sample processing, this application uses DNA proximity connection sequencing data to accurately determine the intracellular or extracellular connection relationship between two genomes, achieving high-precision determination with low false positive and low false negative rates.
[0108] like Figure 1 As shown in the embodiments of this application, the method for determining proximity connectivity signals between different genomes further includes:
[0109] Step S300: Obtain the actual number of neighbor connections for each genome pair, and determine the determination result of the neighbor connection signal generated between each genome pair based on the actual number of neighbor connections and the expected number of neighbor connections for each genome pair.
[0110] This application uses the enrichment of observed values relative to the random background as a filtering criterion, that is, to determine whether the observed cross-genomic proximity connections are significantly higher than the random spatial collision background.
[0111] In this embodiment of the application, step S300 specifically includes:
[0112] Step S310: Count the number of actual neighbor connections observed for each genome pair;
[0113] Step S320: Based on the actual number of neighboring connections and the expected number of neighboring connections for each genome pair, obtain a predetermined index statistic;
[0114] Step S330: Obtain a pre-defined threshold, compare the predetermined indicator statistics of each genome pair with the threshold, and obtain the comparison result;
[0115] Step S340: Based on the comparison results, determine the neighboring connection signal generated between the two genomes in each genome pair.
[0116] The predetermined indicator statistics include at least one of the following: ratio-based indicator statistics, normalized difference-based indicator statistics, log-ratio-based indicator statistics, and standardized indicator statistics. The calculation formula for the predetermined indicator statistics is a unified scoring function reflecting the degree of connection enrichment. Based on the distribution characteristics of the scoring function, simulated sample calibration results or data-driven optimization criteria are used to set a threshold for distinguishing between real spatially adjacent connections and random noise connections, and based on this, cross-genomic connections are systematically filtered.
[0117] For example, if the predetermined indicator statistic For ratio-based statistics, the expected statistic is the ratio of the actual number of neighbor connections to the expected number of neighbor connections (Observed-to-Expected), expressed as:
[0118] ;
[0119] in, This represents the number of observed actual neighbor connections. Indicates the expected number of neighbor connections. It is used to measure whether the proximity connections between a pair of genomes are significantly higher than the level that can be explained by random collisions.
[0120] If the predetermined indicator statistics For a normalized difference statistic, the predetermined statistic is the normalized difference between the actual number of nearest neighbors and the expected number of nearest neighbors, expressed as:
[0121] .
[0122] If the predetermined indicator statistics If the statistic is a log-ratio, then the predetermined statistic is the log-ratio between the actual number of nearest neighbors and the expected number of nearest neighbors, expressed as:
[0123] .
[0124] If the predetermined indicator statistics To standardize the statistic, the predetermined statistic is the standard value between the actual number of neighbor connections and the expected number of neighbor connections, expressed as:
[0125] ;
[0126] in, This represents the variance estimated based on the random collision model.
[0127] In addition, empirical quantiles or ranking indicators can be used as standardized statistical measures, for example, based on the percentile position of a predetermined indicator statistic in all transgenomic connections.
[0128] This application can also use a nonlinear monotonic transformation form. Calculate the statistical measures of the predetermined indicators ,For example .
[0129] Unlike existing technologies that directly use absolute read count thresholds, virus-host connectivity ratios, standard scores (Z-scores), or noise-to-signal ratios, this application uses predetermined statistical indicators as a unified and interpretable filtering scale to make data from different abundance levels and sequencing depths comparable.
[0130] This application effectively eliminates the accumulation of random connections dominated by high-abundance genomes and non-specific connections caused by extracellular spatial proximity or mixing by thresholding the distribution of predetermined indicator statistics. At the same time, it retains connection signals that significantly exceed the expectations of random collisions and effective connections that are more likely to reflect in situ infection or stable spatial proximity relationships between different genomes.
[0131] In one embodiment of this application, step S340 specifically includes:
[0132] If the alignment result is a ratio less than the threshold, then the proximity connection signal generated between the two genomes in the genome pair is an extracellular random connection signal;
[0133] If the alignment result is a ratio greater than or equal to the threshold, then the proximity connection signal generated between the two genomes in the genome pair is the intracellular real spatial proximity signal.
[0134] The threshold is obtained by evaluating the distribution characteristics of the ratio of actual neighbor connections to expected neighbor connections based on known real genomes in several constructed simulated community datasets. The threshold includes at least one of a fixed empirical threshold, a quantile threshold, and a sample adaptive threshold.
[0135] Specifically, based on multiple constructed mock community datasets, the distribution characteristics of predetermined indicator statistics are systematically evaluated to characterize the background distribution range of predetermined indicator statistics under conditions of only random spatial collisions and extracellular nonspecific connections. For example, a fixed threshold (e.g., 30) is set based on the results of simulated community calibration. The sample adaptive threshold is automatically determined by preset rules or algorithms based on the statistical distribution characteristics of predetermined indicator statistics in the current sample to be analyzed or simulated data, and is used to adapt to differences in sequencing depth, genome composition, and abundance structure among different samples.
[0136] This application calculates predetermined index statistics (such as...) for different genome pairs under simulated community conditions. The ratio is used to observe its range under conditions of random spatial collisions and extracellular nonspecific connections, thereby obtaining an empirical threshold that can be used for practical data analysis. In subsequent processing of actual sample data, the statistical value of the predetermined indicator corresponding to each pair of genomes is compared with the threshold, and the neighbor connection type is determined and filtered accordingly.
[0137] For example, an empirical threshold is determined to distinguish between real spatial proximity connections and random noise connections. , making when When the corresponding proximity connections mainly originate from random collisions or non-specific spatial proximity; when At this point, the corresponding connection significantly exceeds the expectation of the random collision background, and is therefore judged as a high-confidence valid connection and retained. Threshold This can be achieved by using real known genomes in simulated communities. The degree of separation of the distributions is empirically calibrated. For example, a threshold. A suitable ratio is the ratio that eliminates the vast majority of random noise connections while still stably preserving real spatial proximity connections. Lower limit.
[0138] This application uses a random collision model based on the relative abundance of the genome to systematically filter non-specific noise connections in DNA proximity sequencing data. This significantly reduces the systematic false positive proximity connection signals introduced by extracellular non-specific connections, effectively suppresses false cross-genome connections generated by the accumulation of random collisions in high-abundance genomes, and improves the accuracy and stability of identifying connections between genomes.
[0139] In another embodiment of this application, step S300 specifically includes:
[0140] Count the number of actual neighbor connections observed for each genome pair;
[0141] Based on the actual number of neighbor connections and the expected number of neighbor connections for each genome pair, a predetermined index statistic is obtained;
[0142] A pre-constructed statistical model is obtained, and the predetermined index statistics are input into the statistical model to obtain the determination result of the proximity connection signal generated between each genome pair.
[0143] Specifically, this application uses predetermined indicator statistics as one of the input features, introduces a statistical model or probabilistic model to determine whether cross-genomic proximity connections are true spatial proximity relationships, and outputs the corresponding confidence level or probability value. The statistical model in this application includes, but is not limited to, the following forms: empirical distribution-based probabilistic models and thresholded probabilistic models. Empirical distribution-based probabilistic models utilize predetermined indicator statistics (such as...) from simulated community data or background data. The empirical distribution of the data is used to calculate the probability that cross-genomic proximity connections fall into the tail of a random background distribution. Thresholding probability models compare the confidence or probability value output by the statistical model with a preset threshold or a threshold determined by a data-driven approach, thereby achieving connection screening that is technically equivalent to fixed-threshold filtering.
[0144] This application significantly reduces the systematic false positive neighbor linkage signals introduced by extracellular nonspecific linkages through a multi-level filtering and random collision modeling process. It effectively suppresses false cross-genome linkages caused by the accumulation of random collisions in high-abundance genomes, and improves the accuracy and stability of identifying linkage relationships between genomes. Without changing the original DNA neighbor linkage experimental procedure or adding additional experimental steps, it realizes automated and robust filtering and analysis of DNA neighbor linkage sequencing data.
[0145] This application starts with the relative abundance of different genomes in the same sample to construct a probabilistic model to characterize the background of random spatial collisions. Based on this random spatial collision model, the expected number of proximity connections between any two genomes is derived. A predetermined statistical index between the observed number of connections and the expected number of connections is used as a unified and scalable filtering criterion to distinguish between real intracellular associations and noisy connections caused by random spatial proximity or extracellular DNA. In other words, this application models the probability of random spatial collisions based on the relative abundance or comparable characteristics of different genomes in a sample, and uses the observed-expected connection ratio or a statistically equivalent one to screen and filter the DNA spatial proximity connection results. This application does not limit the specific experimental platform name (such as Hi-C, Meta-HiC, 3C, Pore-C, or other DNA spatial proximity connection technologies), implementation tools, parameter setting methods, threshold forms, or statistical expression methods. Meta-HiC refers to metagenomic Hi-C technology.
[0146] In one embodiment, such as Figure 2 As shown, based on the above-mentioned method for determining proximity connections between different genomes, this application also provides a device for determining proximity connections between different genomes, comprising:
[0147] The acquisition module 100 is used to acquire the target genome set corresponding to the multi-genome mixed DNA sample, and to determine the relative abundance of each genome in the target genome set and the total number of neighboring connections of the target genome set;
[0148] The estimation module 200 is used to estimate the expected number of neighbor connections for any pair of genomes under completely random spatial collision conditions based on the relative abundance of each genome and the total number of neighbor connections.
[0149] The determination module 300 is used to obtain the actual number of neighbor connections for each genome pair, and to determine the determination result of the neighbor connection signal generated between each genome pair based on the actual number of neighbor connections and the expected number of neighbor connections for each genome pair.
[0150] It should be noted that the explanation of the aforementioned method for determining proximity connection signals between different genomes also applies to the device for determining proximity connection signals between different genomes in this embodiment, and will not be repeated here.
[0151] This application discloses a device for determining proximity connections between different genomes. It acquires a target genome set corresponding to a multi-genome mixed DNA sample and determines the relative abundance of each genome in the target genome set and the total number of proximity connections. Based on the relative abundance of each genome and the total number of proximity connections, it estimates the expected number of proximity connections between any pair of genomes under completely random spatial collision conditions. It then obtains the actual number of proximity connections for each genome pair and, based on the actual and expected number of proximity connections for each genome pair, determines the determination result of the proximity connection signal generated between each genome pair. This application derives the expected number of proximity connections under random collision conditions by using the relative abundance of each genome, and filters out noisy connections by comparing the expected and actual number of proximity connections, retaining true spatial proximity connections. This avoids the problem of false positive connections significantly interfering with the determination of proximity relationships between different genomes in a sample.
[0152] Figure 3 A schematic diagram of the structure of a terminal provided in an embodiment of this application. The terminal may include:
[0153] The memory 501, the processor 502, and the computer program stored on the memory 501 and capable of running on the processor 502.
[0154] When the processor 502 executes the program, it implements the method for determining the proximity connection signal between different genomes provided in the above embodiments.
[0155] Furthermore, the terminal also includes:
[0156] Communication interface 503 is used for communication between memory 501 and processor 502.
[0157] The memory 501 is used to store computer programs that can run on the processor 502.
[0158] The memory 501 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0159] If the memory 501, processor 502, and communication interface 503 are implemented independently, they can be interconnected via a bus to communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one line is used in the diagram, but this does not imply that there is only one bus or one type of bus.
[0160] Optionally, in a specific implementation, if the memory 501, processor 502, and communication interface 503 are integrated on a single chip, then the memory 501, processor 502, and communication interface 503 can communicate with each other through an internal interface.
[0161] Processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0162] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for determining proximity connection signals generated between different genomes.
[0163] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0164] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0165] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0166] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can read and execute instructions from or in conjunction with such an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). In addition, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically by optically scanning paper or other media, then editing, interpreting or otherwise processing them as necessary, and then storing them in computer memory.
[0167] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0168] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware, and the program can be stored in a computer-readable storage medium. When executed, the program includes one or a combination of the steps of the method embodiments.
[0169] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0170] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A method for determining proximity connection signals generated between different genomes, characterized in that, The method includes: Obtain the target genome set corresponding to the multi-genome mixed DNA sample, and determine the relative abundance of each genome in the target genome set and the total number of neighbor connections in the target genome set; The expected number of neighbor connections for any pair of genomes under completely random spatial collision conditions is estimated based on the relative abundance of each genome and the total number of neighbor connections. Obtain the actual number of neighbor connections for each genome pair, and determine the determination result of the neighbor connection signal generated between each genome pair based on the actual number of neighbor connections and the expected number of neighbor connections for each genome pair; Estimate the expected number of neighbor connections for any pair of genomes under completely random spatial collision conditions based on the relative abundance of each genome and the total number of neighbor connections, including: The probability of random spatial collision between any two genomes is calculated based on the relative abundance of each genome. The product of the random spatial collision probability and the total number of neighbor connections is calculated to obtain the expected number of neighbor connections for any genome pair under the condition of completely random spatial collision. The probability of random spatial collision between any two genomes is calculated based on the relative abundance of each genome, including: The weight correction factor for each genome is determined based on the preset weight correction information; The relative abundance weighted value of each genome is calculated based on the relative abundance of each genome and the weight correction factor. Calculate the product of the relative abundance weights of any pair of genomes in the target genome set to obtain the product of the relative abundance weights of each pair of genomes in the target genome set; Summing the products of the relative abundance weights of all genome pairs yields the summation result; The ratio of the product of the relative abundance weighted values to the summation result is used as the probability of random spatial collision between the two genomes; The weight correction information includes at least one of the following: GC content, restriction site density, and adjacent linker fragment recovery efficiency.
2. The method for determining proximity connection signals between different genomes according to claim 1, characterized in that, The probability of random spatial collision between any two genomes is calculated based on the relative abundance of each genome, including: Based on the relative abundance of each genome, the product of the relative abundance of any pair of genomes in the target genome set is calculated to obtain the product of the relative abundance of each pair of genomes in the target genome set. Summing the products of the relative abundance of all genome pairs yields the summation result; The ratio of the product of the relative abundances to the summation result is used as the probability of random spatial collision between the two genomes.
3. The method for determining proximity connection signals between different genomes according to claim 1, characterized in that, Obtain the actual number of neighbor connections for each genome pair. Based on the actual and expected number of neighbor connections for each genome pair, determine the outcome of the neighbor connection signal generated between each genome pair, including: Count the number of actual neighbor connections observed for each genome pair; Based on the actual number of neighbor connections and the expected number of neighbor connections for each genome pair, a predetermined index statistic is obtained; Obtain a pre-defined threshold, and compare the predetermined indicator statistics of each genome pair with the threshold to obtain the comparison result; Based on the comparison results, the determination of the proximity connection signal generated between the two genomes in each genome pair is obtained.
4. The method for determining proximity connection signals between different genomes according to claim 3, characterized in that, Based on the alignment results, the determination of the proximity connection signal between the two genomes in each genome pair is obtained, including: If the alignment result is a ratio less than the threshold, then the proximity connection signal generated between the two genomes in the genome pair is an extracellular random connection signal; If the alignment result is a ratio greater than or equal to the threshold, then the proximity connection signal generated between the two genomes in the genome pair is the intracellular real spatial proximity signal; The threshold is obtained by evaluating the distribution characteristics of the ratio of actual neighbor connections to expected neighbor connections based on known real genomes in several constructed simulated community datasets. The threshold includes at least one of a fixed empirical threshold, a quantile threshold, and a sample adaptive threshold.
5. The method for determining proximity connection signals between different genomes according to claim 1, characterized in that, Obtain the actual number of neighbor connections for each genome pair. Based on the actual and expected number of neighbor connections for each genome pair, determine the outcome of the neighbor connection signal generated between each genome pair, including: Count the number of actual neighbor connections observed for each genome pair; Based on the actual number of neighbor connections and the expected number of neighbor connections for each genome pair, a predetermined index statistic is obtained; A pre-constructed statistical model is obtained, and the predetermined index statistics are input into the statistical model to obtain the determination result of the proximity connection signal generated between each genome pair.
6. A device for determining proximity connection signals generated between different genomes, characterized in that, The device includes: The acquisition module is used to acquire the target genome set corresponding to the multi-genome mixed DNA sample, and to determine the relative abundance of each genome in the target genome set and the total number of neighboring connections of the target genome set; An estimation module is used to estimate the expected number of neighbor connections for any pair of genomes under completely random spatial collision conditions based on the relative abundance of each genome and the total number of neighbor connections. The determination module is used to obtain the actual number of neighbor connections for each genome pair, and based on the actual number of neighbor connections and the expected number of neighbor connections for each genome pair, to determine the determination result of the neighbor connection signal generated between each genome pair; Estimate the expected number of neighbor connections for any pair of genomes under completely random spatial collision conditions based on the relative abundance of each genome and the total number of neighbor connections, including: The probability of random spatial collision between any two genomes is calculated based on the relative abundance of each genome. The product of the random spatial collision probability and the total number of neighbor connections is calculated to obtain the expected number of neighbor connections for any genome pair under the condition of completely random spatial collision. The probability of random spatial collision between any two genomes is calculated based on the relative abundance of each genome, including: The weight correction factor for each genome is determined based on the preset weight correction information; The relative abundance weighted value of each genome is calculated based on the relative abundance of each genome and the weight correction factor. Calculate the product of the relative abundance weights of any pair of genomes in the target genome set to obtain the product of the relative abundance weights of each pair of genomes in the target genome set; Summing the products of the relative abundance weights of all genome pairs yields the summation result; The ratio of the product of the relative abundance weighted values to the summation result is used as the probability of random spatial collision between the two genomes; The weight correction information includes at least one of the following: GC content, restriction site density, and adjacent linker fragment recovery efficiency.
7. A terminal, characterized in that, Includes: The memory, the processor, and the proximity connection signal determination program between different genomes stored in the memory and executable on the processor, wherein when the proximity connection signal determination program between different genomes is executed by the processor, the steps of the proximity connection signal determination method between different genomes as described in any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed to implement the steps of the method for determining proximity connection signals generated between different genomes as described in any one of claims 1 to 5.