Tissue origin inference method and device based on cancer specific chromatin accessibility marker

By using low-cost sWGS technology and multi-dimensional accessibility feature screening, combined with machine learning models, the sensitivity and accuracy issues of cancer tissue tracing in existing technologies have been resolved, achieving efficient and accurate identification of cancer-specific biomarkers and determination of tissue origin.

CN121999877APending Publication Date: 2026-05-08GENESEEQ TECH INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GENESEEQ TECH INC
Filing Date
2025-11-07
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing methods for tracing the origin of cancer tissue mainly rely on copy number, mutation, and methylation signals, which are limited in sensitivity and cost. Fragmentomics is susceptible to sequencing bias, and background modeling and reproducibility are insufficient, making it difficult to achieve efficient and accurate determination and monitoring of tissue origin.

Method used

We employ low-cost shallow whole-genome sequencing (sWGS) technology, characterize chromatin accessibility through EDI (Extended Dispersion Index), a composite index of fragment endpoint dispersion and coverage fluctuation, and screen robust candidate regions by combining global and local significance tests. We construct multi-dimensional accessibility features and use machine learning models for cancer tissue tracing, thereby reducing costs and improving accuracy.

Benefits of technology

It enables efficient screening of cancer-specific chromatin accessibility markers under low-cost conditions, reduces sequencing bias, and improves the discriminative power and accuracy of tissue tracing. It is applicable to clinical scenarios such as early cancer screening, identification of unknown primary lesions, efficacy evaluation, and recurrence warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121999877A_ABST
    Figure CN121999877A_ABST
Patent Text Reader

Abstract

The invention discloses a cancer-specific chromatin accessibility marker-based tissue origin inference method and device and a storage medium, and belongs to the technical field of gene detection. The method aims to solve the problem that low-cost shallow whole genome sequencing data is difficult to carry out accurate cancer traceability. The invention provides a fragment discreteness index which is combined with terminal dispersity and coverage fluctuation of free DNA fragments so as to more accurately characterize chromatin accessibility. And through a global and local combined statistical test strategy, identifying candidate accessibility regions from the data. The method comprises the core step of screening out a unique marker of a specific cancer species by removing a common accessibility region in various cancers and healthy control. Based on the specific markers, multi-dimensional fragment omics features are extracted, a machine learning multi-classification model is constructed, and the probability that a to-be-detected sample comes from different cancer types is predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method, apparatus, and computer storage medium for identifying cancer-specific chromatin accessibility markers based on sWGS cfDNA and for inferring tissue origin, belonging to the field of gene detection technology. Background Technology

[0002] Chromatin accessibility refers to the degree to which DNA can enter and bind to transcription factors, RNA polymerases, and various regulatory complexes in the context of three-dimensional folding of cells. Its essence is determined by nucleosome occupancy and localization, histone modification, and the dynamics of remodeling factors: open regions (such as promoters and enhancers) have sparse nucleosomes and are easily accessible to proteins, while closed regions have dense nucleosomes and restricted binding.

[0003] In recent years, this concept has been successfully applied to the field of plasma cfDNA research. Because nucleosome protection and regulatory protein binding leave "footprints" in fragmentation patterns, cfDNA retains open chromatin information from its source tissue. Researchers utilize the physical principle that "open regions are more easily cleaved by endonucleases, producing shorter fragments with more dispersed endpoints," combining fragment length distribution, endpoint motifs, and nucleosome footprints to non-invasively reconstruct tissue-specific accessibility maps. This is further supplemented by histone modification studies such as cfChIP-seq to characterize regulatory activity at the peripheral blood level. Extensive evidence shows that active chromatin fragments in plasma are highly correlated with promoters / enhancers, RNAPol II stop sites, and gene expression levels in the ENCODE map, distinguishing multiple tissue open lineages and indicating tumor or organ origin. Footprint analysis based on nucleosome occupancy / spacing has also been used for disease subtyping and lesion identification, and has been validated in real-world clinical scenarios. Compared with signals such as mutation, copy number, and methylation, accessibility biomarkers combine functionality (reflecting gene regulatory activity), specificity (sensitive to open lineages of tissues / organs), and real-time availability (suitable for dynamic monitoring), demonstrating higher explanatory power and application potential in tissue tracing, early screening, and efficacy / recurrence assessment.

[0004] Tissue origin tracing (TOO) of cancer is a crucial step in early screening, diagnosis, and determination of the source of metastasis. Current TOO methods mainly rely on three types of signals: copy number / mutation, methylation, and fragmentomics. The first two are limited in sensitivity or cost under low tumor burden, while fragmentomics is easily affected by coverage fluctuations and sequencing bias, and background modeling and reproducibility still need to be improved. Summary of the Invention

[0005] This patent proposes a low-cost, engineered workflow for shallow WGS (Wide Genomics Syndrome). It characterizes the open state using a comprehensive index of "endpoint dispersion × coverage fluctuation," filters robust candidate regions through global and local statistical tests, and merges neighboring windows to generate accessibility features directly usable for learning and modeling. Subsequently, the relevant open signals are integrated with tissue origination models to achieve multi-cancer origin determination and clinical monitoring. This approach reduces reliance on high-depth directional capture, systematically suppresses sequencing and spatial bias, and improves signal stability and statistical reliability. It also possesses good scalability and applicability, suitable for clinical scenarios such as early cancer screening, identification of unknown primary lesions, efficacy evaluation, and recurrence warning, combining novelty and practical value. This method constructs a composite index EDI = EDI × std(coverage) based on fragment endpoint dispersion and coverage fluctuation. A sliding window scan of the entire genome is performed, incorporating coverage, dark areas, and comparability filtering. The normalized FDI is beta-fitted to calculate global significance and corrected using BH. A local normality test is conducted in the ±5kb neighborhood, retaining only doubly significant windows and merging adjacent regions as candidate regions. Cancer-specific candidate regions are screened by removing common regions across cancer types / tissues. A multi-dimensional accessibility feature set is extracted to construct a machine learning model for cancer tissue tracing, reducing costs and improving accuracy and feasibility.

[0006] A method for inferring tissue origin based on cancer-specific chromatin accessibility markers, comprising the following steps:

[0007] a) Provide cell-free DNA sequencing data of a sample to be tested, which has been processed by removing adapters and low-quality sequences, aligning to a reference genome, and removing repetitive sequences;

[0008] b) For a pre-determined set of cancer types, obtain a set of cancer-specific chromatin accessibility markers for each cancer type;

[0009] c) For each cancer-specific chromatin accessibility marker obtained in step b), extract at least one fragment omics feature value for the test sample across all genomic regions defined by the cancer-specific chromatin accessibility marker obtained in step b), and summarize and standardize the multiple feature values ​​extracted for each cancer type to form a multidimensional feature vector representing the test sample.

[0010] d) Input the feature vector obtained in step c) into a pre-trained multi-class machine learning model, output the probability score of the test sample corresponding to each of the cancer types in the group, and infer its tissue origin based on the probability score.

[0011] The cell-free DNA sequencing data were obtained from shallow whole-genome sequencing (sWGS) of plasma samples with a single-sample sequencing depth of 5X.

[0012] The group of cancer types includes breast cancer, cervical cancer, colorectal cancer, esophageal cancer, stomach cancer, head and neck cancer, kidney cancer, liver and gallbladder cancer, lung cancer, melanoma, ovarian cancer, pancreatic cancer, prostate cancer, sarcoma, thyroid cancer, urothelial carcinoma, and endometrial cancer.

[0013] The method for obtaining cancer-specific chromatin accessibility markers in step b) includes:

[0014] i. For each type of cancer and healthy control group, a set of candidate chromatin accessibility regions are identified by analyzing their respective cell-free DNA sequencing data; the cell-free DNA sequencing data is obtained by merging single sample data of multiple diseases to obtain high-depth representative sample data with a depth of not less than 100X.

[0015] ii. For the target cancer type, the candidate chromatin accessibility regions of all other non-target cancer types and the healthy control group are merged to form a background set, and adjacent regions of this background set with a distance of less than 20 bp on the genome are merged and deduplicated.

[0016] iii. From the candidate chromatin accessibility regions of the target cancer type, regions with an overlap ratio exceeding 0.5 with the background set are removed. The overlap ratio is defined as the total length of the bases overlapping between the target candidate region and the background set divided by the base length of the target candidate region itself. The remaining regions are defined as cancer-specific chromatin accessibility markers for the target cancer type.

[0017] The identification of candidate chromatin accessibility regions in step i) is achieved by scanning the genome using a sliding window with a length of 200 bp and a step size of 20 bp, and calculating the fragment dispersion index (FDI) for windows that meet the basic quality filtering criteria. The basic quality filtering criteria are that the sequencing coverage of the window is greater than 0, it does not belong to the dark area composed of genomic repetitive sequences or low-complexity regions, and the average comparability score of all bases in the window is not less than 0.9.

[0018] The fragment dispersion index FDI is defined as the product of the fragment end dispersion index EDI and the fragment coverage standard deviation std(coverage), i.e., FDI = EDI × std(coverage).

[0019] The fragment end dispersion index (EDI) is used to measure the spatial dispersion of free DNA fragment ends within a region, and it is expressed by the formula... Perform the calculation, where:

[0020] n is the total number of free DNA fragment ends in this region;

[0021] j is the last index in the region from 1 to n;

[0022] w is the preset local cell width, with a value of 20bp;

[0023] x is a preset constant with a value of 0.5;

[0024] C j,x×w Let w be the total number of ends in a subinterval centered at the j-th end and with a width of w.

[0025] The identification of candidate chromatin accessibility regions further includes a statistical test step that combines global and local significance tests to screen for windows with significantly high FDI values.

[0026] The global significance test in the statistical testing step includes:

[0027] a. Normalize the FDI values ​​of all windows that pass through the basic quality filter to obtain FDI_norm;

[0028] b. At a single chromosome scale, fit the FDI_norm values ​​of all windows to a Beta distribution and calculate the right-tailed cumulative probability of the FDI_norm value for each window as the global significance p_global.

[0029] c. The Benjamini-Hochberg method was used to perform multiple tests to correct all p_global values ​​to control the false discovery rate (FDR), and windows that simultaneously satisfy p_global < 0.05 and FDR < 0.05 were retained as global significant candidate windows.

[0030] The statistical testing step further includes a local significance test, which performs the following operation on each globally significant candidate window:

[0031] d. Using all windows within a 5kb range upstream and downstream of the center of the candidate window as the local background, approximate the FDI value of the local background with a normal distribution, and calculate the mean μ and standard deviation σ.

[0032] e. Calculate the right-tail local saliency p_ of the candidate window relative to the local background. local = 1 - Φ((FDI - μ) / σ), where Φ is the cumulative distribution function of the standard normal distribution;

[0033] f. Retain windows with p_local < 0.05.

[0034] The statistical testing step concludes with a region integration step, which merges windows that pass both global and local significance tests and have an adjacency distance of less than or equal to 200 bp on the genome to form the final candidate chromatin accessibility regions.

[0035] The fragment omics features in step c) include the fragment discreteness index (FDI), the orientation-aware fragmentation feature (OCF), the transcription start site nearest neighbor coverage (TSS), and the nucleosome depletion region coverage (NDR). The construction of the feature vector specifically includes: calculating four features for each identified cancer-specific chromatin accessibility marker interval; then grouping by cancer type; calculating the arithmetic mean of the four features for each cancer type; thus obtaining a single feature value representing each cancer type in that feature dimension; and finally obtaining the 4-dimensional feature mean vector corresponding to each cancer type, which together constitute the feature vector.

[0036] The Directional Sensing Fragmentation Feature (OCF) is obtained by calculating the directional difference or ratio of the number of 5' ends of cfDNA fragments originating from the positive and negative strands of the genome within a specific region, where a higher OCF value indicates higher accessibility.

[0037] The OCF value is calculated using the following formula:

[0038]

[0039] Where U is the upstream signal of the segment and D is the downstream signal of the segment.

[0040] The transcription start site nearest neighbor coverage (TSS) is obtained by calculating the average coverage depth of cfDNA within a 2kb window consisting of a 1kb range upstream and downstream of the center of each cancer-specific chromatin accessibility marker region, and then normalizing it using RPKM. A lower TSS value indicates higher accessibility.

[0041] The nucleosome depletion region coverage (NDR) is obtained by calculating the average cfDNA coverage depth within a 500bp window of 250bp upstream and downstream of the center of each cancer-specific chromatin accessibility marker region, and calculating its ratio to the total coverage depth of the TSS interval corresponding to that marker region as calculated according to claim 14, wherein a lower NDR value indicates higher accessibility.

[0042] The multi-class machine learning model is built on the H2O AutoML framework.

[0043] The inference of its tissue origin further includes outputting a Top-2 cancer type prediction result.

[0044] An apparatus for using any one of the methods described herein, characterized in that it comprises:

[0045] One or more processors;

[0046] A memory having stored thereon computer-executable instructions that, when executed by the one or more processors, enable the apparatus to perform the method of any one of claims 1 to 17.

[0047] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the methods described herein.

[0048] The beneficial effects of this invention are:

[0049] The novelty of this patent is reflected in four aspects: First, it uses the physical statistical composite index FDI, which is "endpoint dispersion × coverage fluctuation," to characterize accessibility, weakening the dependence on absolute coverage depth and mutation frequency, and adapting to the lower-cost sWGS. Second, it introduces a hierarchical significance test of "global beta + local normality," and works with mappability, dark area masking, and distance merging strategies to systematically suppress false positives caused by sequencing bias and spatial autocorrelation. Third, it screens cancer-specific accessibility biomarkers, effectively removing interference signals caused by common regions across cancer types / tissues, which can significantly improve the discriminative power and accuracy of source tracing. Fourth, it integrates multi-dimensional accessibility features such as FDI and OCF / TSS / NDR into a machine learning-enabled feature set, providing a universal and scalable engineering path for source tracing of multiple cancer types. Overall, this approach has practical advantages in terms of broad sample sources, user-friendly process, statistical robustness, and low implementation cost. It can be used in the entire process of early screening positive localization, metastasis tracing, and efficacy / relapse monitoring, providing a new methodological support for the clinical translation of cfDNA epigenetics and fragmentomics. Attached Figure Description

[0050] Figure 1 A flowchart of this patent is shown.

[0051] Figure 2 shows the number (2a) and length distribution (2b) of the 17 cancer chromatin accessibility markers identified.

[0052] Figure 3 The comparison of FDI values ​​for the identified chromatin accessibility region and the genomic random region is shown.

[0053] Figure 4 The identified chromatin accessibility regions (left) and their FDI values ​​(right) are shown to be enriched around the transcription start site.

[0054] Figure 5 The sequence distribution around the chromatin accessibility region visualized by IGV is shown.

[0055] Figure 6 This shows that there is overlap in chromatin accessibility regions across cancer types.

[0056] Figure 7 The number of cancer-specific chromatin accessibility markers identified and their ability to distinguish them from non-target categories are shown.

[0057] Figure 8 The distribution of accessibility features such as FDI, OCF, TSS, and NDR on cancer-specific biomarkers is shown.

[0058] Figure 9 shows the Top-2 accuracy (9a) and confusion matrix (9b) of the organization origination model training set.

[0059] Figure 10 shows the Top-2 accuracy (10a) and confusion matrix (10b) of the organization traceability model validation set.

[0060] Figure 11 This demonstrates the full-sample Top-2 accuracy of this patent compared to the comparative example.

[0061] Figure 12 This paper shows the Top-2 accuracy on the training set compared to the comparative example.

[0062] Figure 13 The Top-2 accuracy on the validation set of this patent is shown compared to the comparative example. Detailed Implementation

[0063] This invention discloses a method, apparatus, and computer storage medium for identifying cancer-specific chromatin accessibility markers and inferring tissue origin based on shallow whole-genome sequencing (sWGS) cfDNA, belonging to the field of gene detection technology. For 17 types of cancer and healthy individuals, multiple sWGS datasets from each category are merged at the sample level to obtain highly representative samples. On a whole-genome isolength window, the open state is characterized by a comprehensive index of "endpoint dispersion × coverage fluctuation," and robust candidate regions are screened through global and local statistical tests. The authenticity of the candidate regions is verified through methods such as comparison with random intervals, functional element enrichment annotation, and visualization.

[0064] To obtain cancer-specific biomarkers, a category-specific interval identification process is proposed: non-target category candidates are treated as a background set; adjacent intervals with a distance <20bp in the background are merged and deduplicated; intervals with an overlap ratio >0.5 with the background are removed from the target cancer candidate intervals; the remaining intervals are defined as "unique biomarkers". Four types of fragment omics features (FDI, OCF, TSS, and NDR) are extracted from the cancer-specific intervals, and robustness and standardization are performed to generate a 68-dimensional feature matrix of "17 cancers × 4 features" for the sample. A multi-class tissue origination model is trained based on the H2O framework to predict the probability of 17 cancer types and the top-2 possible tissue sources. This invention uses plasma cfDNA as the detection target, achieving efficient discovery and tissue origination of cancer-specific accessibility biomarkers at low depth, and is suitable for non-invasive early screening, identification of unknown primary lesions, and efficacy / prognostic assessment.

[0065] The method of the present invention is implemented in the following manner:

[0066] A method for identifying cancer-specific sWGS cfDNA chromatin accessibility markers, comprising the following steps:

[0067] Step 1: cfDNA extraction, library construction, and sequencing.

[0068] Step 2: Remove adapters and low-quality sequences from the sequencing data to obtain high-quality data.

[0069] Step 3: Align high-quality data to the human reference genome, remove repetitive sequences, and obtain the sequence alignment file.

[0070] Step 4: Merge samples of the same disease type to obtain high-depth samples (bam), segment by chromosome and convert to bed.

[0071] Step 5 involves de novo scanning of the genome using a sliding window with a length of 200 bp and a step size of 20 bp, filtering out windows with poor baseline quality. For each window, the fragment dispersity index (FDI) is defined as the core score. By developing a beta distribution test probability model, regions with significantly high FDI are systematically identified as candidate biomarkers for chromatin accessibility.

[0072] Step 6: Verify the reliability of the candidate regions.

[0073] Step 7: Merge the intervals of all categories except the target category into a background set. Merge adjacent intervals in the background set with a distance <20bp and remove duplicates. Remove subsets from the target category candidate regions that have an overlap ratio >0.5 with the background set. The remaining regions are defined as "unique accessibility markers".

[0074] In step 4, the disease type is either cancer or healthy. Cancer types include breast cancer, cervical cancer, colorectal cancer, esophageal cancer, gastric cancer, head and neck cancer, kidney cancer, hepatobiliary cancer, lung cancer, melanoma, ovarian cancer, pancreatic cancer, prostate cancer, sarcoma, thyroid cancer, urothelial carcinoma, and endometrial cancer. Single-sample sequencing depth is 5X, and pooled high-depth samples are no less than 100X.

[0075] In step 5, the basic quality filtering thresholds are coverage > 0, non-dark areas, and average mappability ≥ 0.9.

[0076] In step 5, FDI = EDI × std(coverage), which is the product of the endpoint dispersity index (EDI) and the variance of cfDNA coverage. Regions with more dispersed endpoints and greater coverage variation have higher FDI. The fragment dispersity index, FDI, is a composite indicator designed to accurately quantify chromatin accessibility within a specific genomic region by integrating information on two complementary cfDNA fragmentation patterns. This index is obtained by multiplying two core components, and its calculation formula is: FDI = EDI × std(coverage). Here, std(coverage) is the standard deviation of fragment coverage, which aims to quantify the variability or inhomogeneity of cfDNA fragment stacking depth within a region. In the calculation, a genomic region is first identified, and then the number of cfDNA sequencing fragments covering each base position within that region is counted, resulting in a coverage depth vector. `std(coverage)` is the standard deviation of this coverage depth vector. Its significance lies in the fact that in open chromatin regions, due to the alternating presence of nucleosomes and naked adapter DNA, nuclease cleavage is uneven, leading to a distinct peak-and-valley structure in the distribution of sequencing fragments in these regions, i.e., dramatic fluctuations in coverage depth, hence a higher standard deviation. Conversely, in tightly packed heterochromatin regions, due to the consistent protection of nucleosomes, fragment coverage is relatively flat, resulting in a lower `std(coverage)` value. `EDI` is the endpoint dispersion index, designed to measure the uniformity or dispersion of the spatial distribution of the 5'-ends of all cfDNA fragments within a region. Its calculation formula is: In this formula, n represents the total number of 5' endpoints of all cfDNA fragments within the genomic region; the summation operator iterates through each endpoint j from 1 to n; for each endpoint j, w defines a small window width centered on it, for example, w takes the value of 20 base pairs; C j,x×w This represents the total number of endpoints falling within this window centered at j and with a width of w. Therefore, the ratio... The proportion of endpoint clustering within the local neighborhood of endpoint j relative to the total number of endpoints in the region is quantified. This ratio is exponentially transformed by a constant x (0.5 in this invention) to further adjust its sensitivity to the degree of clustering. By subtracting this term from 1, the local clustering degree is converted into a local sparsity or dispersion score. Finally, the final EDI value is obtained by summing the local dispersion scores of all n endpoints, dividing by n, and taking the average. A high EDI value indicates a uniform distribution of cfDNA endpoints within the region, corresponding to the physical phenomenon that nuclease cleavage sites tend to be more random in open chromatin regions. Conversely, a low EDI value indicates that endpoints tend to cluster in a few specific locations, which is consistent with the protected and less cleavable characteristics of closed chromatin regions. Finally, by multiplying the EDI by std (coverage), the FDI index requires a region to simultaneously meet two conditions: dispersed endpoint distribution and large fluctuations in coverage depth, in order to obtain a high score. This enables FDI to identify true open chromatin regions more robustly and specifically, while effectively suppressing false positive signals for a single indicator that may be caused by technical noise such as low sequencing depth.

[0077] In step 5, a chromatin accessibility region identification process is designed. The purpose of identifying candidate chromatin accessibility regions in this step is to extract signal regions representing the open state of chromatin with high accuracy and high reliability from whole-genome sequencing data containing a large amount of noise. Specific steps include: 1. Statistical definition: Calculate the FDI for windows that pass the basic quality filter and normalize it to FDI_norm. In the initial stage of this process, the fragment dispersion index FDI is calculated for all genomic windows that pass the basic quality filter, and the FDI values ​​are linearly mapped to the [0,1] interval through normalization. The purpose of this step is to standardize the input data so that it meets the mathematical requirements of the subsequent probability distribution model, ensuring the comparability between different samples or genomic regions. 2. Global background modeling and significance determination: Perform Beta distribution fitting on all FDI_norms across the chromosome range, and calculate the right tail global significance p_global = 1 - CDF. β (FDI normThe core objective of this invention is to initially screen out windows with abnormal FDI values ​​on a macro scale. To this end, within a single chromosome, this invention uses a Beta distribution to fit the probability density of the FDI_norm values ​​of all windows, constructing a global background model that accurately describes the baseline level or "background noise" of the FDI values ​​for that chromosome. Based on this model, the right-tailed cumulative probability of the FDI_norm value for each window is calculated, i.e., p_global, thereby quantifying its significance in deviating from the global background. 3. Multiple Test Correction: Multiple test correction is performed on all p_global values ​​using the Benjamini-Hochberg method to control FDR. Windows that simultaneously satisfy p_global < 0.05 and FDR < 0.05 are retained as globally significant candidates. This step addresses the problem of false positive accumulation (i.e., the multiple test problem) inevitably caused by millions of independent tests. The Benjamini-Hochberg method is introduced to correct all p_global values, significantly improving the statistical robustness of the screening results by controlling the overall false discovery rate (FDE), rather than the single false positive rate, ensuring that all globally significant candidates passing this step have high statistical significance. 4. Local Background Significance Test: For each globally significant candidate, the FDI within a window of ±5kb from its center is taken as the neighborhood background. The mean μ and standard deviation σ of the neighborhood FDI are approximated by a normal distribution, and the right-tailed local significance p_ is calculated. local=1-Φ((FDI-μ) / σ), retaining windows with p_local<0.05. Considering that global significance is not equivalent to local functional hotspots, to exclude cases where certain genomic regions (such as high GC content regions) have inherently high signal values, a local background significance test is used to verify whether the candidate window also constitutes a significant peak in its immediate neighborhood. This test uses the neighboring region of the candidate window as the local background, models the signal distribution of this local environment using a normal distribution approximation, and calculates the significance p_local of the candidate window in this local background. This dual-test strategy ensures that the finally selected signal is not only anomalous in the global background but also prominent in the local environment, thereby improving the specificity and reliability of the signal. 5. Region Integration: For windows that pass both global and local significance tests, those with an adjacent distance ≤200bp are merged as the final candidate segments. The std(coverage), EDI, and FDI of the merged segments are recalculated to characterize the segment signal intensity. The aim is to merge adjacent or nearby significant windows into a continuous segment to better correspond to actual DNA functional elements such as promoters and enhancers. After merging, the signal intensity index of the integrated segment is recalculated, thereby assigning a unified and more representative quantitative characterization to each candidate chromatin accessibility region, providing high-quality input for downstream cancer-specific biomarker screening. Step 6 verifies the reliability of candidate regions from three levels: 1. Randomly select genomic background regions with the same number and length as candidate regions, calculate the FDI, and compare it with the candidate regions. 2. Perform functional annotation on the candidate regions to confirm enrichment. 3. Visual verification.

[0078] Step 7 aims to further purify and identify unique accessibility markers specific to the target cancer type from the candidate chromatin accessibility region set selected in previous steps. This step is crucial for achieving the high-specificity cancer tissue tracing of this invention. Its core principle is to construct a comprehensive set of background signals for non-target categories and use this background set to negatively screen candidate regions for the target category, thereby eliminating interference from common, non-specific accessibility regions that are prevalent in various cancer types or healthy tissues. Specifically, firstly, candidate chromatin accessibility regions from all categories other than the target category, including the remaining 16 cancers and healthy control groups, are merged to construct a comprehensive set of genomic regions representing non-specific signals. Biologically, this set symbolizes chromatin open maps widely present in other tissues or disease states, such as background noise containing promoter regions of housekeeping genes shared by multiple tissues that lack discriminative value. Subsequently, a key data standardization and integration operation is performed on this background set, namely, merging adjacent intervals less than 20 bp apart and removing duplicates. The principle behind this operation is that it effectively addresses the problem of regions belonging to the same biological functional element being identified as multiple neighboring fragments due to experimental randomness or algorithmic boundary effects. By merging these fragments to form more complete and representative continuous regions, it reduces the redundancy of the background set and accurately depicts the genomic extent of non-specific regions. After constructing this integrated and optimized background set, the core filtering operation of this step is then performed: removing subsets that significantly overlap with the background set from the candidate regions of the target category. The judgment criterion uses a precise quantitative indicator, the overlap ratio, defined as the overlap ratio of the target candidate region R. target The total length L of the genomic segments overlapping with the background set overlap Divide by the length of the candidate region itself (Length(R)) target This invention sets a screening threshold of 0.5, meaning that if more than half of the length of a target candidate region overlaps with the background set, it is deemed lacking in specificity and removed. This threshold balances robustness and specificity, tolerating a small amount of random overlap to avoid falsely eliminating true signals, while effectively excluding candidates whose main body is a shared region, thereby reducing the risk of false positives in downstream model analysis. Finally, after the above rigorous construction and screening process, the subset of regions remaining from the target category candidate region set, because they simultaneously meet the two conditions of being significantly accessible in the target category and having no significant overlap with other categories, are defined as unique accessibility markers. These markers can maximize the signal-to-noise ratio of the classification model, thereby achieving accurate inference of the cancer tissue origin of cfDNA samples.

[0079] The tissue tracing method based on cancer-specific chromatin accessibility markers in this patent includes the following steps in its overall process:

[0080] Step 1: Extract four types of fragment omics features (FDI, OCF, TSS, and NDR) from the chromatin accessibility regions specific to each cancer obtained through screening, and perform robustness and standardization processing to generate a 68-dimensional feature matrix of "17 cancers × 4 features" for the sample.

[0081] Step 2: Split the dataset into a training set and a validation set. Use the real cancer categories of the training set samples as the dependent variable Y and the 68-dimensional feature matrix as the input X to construct the Y~X machine learning model.

[0082] Step 3: Predict the validation set and output the probabilities of 17 cancer types and the top-2 cancer types.

[0083] In the above process, the calculation method for generating the 68-dimensional feature matrix is ​​as follows: Taking FDI as an example, for any target cancer among the 17 cancer types (taking breast cancer as an example), firstly, all cancer-specific chromatin accessibility marker regions are obtained. Then, for a test sample, an FDI value is calculated for each breast cancer-specific marker region on the sequencing data of that sample. Finally, all FDI values ​​calculated for all breast cancer-specific marker regions are summarized, and their arithmetic mean is calculated. This single mean is used as the FDI feature value of the test sample in the breast cancer dimension. The same summarization and calculation method is used for the other 16 cancer types and other features such as OCF, TSS, and NDR, ultimately forming the 68-dimensional feature vector of the test sample.

[0084] Example 1

[0085] This embodiment illustrates the generation and processing of the test dataset. It uses 1720 plasma samples: 100 cases from each of the 17 aforementioned cancer types, and 20 healthy adults. The data originates from clinical samples actually obtained by the applicant; blood was collected and plasma separated on the same day, and the plasma underwent WGS analysis with a sequencing depth of 5X.

[0086] WGS Testing Process

[0087] The process involves plasma cfDNA extraction, library construction, whole-genome sequencing, and bioinformatics analysis, and the steps are as follows:

[0088] 1) Centrifuge the whole blood sample to obtain plasma containing cfDNA. Centrifuge the plasma at 4°C for 10 min to obtain the supernatant;

[0089] 2) Add lysis binding buffer, proteinase K and magnetic beads to the supernatant, vortex for 30 seconds, and incubate at room temperature for 10 minutes;

[0090] 3) Place on a magnetic rack and let stand for 3 minutes. Remove the supernatant, add washing buffer 1, vortex to mix for 10 seconds, and incubate at room temperature for 2 minutes.

[0091] 4) Place on a magnetic rack and let stand for 3 minutes. Remove the supernatant, add washing buffer 2, vortex to mix for 10 seconds, and incubate at room temperature for 2 minutes.

[0092] 5) Place on a magnetic rack and let stand for 3 minutes. Remove the supernatant, add 100% ethanol, vortex to mix for 10 seconds, and incubate at room temperature for 2 minutes.

[0093] 6) Place on a magnetic rack and let stand for 3 minutes, remove the supernatant, dry at room temperature for 45-55 minutes, wash with enzyme-free water, and incubate at room temperature for 5 minutes;

[0094] 7) Place on a magnetic rack and let stand for 2 minutes, then transfer the supernatant to a new 1.5 ml centrifuge tube;

[0095] 8) Add Qiager end repair reagent to the cfDNA sample, mix well (must be done on ice), incubate until stable at 4°C, then add IDT 8nt UDI Adapter, mix well, and incubate at 20°C for 15 min;

[0096] 9) Purify the above reaction solution with 0.85X magnetic beads, mix well, and incubate at room temperature for 5 min to allow DNA to bind to the magnetic beads. Place the 96-well PCR plate on a magnetic rack and incubate at room temperature for 5 min until the magnetic beads adhere to the plate. Once the supernatant is clear, discard the supernatant.

[0097] 10) Add 80% ethanol, incubate at room temperature for 30 seconds, remove the ethanol, and dry at room temperature for 2-5 minutes. Repeat twice.

[0098] 11) Add enzyme-free water to wash away the DNA bound to the magnetic beads, incubate at room temperature for 5 min until the supernatant is clear, and transfer the supernatant for instrumentation;

[0099] 12) For the cfDNA library obtained above, whole genome sequencing was performed using the Illumina Hiseq Xten sequencer. After sequencing was completed, a fastq file was generated using bcl2fastq.

[0100] 13) Use Fastp software to perform quality control on the data, remove adapters and low-quality sequences, and align the high-quality data to the human reference genome hg19 using bwa to generate a bam file.

[0101] Example 2

[0102] This embodiment illustrates the process of identifying cancer-specific chromatin accessibility markers.

[0103] Because nucleosomes are sparsely distributed in chromatin-accessible regions, nucleases can more easily enter and produce more random cleavage, resulting in more dispersed fragment endpoints in local space and greater coverage fluctuations with coordinates. In contrast, endpoints are more concentrated and coverage is smoother in closed regions. Therefore, the chromatin accessibility state can be inferred by quantifying the accessibility physical signals during cfDNA fragmentation. A recent study proposed the Fragment Dispersion Index (FDI), which integrates the dispersion of fragment endpoints and the variability of coverage. The magnitude of the FDI reflects the open state of "more random endpoints and more fluctuating coverage."

[0104] FDI is defined as the product of the dispersion of cfDNA endpoints and the variance of coverage. Regions with more dispersed endpoints and greater variation in coverage have higher FDI. The formula is:

[0105] FDI = EDI × std(coverage) defines the endpoint dispersion index. EDI measures the degree of dispersion of the endpoints of a segment, and the formula is:

[0106]

[0107] Where n is the total number of free DNA fragment ends in the region; j is the end index from 1 to n in the region; w

[0108] The preset local interval width is 20bp; x is a preset constant with a value of 0.5; C j,x×w Let EDI be the total number of endpoints within a small interval of width w centered at the j-th endpoint. For a genomic region, the more uniform the endpoint distribution, the larger the EDI; conversely, the more uneven the endpoint distribution, the smaller the EDI.

[0109] Based on the FDI index, chromatin accessibility regions were identified de novo across the entire genome.

[0110] The 1700 cancer samples in this example were randomly stratified into training and validation sets according to cancer type in a 6:4 ratio. Twenty samples from each cancer type were randomly selected from the training set as BAMs and merged into a high-depth representative BAM, with an average sequencing depth of 133X after merging. The 17 cancer type BAMs were segmented by chromosome and converted into BED files. A sliding window with a length of 200 bp and a step size of 20 bp was used to scan the entire genome de novo, retaining only windows that passed the basic quality filter (coverage > 0, non-dark areas, average mappability ≥ 0.9). The FDI of each window was calculated as the core score according to the above formula and normalized to FDI_norm. Then, a beta distribution test probability model was used to systematically identify regions with significantly high FDI as candidate chromatin accessibility markers. 1. Global background modeling and significance determination: Beta distribution fitting was performed on all FDI_norms across the chromosome range, and the right tail global significance p_global = 1 - CDF was calculated. β (FDI norm 2. Multiple Test Correction: Multiple tests are performed on all p_global values ​​using the BH method to control FDR. Windows that simultaneously satisfy p_global < 0.05 and FDR < 0.05 are retained as global significance candidates. 3. Local Background Significance Test: For each global significance candidate, the FDI of the window within ±5kb of its center is taken as the neighborhood background. The mean μ and standard deviation σ of the neighborhood FDI are approximated using a normal distribution, and the right-tailed local significance p_ is calculated. local =1-Φ((FDI-μ) / σ), retaining windows with p_local<0.05. 4. Region integration: For windows that pass both global and local significance tests, merge adjacent segments with a distance ≤200bp as the final candidate segments, and recalculate std(coverage), EDI, and FDI for the merged segments to characterize the segment signal strength.

[0111] Count the number of candidate accessible regions for each type of cancer ( Figure 2a ) and length distribution ( Figure 2b The number of candidate regions was negatively correlated with sequencing depth, with lung cancer having the fewest at 26,575 and thyroid cancer having the most at 113,081, with an average of 43,689.

[0112] To confirm the existence of candidate regions, a three-dimensional scheme was designed to verify their reliability: 1. Randomly select genomic random intervals with the same number and length as candidate regions, calculate the FDI value, and compare it with the candidate regions. Figure 3 Candidate regions (red) have a significantly higher FDI than random regions (blue). 2. Candidate regions are annotated near active transcription start sites (TSS), and candidate regions ( Figure 4 (Left) and its FDI value ( Figure 4(Right) Clearly enriched around the TSS. 3. IGV visualization of sequence distribution around candidate regions ( Figure 5 The accessible interval sequence is less visible to the naked eye than the surrounding area.

[0113] Because there is overlap in the reachable range of candidates between different categories ( Figure 6 Therefore, the following method was used to remove shared intervals across cancer types / categories, searching for chromatin-accessible intervals unique to each cancer type, i.e., intervals that only appear in the target cancer type and are not covered by other cancer types or healthy individuals. First, for 20 healthy samples, candidate accessible intervals for healthy individuals were identified using the same method. Then, for each cancer type: 1. Candidate intervals from other cancers and healthy individuals, excluding the target category, were merged into a background set; 2. Adjacent intervals with a distance <20bp in the background set were merged and deduplicated; 3. A subset of candidate intervals from the target category with an overlap ratio >0.5 with the background set was removed, and the remaining intervals were considered "unique accessible intervals". The difference in FDI between the target cancer type and other categories on their "unique intervals" was assessed: the AUC for all categories was higher than 80%, with the AUC for the thyroid cancer interval reaching as high as 94.67% (…). Figure 7 Above). Count the number of cancer-specific reachable intervals ( Figure 7 The ranges from 1562 to 46530 represent 5.88% to 41.15% of their respective candidate intervals.

[0114] As research on chromatin accessibility continues to deepen, researchers have proposed more and more fragment omics indicators that can characterize accessibility, among which the most representative are OCF, TSS, and NDR.

[0115] OCF (Orientation-aware cell-free fragmentation) is a feature of cell-free cfDNA fragmentation that uses the orientation and position information of the 5' ends of fragments to characterize the "endpoint orientation preference / asymmetry" in open chromatin. Open regions, due to the sparse nucleosomes and easier access by nucleases, typically exhibit a characteristic distribution of fragment endpoints in the upstream and downstream directions. OCF thus serves as a quantitative indicator of chromatin accessibility. Within a known open region, the OCF value is obtained by calculating the directional difference or ratio of positive and negative strand fragment endpoints / coverage on genomic coordinates. An elevated OCF usually indicates a more open region. The calculation principle is as follows: In open regions, the upstream (U) and downstream (D) ends of fragments exhibit the highest read density approximately 60 bp from the center, with the peaks at the U and D ends located on the right and left sides, respectively. Conversely, this pattern does not occur in tissue-specific open chromatin regions where the tissue does not contribute DNA to the plasma. Therefore, the difference in U and D end signals within a 20 bp window is used as the corresponding tissue's OCF value, using the formula:

[0116]

[0117] The TSS (Transcription Start Site) is the anchor point of the gene promoter. When chromatin is in an activated state, nucleosome deletion regions often appear near the TSS, making this region more accessible to the transcription complex and nucleases. Therefore, using the TSS as an anchor point to assess the coverage of local cfDNA fragments can indirectly reflect promoter accessibility. Calculation method: Define a 2kb window, 1kb upstream and downstream of the center of the open region, as the TSS interval. Use bedtools to obtain the coverage depth of all TSS intervals and perform RPKM normalization; this is the TSS characteristic value. The lower the TSS value, the more open the interval.

[0118] NDR (Nucleosome-Depleted Region) refers to DNA segments with sparse or absent nucleosomes, typically located approximately 250 bp upstream of the TSS (Transmission Stratum Corner). NDR is a core marker of chromatin accessibility: absence of nucleosomes → increased susceptibility to nuclease cleavage → reduced cfDNA coverage and more dispersed endpoints. NDR scores are obtained by quantifying the "depletion level" within a local TSS window using coverage depth; lower coverage indicates more significant depletion, and a lower NDR suggests a more open region. Calculation method: A 2kb window (1kb upstream and downstream of the center of the open region) is defined as the TSS interval. The coverage depth of all TSS intervals is obtained using bedtools and normalized using RPKM (i.e., TSS value). A 500bp window (250 bp upstream and downstream of the TSS center) is defined as the central region. The NDR score is calculated as the ratio of the central region's coverage depth to the entire TSS interval.

[0119] Each of these three functions has its own role and they corroborate each other—OCF focuses on endpoint directional patterns, TSS is used for anchored promoter accessibility analysis, and NDR directly characterizes the open state of locally absent nucleosomes. In cfDNA, open regions are usually characterized by more dispersed endpoints, more undulating coverage, and more prominent NDRs adjacent to TSSs. These signals are closely related to gene expression and tissue / tumor-specific accessibility.

[0120] For 17 pooled high-depth samples, four fragment omics features representing accessibility—FDI, OCF, TSS, and NDR—were extracted from the cancer-specific accessibility regions mentioned above. Figure 8 The FDI and OCF of the target cancer type (red dot) were significantly higher than those of other categories (black box), while the TSS and NDR were significantly lower than those of other categories, which is in line with expectations and indicates that the cancer specificity of the indicated area is obvious.

[0121] Example 3

[0122] This embodiment is used to illustrate the tissue tracing process based on the cancer-specific chromatin accessibility markers described in Example 2.

[0123] To evaluate the potential of cancer-specific chromatin accessibility biomarkers in tissue tracing, this invention specifically developed a machine learning model. For the 1700 cancer samples described in Example 1, four fragment omics features (FDI, OCF, TSS, and NDR) were analyzed in each cancer-specific chromatin accessibility region and standardized to generate a 68-dimensional feature matrix of "17 cancers × 4 features". Using the 68-dimensional feature matrix as input X and the true cancer type of the sample as the response variable Y, a machine learning multi-classification model (Y~X) was constructed on the training set using the H2O AutoML framework. The model was then locked, a validation set was predicted, and performance was evaluated.

[0124] The multi-classification model consistently demonstrates high accuracy for most cancer types: the overall Top1Call accuracy on the training set is 0.77 ( Figure 9a The validation set is 0.74. Figure 10a When Top2Call is allowed, the overall accuracy improves to 0.86. Analysis of the confusion matrix (training set - Figure 9b Validation set - Figure 10b Rows represent predicted categories, columns represent true categories, and the color of diagonal numbers indicates the number and proportion of correctly identified categories; off-diagonal numbers represent false positives. The model's Top1Call accuracy remains high (left chart), and Top2Call accuracy has further improved (right chart). A positive prediction value (PPV) is defined as TP / (TP+FP), which represents the proportion of true positives (TP) among samples classified as belonging to that category (percentage on the right side of the confusion matrix). It measures the "purity" of the category prediction; a higher PPV indicates fewer false positives (FP). The PPV on both the training and validation sets remains high.

[0125] These results demonstrate that automated machine learning can effectively capture cancer-specific chromatin-accessible signals even in low-depth samples, achieving highly accurate source tracing results. Consistency between the training and validation sets proves the model's excellent generalization ability. Furthermore, reliable Top2Call predictions indicate the model's ability to effectively identify robust alternatives when the primary choice is incorrect, making it particularly valuable when dealing with complex and heterogeneous datasets.

[0126] Comparative Example

[0127] Compared with Example 3, in order to explore the effect of different feature matrices on the model’s ability to trace the source of the virus, matrices composed of three features, namely OCF, TSS and NDR, were used on the same chromosome accessibility region, and models were built. The Top 2 traceability accuracy of the model was statistically analyzed.

[0128] The model built using only three classic features (OCF, TSS, and NDR) achieved a Top1Call score of 0.5 and a Top2Call score of 0.64 on the test set. When the FDI feature was added, creating a model with four features, the Top1Call score improved to 0.74, and the Top2Call score reached 0.86. Calculations show that the Top2Call accuracy improved by more than 20% compared to using only the three classic features. This significant improvement demonstrates that the FDI feature exhibits excellent performance in organizational attribution tasks and can effectively enhance the model's attribution capabilities.

Claims

1. A method for inferring tissue origin based on cancer-specific chromatin accessibility markers, characterized in that, Its use for non-therapeutic and non-diagnostic purposes includes the following steps: a) Provide cell-free DNA sequencing data of a sample to be tested, which has been processed by removing adapters and low-quality sequences, aligning to a reference genome, and removing repetitive sequences; b) For a pre-determined set of cancer types, obtain a set of cancer-specific chromatin accessibility markers for each cancer type; c) For each cancer-specific chromatin accessibility marker obtained in step b), extract at least one fragment omics feature value for the test sample across all genomic regions defined by the cancer-specific chromatin accessibility marker obtained in step b), and summarize and standardize the multiple feature values ​​extracted for each cancer type to form a multidimensional feature vector representing the test sample. d) Input the feature vector obtained in step c) into a pre-trained multi-class machine learning model, output the probability score of the test sample corresponding to each of the cancer types in the group, and infer its tissue origin based on the probability score.

2. A device for inferring tissue origin based on cancer-specific chromatin accessibility markers, characterized in that, include: The sequencing data acquisition module provides cell-free DNA sequencing data of a sample to be tested, which has been processed by removing adapters and low-quality sequences, aligning to a reference genome, and removing repetitive sequences. The biomarker screening module is used to obtain a set of cancer-specific chromatin accessibility biomarkers for each of a pre-determined set of cancer types. The feature vector construction module is used to extract values ​​of at least one fragment omics feature for the test sample on all genomic regions defined by the cancer-specific chromatin accessibility markers for each cancer type obtained by the marker screening module, and to summarize and standardize the multiple feature values ​​extracted for each cancer type to form a multidimensional feature vector representing the test sample. The source analysis module is used to input the feature vectors obtained by the feature vector construction module into a pre-trained multi-class machine learning model, output the probability score of the test sample corresponding to each of the cancer types in the group, and infer its tissue origin based on the probability score.

3. The apparatus according to claim 2, characterized in that, The cell-free DNA sequencing data was obtained from shallow whole-genome sequencing of plasma samples with a single-sample sequencing depth of 5X. The group of cancer types includes breast cancer, cervical cancer, colorectal cancer, esophageal cancer, stomach cancer, head and neck cancer, kidney cancer, liver and gallbladder cancer, lung cancer, melanoma, ovarian cancer, pancreatic cancer, prostate cancer, sarcoma, thyroid cancer, urothelial carcinoma, and endometrial cancer.

4. The apparatus according to claim 2, characterized in that, Methods for obtaining markers of chromatin accessibility include: i. For each type of cancer and healthy control group, a set of candidate chromatin accessibility regions are identified by analyzing their respective cell-free DNA sequencing data; the cell-free DNA sequencing data is obtained by merging single sample data of multiple diseases to obtain high-depth representative sample data with a depth of not less than 100X. ii. For the target cancer type, the candidate chromatin accessibility regions of all other non-target cancer types and the healthy control group are merged to form a background set, and adjacent regions of this background set with a distance of less than 20 bp on the genome are merged and deduplicated. iii. From the candidate chromatin accessibility regions of the target cancer type, regions with an overlap ratio exceeding 0.5 with the background set are removed. The overlap ratio is defined as the total length of the bases overlapping between the target candidate region and the background set divided by the base length of the target candidate region itself. The remaining regions are defined as cancer-specific chromatin accessibility markers for the target cancer type. The identification of candidate chromatin accessibility regions in step i) is achieved by scanning the genome using a sliding window with a length of 200 bp and a step size of 20 bp, and calculating the fragment dispersion index (FDI) for windows that meet the basic quality filtering criteria. The basic quality filtering criteria are that the sequencing coverage of the window is greater than 0, it does not belong to the dark area composed of repetitive sequences or low-complexity regions of the genome, and the average comparability score of all bases in the window is not less than 0.

9. The fragment dispersion index FDI is defined as the product of the fragment end dispersion index EDI and the fragment coverage standard deviation std(coverage), i.e., FDI = EDI × std(coverage); The fragment end dispersion index (EDI) is used to measure the spatial dispersion of free DNA fragment ends within a region, and it is expressed by the formula... Perform the calculation, where: n is the total number of free DNA fragment ends in this region; j is the last index in the region from 1 to n; w is the preset local cell width, with a value of 20bp; x is a preset constant with a value of 0.5; C j,x×w Let w be the total number of ends in a subinterval centered at the j-th end and with a width of w.

5. The apparatus according to claim 2, characterized in that, The identification of candidate chromatin accessibility regions further includes a statistical test step that combines global and local significance tests to screen windows with significantly high FDI values. The global significance test in the statistical testing step includes: a. Normalize the FDI values ​​of all windows that pass through the basic quality filter to obtain FDI_norm; b. At a single chromosome scale, fit the FDI_norm values ​​of all windows to a Beta distribution and calculate the right-tailed cumulative probability of the FDI_norm value for each window as the global significance p_global. c. The Benjamini-Hochberg method was used to perform multiple tests to correct all p_global values ​​to control the false discovery rate (FDR), and windows that simultaneously satisfy p_global < 0.05 and FDR < 0.05 were retained as global significant candidate windows. The statistical testing step further includes a local significance test, which performs the following operation on each globally significant candidate window: a. Using all windows within a 5kb range upstream and downstream of the center of the candidate window as the local background, the FDI value of the local background is approximated by a normal distribution, and the mean μ and standard deviation σ are calculated. b. Calculate the right-tailed local significance of the candidate window relative to the local background: p_local = 1 - Φ((EDI-μ) / σ), where Φ is the cumulative distribution function of the standard normal distribution; c. Retain windows with p_local < 0.05; The statistical testing step concludes with a region integration step, which merges windows that pass both global and local significance tests and have an adjacency distance of less than or equal to 200 bp on the genome to form the final candidate chromatin accessibility regions.

6. The apparatus according to claim 2, characterized in that, Fragmentomics features include the fragment dispersion index (FDI), directed sensing fragmentation feature (OCF), transcription start site nearest neighbor coverage (TSS), and nucleosome depletion region coverage (NDR); and The construction of the feature vector specifically includes: calculating four features for each identified cancer-specific chromatin accessibility marker interval, then grouping by cancer type, calculating the arithmetic mean of the four features for each cancer type, thereby obtaining a single feature value representing each cancer type in that feature dimension, and finally obtaining the 4-dimensional feature mean vector corresponding to each cancer type, which together constitute the feature vector.

7. The apparatus according to claim 6, characterized in that, The Directional Sensing Fragmentation Feature (OCF) is obtained by calculating the directional difference or ratio of the number of 5' ends of cfDNA fragments originating from the positive and negative strands of the genome within a specific region, where a higher OCF value indicates higher accessibility.

8. The apparatus according to claim 6, characterized in that, The transcription start site nearest neighbor coverage (TSS) is obtained by calculating the average cfDNA coverage depth within a 2kb window consisting of a 1kb range upstream and downstream of the center of each cancer-specific chromatin accessibility marker region, and then normalizing it using RPKM. A lower TSS value indicates higher accessibility. The nucleosome depletion region coverage (NDR) is obtained by calculating the average cfDNA coverage depth within a 500bp window consisting of a 250bp range upstream and downstream of the center of each cancer-specific chromatin accessibility marker region, and then calculating the ratio of the average cfDNA coverage depth within this 500bp window to the total coverage depth of the corresponding TSS interval for that marker region. A lower NDR value indicates higher accessibility.

9. The apparatus according to claim 2, characterized in that, The multi-class machine learning model is built on the H2OAutoML framework; the inference of its tissue origin further includes outputting a Top-2 cancer type prediction result.

10. A computer-readable storage medium having a computer program stored thereon, the computer program, when executed by a processor, implementing the method of claim 1.