A tumor risk prediction method and device based on cfDNA nucleosome fingerprint
By using a cfDNA nucleosome fingerprinting method and whole-genome sequencing data and machine learning models, tumor risk can be directly predicted, solving the problems of complexity and high cost in early cancer detection in existing technologies, and achieving highly sensitive and specific non-invasive cancer screening.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YUNKANG INFORMATION TECH (SHANGHAI) CO LTD
- Filing Date
- 2025-11-17
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies for early cancer screening suffer from problems such as complex detection, low sensitivity, unstable specificity, high cost, and limited detection range. In particular, the low proportion of tumor-specific mutations in cfDNA makes it difficult to effectively detect early-stage cancer.
Using a cfDNA nucleosome fingerprinting method, and through whole-genome sequencing data quality control processing, we identify nucleosome localization boundaries and calculate nucleosome localization strength scores. Combining convolutional neural networks and LASSO penalized logistic regression models, we can directly predict tumor risk, eliminate interference from individual background mutations, and improve detection accuracy.
It enables non-invasive, low-cost early cancer detection, improves the sensitivity and specificity of detection, covers the diverse characteristics of tumor types, and reduces the requirements for equipment and laboratory environment.
Smart Images

Figure CN122117344A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of in vitro diagnostic technology, specifically relating to a method and device for predicting tumor risk based on cfDNA nucleosome fingerprinting. Background Technology
[0002] Because tumor recurrence, metastasis, treatment resistance, and often late diagnosis lead to poor treatment outcomes and even death in most cancer patients, early screening for cancer is considered to have significant clinical benefits in clinical practice across multiple cancer types. However, the implementation of screening methods remains challenging. For example, in the United States and China, clinical experts recommend low-dose computed tomography (LDCT) screening for lung cancer in individuals aged 50-80 years with a smoking history of at least 20 pack-years and who are currently smoking or have quit smoking within the past 15 years. Although LDCT screening has been shown to reduce mortality, adherence in high-risk populations is low (<6%), partly due to its low specificity, radiation exposure, and the potential harm from unnecessary diagnostic procedures caused by overdiagnosis. For other cancers, while early detection can improve patient prognosis, there are currently no effective screening methods. Liquid biopsy may overcome these challenges and provide an attractive non-invasive method for the detection of lung cancer and other malignancies.
[0003] Liquid biopsy technology is revolutionizing cancer diagnosis and treatment. While research on cancer surveillance targeting various biomarkers (such as circulating tumor cells, exosomes, and microRNAs) is actively underway, the detection of somatic mutations in circulating tumor DNA (ctDNA) in plasma circulating cell-free DNA (cfDNA) has been proven to provide non-invasive characterization of somatic malignant genomes. Current applications primarily focus on metastatic diseases with high tumor burden. In this context, cfDNA tumor fractions (TF) are high, and effective mutation characterization can be achieved with minimal modifications to next-generation sequencing technologies (such as whole-exome sequencing or a combination of targeted sequencing).
[0004] Cancer genomes are rich in sequence alterations, but the proportion of tumor-specific (somatic) mutations in cell-free DNA (cfDNA) is typically low, making it difficult to detect true variations amidst the background noise introduced during library construction and sequencing. Significant efforts have been made to detect low-frequency mutations in cfDNA. However, these methods often rely on deep sequencing and are limited to examining a small subset of specific genes within the genome. Due to the small number of tumor genomic equivalents in cfDNA, such methods have limited effectiveness in detecting cancer, especially early-stage cancer. Furthermore, cfDNA sequence alterations can originate from leukocytes, which can interfere with cancer detection. Recent analyses suggest that genome-wide mutational variation, genome-wide fragmentation, and methylation analysis can be used for non-invasive early cancer detection. However, the above methods have the following problems: 1) The experiments are complex and have stringent requirements for equipment and laboratory environment; 2) The detection sensitivity is not high and the specificity is unstable; 3) The detection cost is high. Due to the high cost and complex infrastructure required for detection, ctDNA mutations, DNA methylation or plasma proteomics cannot be routinely obtained; 4) The detection indicators and scope are limited to local target range, which cannot fully cover the diverse types of tumors and also directly affects the room for improvement of sensitivity and specificity. Summary of the Invention
[0005] To address the aforementioned problems in the existing technology, this invention provides a method and apparatus for predicting tumor risk based on cfDNA nucleosome fingerprinting. The objective of this invention can be achieved through the following technical solutions: A tumor risk prediction method based on cfDNA nucleosome fingerprinting includes: S1: Obtain whole-genome sequencing data of plasma or serum cfDNA; S2: The whole genome sequencing data is subjected to quality control processing to remove non-human sequences and obtain human cfDNA sequencing fragments; the human cfDNA sequencing fragments are aligned to the human reference genome to obtain the genomic coordinates of each fragment and its coverage information on the reference genome; S3: Based on the whole genome coverage signal, the enriched region identification algorithm is used to determine the nucleosome localization boundary, and the peak signal intensity is used as the nucleosome localization intensity score; S4: Input the nucleosome localization strength score into the trained tumor risk prediction model and output the predicted risk probability of the subject developing a tumor.
[0006] Specifically, the comparison method described in S2 includes: The human cfDNA sequencing fragment was initially aligned to the human reference genome using an alignment tool to obtain an initial alignment mapping file; the alignment weights of the multiple alignment sequences in the initial alignment mapping file were then reassigned, specifically as follows: Input the initial alignment mapping file containing multiple alignment sequences; The core of the expectation-maximization algorithm is invoked to iteratively calculate the posterior probability of each sequence belonging to each alignment site, and the alignment quality value is updated accordingly. Generate a weighted alignment mapping file, in which each sequence is no longer counted as an integer 1 in subsequent coverage statistics, but is weighted and accumulated with its expected maximum weight; The whole genome coverage is calculated using the weighted alignment mapping file, and in highly repetitive sequence regions, a detectable nucleosome enrichment peak is formed by the weighted signal of the unique supporting sequence.
[0007] Specifically, the calculation method for the nucleosome localization intensity score is as follows: Each chromosome is divided into non-overlapping window segments of fixed length, and the coverage depth of sequence fragments is counted segment by segment. The whole genome was scanned using a sliding window to identify local maxima and to merge adjacent high-abundance regions to form candidate peaks. The Poisson distribution test is used to compare the fragment count of the candidate peak with the background value generated by the randomly shuffled sequence fragment or the adjacent regions on both sides, and the peaks with statistical significance reaching the preset threshold are retained. For significant peaks, expand to both sides, and then determine the precise boundary in one go using local minima or gradient descent method; The total number of sequence fragments within the final boundary is directly recorded as the nucleosome localization strength score.
[0008] Specifically, the rich region identification algorithm described in S3 is implemented as follows: Nucleosome localization data, experimentally validated based on nucleosome sequencing data, was used as a real positive dataset. Peaks in regions on the genome that do not overlap with the real peaks were generated using a random method. Signals were extracted from the alignment mapping file and the signal peaks were standardized. Construct an input matrix, where each peak is represented as a one-dimensional vector. Stack all vectors together to form a matrix, which serves as the input to the CNN model. The trained CNN model calculates the probability value of the nucleosome localization intensity score. Set a probability threshold, retain peaks with probabilities higher than the threshold, and filter out peaks with low probabilities; use the list of filtered peaks as the final high-confidence result.
[0009] Specifically, nucleosome localization signal intensity scores were calculated independently for long cfDNA fragments with a length greater than 120 base pairs and short cfDNA fragments with a length of no more than 120 base pairs. The sliding window width for the long cfDNA fragment is defined as 35 to 80 base pairs; the sliding window width for the short cfDNA fragment is defined as 120 to 180 base pairs.
[0010] Specifically, after calculating the nucleosome localization signal strength score, preprocessing is performed: The genome was divided into continuous 1kb intervals, and the nucleosome localization signal intensity scores within each interval were normalized to a median of 0. A Savitzky-Golay filter with a window size of 21 was used to smooth the intensity score of the normalized nucleosome localization signal.
[0011] Specifically, S3 also includes the calculation of the nucleosome occupancy peak: Within the genomic region, identify continuous windows where the nucleosome localization signal intensity score is greater than the median value within the range of 50-450 base pairs upstream and downstream of it; The center position of the area covered by the continuous window is determined as the position of the nucleosome occupancy peak. The peak value of nucleosome occupancy within a genomic region is calculated as the difference between the maximum nucleosome localization signal intensity score within the region and the average of the nucleosome localization signal intensity scores of the two nearest minimum nucleosome localization points.
[0012] Specifically, the enriched region identification algorithm calculates nucleosome localization signal intensity scores and nucleosome occupancy peaks in the following 25 preset functional genomic regions: Gene structure-related regions: promoter region, exon-intron boundary, CTCF binding site, chromatin compartments; Epigenetic marker regions: DNase hypersensitive region, H3K4me1, H3K4me3, H3K9me, H3K9me2, H3K9me3, H3K9ac, H3K27me, H3K27me2, H3K27me3, H3K27ac, H3K36me3, H3K79me, H3K79me3, H2BK5me, H2BK20ac and ATAC open region.
[0013] Specifically, the tumor risk prediction model is the LASSO penalized logistic regression model; The nucleosome localization signal intensity score and nucleosome occupancy peak value, which characterize nucleosome features, are integrated with the short fragment ratio and terminal motif, which characterize fragment omics features, to construct a feature set.
[0014] Specifically, the feature extraction stage of the data acquisition and feature extraction of the tumor risk prediction model includes a nucleosome fingerprint feature system; a unified data quality control process is implemented, and multiple imputation methods are used to fill in missing values in the data.
[0015] Specifically, the model structure of the tumor risk prediction model includes a feature encoder and a classifier. The feature encoder is an encoding network designed for nucleosome features, which includes a linear transformation layer, a ReLU activation function, and a Dropout layer to map the original high-dimensional features to a unified 128-dimensional semantic space. The classifier uses a two-layer fully connected network to achieve the final risk probability prediction. The output layer uses the Sigmoid function to generate tumor risk probabilities between 0 and 1. An auxiliary learning task is introduced, with the main task set as binary tumor risk prediction and the auxiliary tasks being multi-class cancer type identification and tissue origin classification.
[0016] A tumor risk prediction device based on cfDNA nucleosome fingerprinting, comprising: The data acquisition module is used to acquire whole-genome sequencing data of plasma or serum cfDNA; The data preprocessing module is used to perform quality control processing on the whole genome sequencing data to remove non-human sequences and obtain human cfDNA sequencing fragments; the human cfDNA sequencing fragments are aligned to the human reference genome to obtain the genomic coordinates of each fragment and its coverage information on the reference genome; The nucleosome feature calculation module uses a rich region identification algorithm to determine the nucleosome localization boundary based on the whole genome coverage signal, and uses the peak signal intensity as the nucleosome localization intensity score. The prediction module is used to input the nucleosome localization intensity score into the trained tumor risk prediction model and output the predicted risk probability of the subject developing a tumor.
[0017] The beneficial effects of this invention are as follows: The major innovation of this invention lies in the direct use of low-depth whole-genome sequencing data of cfDNA, which can predict tumor risk by directly comparing it with the healthy baseline cfDNA data of the training set without matching tumor tissue or white blood cell control samples, thereby eliminating interference from individual background mutations. Secondly, compared with traditional methylation markers, cfDNA fragmentation distribution, or sequence biomarkers such as ctDNA mutations and CNV copy number variations, this application develops a novel and more sensitive type of biomarker for training parameters in tumor prediction. The nucleosome localization of the tumor genome identified through the training set (cfDNA samples), including the nucleosome localization signal intensity of regulated and non-regulated regions within the genome, NPS scores, and further conversion of NPS into nucleosome peak values (NPKS), are used as the main training parameters. Secondly, considering the differences in cfDNA fragment distribution, NPS and NPKS are calculated separately for long and short fragments to obtain derived nucleosome localization feature values NPSL and NPSs, as well as NPKSL and NPKSs. Finally, a LASSO penalized logistic regression model is used to rigorously train the quantified genome-wide nucleosome localization master features NPS and NPKS, along with other derived features, to obtain the tumor prediction model.
[0018] Secondly, this application employs a convolutional neural network (CNN) to assist in filtering nucleosome localization of significant peaks. This method retains peaks with "real" signal characteristics (such as sharpness, symmetry, and specific patterns) while eliminating peaks with odd shapes or those more likely to be false positives caused by noise. CNN-assisted filtering is a prime example of the successful application of deep learning computer vision technology to bioinformatics. It does not focus on the sequence itself but treats sequencing coverage depth as a "signal image," using the powerful feature extraction capabilities of convolutional networks to distinguish the shape patterns of "real biological signals" from "technical noise," thereby significantly improving the accuracy of nucleosome localization.
[0019] This application addresses the noise problem in detecting low-frequency mutations in cfDNA by quantifying a series of nucleosome localization features inferred from whole-genome cfDNA. Its core innovation lies in designing and identifying novel biomarkers for non-invasive cancer screening based on the fundamental facts and scientific theories of spatial genome collapse and damage during cancer development, providing a new paradigm for non-invasive tumor screening. Attached Figure Description
[0020] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.
[0021] Figure 1 This is a flowchart illustrating a tumor risk prediction method based on cfDNA nucleosome fingerprinting according to the present invention. Figure 2 This is a schematic diagram of the tumor risk prediction process of the present invention. Detailed Implementation
[0022] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of example embodiments to those skilled in the art. Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this disclosure. The blocks shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices. The flowcharts shown in the drawings are merely illustrative and do not necessarily include all contents and operations / steps, nor do they necessarily have to be performed in the order described. For example, some operations / steps can be broken down, while others can be combined or partially combined. Therefore, the actual execution order may change depending on the actual situation.
[0023] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided.
[0024] Please see Figure 1-2 A tumor risk prediction method based on cfDNA nucleosome fingerprinting includes: S1: Obtain whole-genome sequencing data of plasma or serum cfDNA; S2: The whole genome sequencing data is subjected to quality control processing to remove non-human sequences and obtain human cfDNA sequencing fragments; the human cfDNA sequencing fragments are aligned to the human reference genome to obtain the genomic coordinates of each fragment and its coverage information on the reference genome; S3: Based on the whole genome coverage signal, the enriched region identification algorithm is used to determine the nucleosome localization boundary, and the peak signal intensity is used as the nucleosome localization intensity score; S4: Input the nucleosome localization strength score into the trained tumor risk prediction model and output the predicted risk probability of the subject developing a tumor.
[0025] Specifically, the comparison method described in S2 includes: The human cfDNA sequencing fragment was initially aligned to the human reference genome using an alignment tool to obtain an initial alignment mapping file; the alignment weights of the multiple alignment sequences in the initial alignment mapping file were then reassigned, specifically as follows: Input the initial alignment mapping file containing multiple alignment sequences; The core of the expectation-maximization algorithm is invoked to iteratively calculate the posterior probability of each sequence belonging to each alignment site, and the alignment quality value is updated accordingly. Generate a weighted alignment mapping file, in which each sequence is no longer counted as an integer 1 in subsequent coverage statistics, but is weighted and accumulated with its expected maximum weight; The whole genome coverage is calculated using the weighted alignment mapping file, and in highly repetitive sequence regions, a detectable nucleosome enrichment peak is formed by the weighted signal of the unique supporting sequence.
[0026] Specifically, the calculation method for the nucleosome localization intensity score is as follows: Each chromosome is divided into non-overlapping window segments of fixed length, and the coverage depth of sequence fragments is counted segment by segment. The whole genome was scanned using a sliding window to identify local maxima and to merge adjacent high-abundance regions to form candidate peaks. The Poisson distribution test is used to compare the fragment count of the candidate peak with the background value generated by the randomly shuffled sequence fragment or the adjacent regions on both sides, and the peaks with statistical significance reaching the preset threshold are retained. For significant peaks, expand to both sides, and then determine the precise boundary in one go using local minima or gradient descent method; The total number of sequence fragments within the final boundary is directly recorded as the nucleosome localization strength score.
[0027] Specifically, the rich region identification algorithm described in S3 is implemented as follows: Nucleosome localization data, experimentally validated based on nucleosome sequencing data, was used as a real positive dataset. Peaks in regions on the genome that do not overlap with the real peaks were generated using a random method. Signals were extracted from the alignment mapping file and the signal peaks were standardized. Construct an input matrix, where each peak is represented as a one-dimensional vector. Stack all vectors together to form a matrix, which serves as the input to the CNN model. The trained CNN model calculates the probability value of the nucleosome localization intensity score. Set a probability threshold, retain peaks with probabilities higher than the threshold, and filter out peaks with low probabilities; use the list of filtered peaks as the final high-confidence result.
[0028] Specifically, nucleosome localization signal intensity scores were calculated independently for long cfDNA fragments with a length greater than 120 base pairs and short cfDNA fragments with a length of no more than 120 base pairs. The sliding window width for the long cfDNA fragment is defined as 35 to 80 base pairs; the sliding window width for the short cfDNA fragment is defined as 120 to 180 base pairs.
[0029] Specifically, after calculating the nucleosome localization signal strength score, preprocessing is performed: The genome was divided into continuous 1kb intervals, and the nucleosome localization signal intensity scores within each interval were normalized to a median of 0. A Savitzky-Golay filter with a window size of 21 was used to smooth the intensity score of the normalized nucleosome localization signal.
[0030] Specifically, S3 also includes the calculation of the nucleosome occupancy peak: Within the genomic region, identify continuous windows where the nucleosome localization signal intensity score is greater than the median value within the range of 50-450 base pairs upstream and downstream of it; The center position of the area covered by the continuous window is determined as the position of the nucleosome occupancy peak. The peak value of nucleosome occupancy within a genomic region is calculated as the difference between the maximum nucleosome localization signal intensity score within the region and the average of the nucleosome localization signal intensity scores of the two nearest minimum nucleosome localization points.
[0031] Furthermore, the enriched region identification algorithm calculates the nucleosome localization signal intensity score and nucleosome occupancy peak value in the following 25 preset functional genomic regions: Gene structure-related regions: promoter region, exon-intron boundary, CTCF binding site, chromatin compartments; Epigenetic marker regions: DNase hypersensitive region, H3K4me1, H3K4me3, H3K9me, H3K9me2, H3K9me3, H3K9ac, H3K27me, H3K27me2, H3K27me3, H3K27ac, H3K36me3, H3K79me, H3K79me3, H2BK5me, H2BK20ac and ATAC open region.
[0032] Specifically, the tumor risk prediction model is the LASSO penalized logistic regression model; The nucleosome localization signal intensity score and nucleosome occupancy peak value, which characterize nucleosome features, are integrated with the short fragment ratio and terminal motif, which characterize fragment omics features, to construct a feature set.
[0033] Specifically, the feature extraction stage of the data acquisition and feature extraction of the tumor risk prediction model includes a nucleosome fingerprint feature system; a unified data quality control process is implemented, and multiple imputation methods are used to fill in missing values in the data.
[0034] Specifically, the model structure of the tumor risk prediction model includes a feature encoder and a classifier. The feature encoder is an encoding network designed for nucleosome features, which includes a linear transformation layer, a ReLU activation function, and a Dropout layer to map the original high-dimensional features to a unified 128-dimensional semantic space. The classifier uses a two-layer fully connected network to achieve the final risk probability prediction. The output layer uses the Sigmoid function to generate tumor risk probabilities between 0 and 1. An auxiliary learning task is introduced, with the main task set as binary tumor risk prediction and the auxiliary tasks being multi-class cancer type identification and tissue origin classification.
[0035] A tumor risk prediction device based on cfDNA nucleosome fingerprinting, comprising: The data acquisition module is used to acquire whole-genome sequencing data of plasma or serum cfDNA; The data preprocessing module is used to perform quality control processing on the whole genome sequencing data to remove non-human sequences and obtain human cfDNA sequencing fragments; the human cfDNA sequencing fragments are aligned to the human reference genome to obtain the genomic coordinates of each fragment and its coverage information on the reference genome; The nucleosome feature calculation module uses a rich region identification algorithm to determine the nucleosome localization boundary based on the whole genome coverage signal, and uses the peak signal intensity as the nucleosome localization intensity score. The prediction module is used to input the nucleosome localization intensity score into the trained tumor risk prediction model and output the predicted risk probability of the subject developing a tumor.
[0036] In this embodiment, plasma cfDNA whole genome sequencing data of patient samples and healthy samples are obtained, and cfDNA is extracted from the collected plasma, followed by quality inspection, whole genome library construction, and sequencing. Adapters and low-quality bases were removed using FASTP or CutAdapt; non-human sequences from mitochondria, bacteria, viruses, etc., were removed; the sequencing fragments were aligned to the reference genome, and fragment coordinates and coverage were extracted; the cfDNA was aligned to the reference human genome sequence using BWA alignment software (and aligned to the human reference genome hg19 in the UCSC database to obtain fragment information files) to obtain binary alignment files (BAM); fragment coverage and fragment coordinates for each region of the genome were extracted from BAM using Samtools; PE mode sequencing (paired-end PE150) was used to extract the coordinates of correctly matched paired ends; and fragment coverage, i.e., all positions (including endpoints) between two fragment endpoints, was calculated. The EM algorithm is used to further re-align initially aligned reads to improve the accuracy of multiple alignment reads and handle special repetitive sequence regions. This includes extracting BAM files containing multiple alignment reads; running the EM algorithm (using the algorithm core of tools such as RSEM and Salmon); outputting a new BAM file or counting matrix; in the new BAM file, the alignment quality value (MAPQ) for each multiple alignment read is modified. The MAPQ value is set to a value related to P_{ij}, representing the confidence level of the assignment. Nucleosome positioning score (NPS) calculation based on Poisson distribution and convolutional neural network (CNN): (1) Calculate the coverage depth vector: Divide each chromosome into non-overlapping window intervals bins and count the coverage depth of reads in each bin; (2) Identify candidate peak regions: Calculate local maxima using a sliding window. Merge adjacent significantly enriched regions; (3) Background modeling and significance test: Use the Poisson test or negative binomial distribution test to evaluate whether each candidate peak is significantly higher than the background; the background region can be generated by randomly shuffling reads or using flanking regions; (4) Precise boundary location: For each significant peak, extend a certain range on both sides and use the local minimum or gradient descent method to determine the precise boundary; (5) The cumulative read count within the significant peak range is taken as the NPS, i.e., the nucleosome localization signal strength score, based on the precise positioning boundary. (6) Convolutional Neural Network (CNN) assisted filtering of kernel bodies to locate significant peaks.
[0037] (7) Construction of training set: Nucleosome localization data rigorously verified by nucleosome sequencing data is used as the real positive dataset. Peaks on the genome that do not overlap with the real peaks are generated using a random method. Then, the signal is extracted from the aligned BAM and the signal peaks are Z-score normalized.
[0038] (8) Constructing the input matrix: Each peak is represented as a one-dimensional vector X i =[x1,x2,x3,...,x L ], where L is the sequence length (e.g., 101 bp), x i This is the normalized coverage depth at that location. All vectors are stacked to form a matrix of size [N_samples,L,1], which serves as the input to the CNN. This can be viewed as a "single-channel" one-dimensional image.
[0039] (7) CNN model architecture design: Input layer: Receives a vector of shape (L,1), where L is the length of the input sequence.
[0040] Convolutional layers: Multiple one-dimensional convolutional kernels (e.g., kernel_size=3,5,7) slide across the sequence to extract local features. The first convolutional layer may learn basic shapes such as "spikes," "double peaks," "broad peaks," and "slopes." Subsequent convolutional layers combine these basic features into more complex patterns. ReLU is used as the activation function to introduce non-linearity.
[0041] Fully connected layers: After high-level features are extracted by convolutional and pooling layers, these features are flattened and fed into one or more fully connected layers. The role of fully connected layers is to synthesize all features and perform classification.
[0042] Output layer: Typically a single neuron with a sigmoid activation function, outputting a probability value between 0 and 1, representing the probability that the input region is a "true peak".
[0043] Regularization: Use Dropout layers to randomly drop a portion of neurons between fully connected layers to prevent the model from overfitting.
[0044] (8) Model training and optimization: using the binary cross-entropy loss function: Loss = -[y*log(p) + (1-y)*log(1-p)]; where y is the true label (0 or 1), and p is the probability predicted by the model. The optimizer uses optimizers such as Adam or SGD to minimize the loss function and update the network weights.
[0045] (9) Training process: Divide the dataset into training, validation, and test sets. Train the model on the training set and monitor performance (such as accuracy and AUC) on the validation set to prevent overfitting. Stop early if the validation set performance no longer improves. Finally, evaluate the model's final performance on the test set.
[0046] (10) Filtering nucleosome localization peaks: Standardize the NPS based on the Poisson test by inputting the vector into the trained CNN model to obtain the probability P value of the NPS. Set a probability threshold. Keep peaks with probabilities higher than the threshold and filter out peaks with low probabilities. Use the list of filtered peaks as the final high-confidence result.
[0047] NPS is calculated independently for both long and short cfDNA fragments. Considering the potential for two distinct distribution patterns of long and short cfDNA fragments, sliding window widths are applied to fragments of different sizes to calculate independent nucleosome localization scores on the genome: the long fragment NPS score is defined as NPSL, and the short fragment NPS score as NPSs, where subscript L represents a long fragment and subscript S represents a short fragment; based on prior knowledge, the sliding window widths for long and short fragments are set to kL = 35-80 and ks = 120-180, in units of (bp). NPSL and NPSs were standardized within a 1kb range (median of 0). Next, the NPSL and NPSs within the 1kb range were smoothed to preserve nucleosome signal characteristics and enhance the signal-to-noise ratio. Smoothing methods included, but were not limited to, wavelet denoising, Savitzky-Golay filters, wavelet analysis, local regression smoothing, and median filtering. This invention compared and evaluated the effects of different smoothing methods on NPSs, finding that in cfDNA nucleosome mapping studies, Savitzky-Golay was optimal due to its balance between preserving peak width accuracy (a key indicator of nucleosome occupancy) and suppressing oscillatory noise. The second-order polynomial fitting of Savitzky-Golay accurately described the parabolic distribution of nucleosome protection strength (strong at the center → weak at the edges), and the empirical window size for Savitzky-Golay parameters was set to 21. Internal smoothing can also be performed by combining wavelet denoising and Savitzky-Golay processing. The processed NPSL and NPSs are labeled as NPSL-sig and NPSs-sig, respectively (the subscript sig indicates the signal after smoothing). NPSL-sig and NPSs-sig are used as multidimensional input parameter variables for the artificial intelligence model for predicting tumors.
[0048] Nucleosome occupancy peaks refer to the local maximum values of nucleic acid sequences protected by nucleosomes, reflecting the occupancy of nucleosomes in the core region. The coordinate search method for nucleosome occupancy peaks is defined as follows: Assuming the genomic coordinates of the nucleosome occupancy peak are P{S,E}, where S represents the start position and E represents the end position, then the start position of the nucleosome occupancy peak is located within the range of (S,E) and above the median NPS within this region. The position {S,E} of the nucleosome occupancy peak must satisfy the following: within the range of S and E, there exists a position where NPSi > Median (NPSr), where position i is between S and E, and r = 50~450, where r represents the length of the region considered for peak calculation, with a minimum value of 50bp and a maximum value of 450bp. Therefore, the nucleosome occupancy peak position P = (S+E) / 2, which is the center coordinate of the region {S,E}. Subsequent steps will use P (the center coordinates of the peak) for calculating the nucleosome spacing density in the nucleosome occupancy map and for other analyses.
[0049] The peak value of nucleosome occupancy is defined as the difference between the maximum value (denoted by NPSi, where coordinate i belongs to the range {S,..,E}) within a window of coordinates {S,E} and the average of the two lowest points of the nearest NPS (denoted by NPSaround). NPKS is calculated independently for a single range (<150bp) or multiple times for a larger region (greater than 150bp and less than 450bp). Peak value calculations for excessively short or long regions (<50bp or >450bp) are not considered. NPKS is calculated separately for long and short fragments, labeled as NPKSL and NPKSs respectively, and used as multidimensional input parameter variables for the artificial intelligence model predicting tumors.
[0050] This invention targets whole-genome regulatory sequences marked by authoritative databases (including promoter regions (TSS), exon-intron boundary regions, CTCF binding sites, chromatin compartments (A / B compartments), and epigenetic regulatory marker regions (including DNase regions, H3K4me1, H3K4me3, H3K4ac, H3K9me, H3K9me2, H3K9me3, H3K9ac, H3K14ac, H3K27me, H3K27me2, H3K27me3, and H3K27ac regions) analyzed from data from various experimental techniques such as transcriptomics, ChIP-seq, DNase-seq, and ATAC-seq integrated by the ENCODE project). Nucleosome occupancy signal score (NPS) and nucleosome peak NPKS were calculated sequentially for functional sequence regions such as H3K79me, H3K79me3, H3K36me3, H3K40me, H3K40me3, H2BK5me, H2BK20ac, and ATAC chromosome proximity region. The NPS and NPKS scores for these 25 types of regions were calculated cumulatively.
[0051] Downloaded the BED format coordinates of the 25 defined regions from the USCS and ENCODE databases. Then, cfDNA fragments were extracted from the cfDNA alignment file (.bam). Finally, NPS and NPKS were calculated according to steps S6-S7. The NPS and NPKS obtained in this step were also used as input parameters for subsequent early cancer screening risk models.
[0052] A tumor risk model was constructed using a LASSO penalized logistic regression model trained on nucleosome localization atlas data. To address the sparsity issue in high-dimensional data, the LASSO model was used to train the nucleosomes, yielding the NPS and NPKS features obtained from the previous steps. An L1 norm penalty term was included in the LASSO penalized logistic regression model. This is to achieve feature selection and prevent overfitting. Specifically, the tumor risk model is trained in the form of a model: , Where p represents the probability that the sample is cancerous; β0 is the intercept term; β k is the regression coefficient of the k-th feature; the feature includes nucleosome occupancy features, which are derived from NPS and NPKS scores.
[0053] The objective function after adding the LASSO penalty term is: , Where L(β) is the loss function (negative log-likelihood) of standard logistic regression, and its specific expression is: , Where p is the sigmoid function and β is the coefficient vector; The model is trained using LASSO penalized logistic regression. Specific configuration: Method for selecting the penalty strength λ: Implemented using the R software caret package, and optimized through 10 repeated 5-fold cross-validation; select the λ that minimizes the bias; Logistic regression prediction operation: Input the nucleosome occupancy map NPS and NPKS into the trained LASSO model: , Finally, log-odds are converted into a PlasmaAtlas Score (PAS), which represents the probability of tumor risk. The formula for calculating the PlasmaAtlas Score is as follows: , The probability range of PAS is 0 to 1. Generally, log-odds=1.0 is used as the risk threshold, i.e., PAS=0.5. When PAS>0.5, the sample is classified as high-risk for cancer. The sensitivity, specificity, and AUC of the cfDNA model are evaluated in independent validation set samples.
[0054] In this embodiment, the nucleosome fingerprint feature system constructs a complete nucleosome feature matrix based on the original NPS (nucleosome localization signal strength score) and NPKS (nucleosome occupancy peak value): Basic features: NPS and NPKS calculations across the entire genome; Long and short segment separation features: NPSL (long segment localization score), NPSs (short segment localization score), NPKSL (long segment occupancy peak), NPKSs (short segment occupancy peak). Functional region characteristics: NPS and NPKS were calculated for 25 pre-defined functional genomic regions (including promoter regions, exon-intron boundaries, CTCF binding sites, chromatin compartments, and various histone modification regions); The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A tumor risk prediction method based on cfDNA nucleosome fingerprinting, characterized in that, include: S1: Obtain whole-genome sequencing data of plasma or serum cfDNA; S2: The whole genome sequencing data is subjected to quality control processing to remove non-human sequences and obtain human cfDNA sequencing fragments; the human cfDNA sequencing fragments are aligned to the human reference genome to obtain the genomic coordinates of each fragment and its coverage information on the reference genome; S3: Based on the whole genome coverage signal, the enriched region identification algorithm is used to determine the nucleosome localization boundary, and the peak signal intensity is used as the nucleosome localization intensity score; S4: Input the nucleosome localization strength score into the trained tumor risk prediction model and output the predicted risk probability of the subject developing a tumor.
2. The method according to claim 1, characterized in that, The comparison method described in S2 includes: The human cfDNA sequencing fragment was initially aligned to the human reference genome using an alignment tool to obtain an initial alignment mapping file; the alignment weights of the multiple alignment sequences in the initial alignment mapping file were then reassigned, specifically as follows: Input the initial alignment mapping file containing multiple alignment sequences; The core of the expectation-maximization algorithm is invoked to iteratively calculate the posterior probability of each sequence belonging to each alignment site, and the alignment quality value is updated accordingly. Generate a weighted alignment mapping file, in which each sequence is no longer counted as an integer 1 in subsequent coverage statistics, but is weighted and accumulated with its expected maximum weight; The whole genome coverage is calculated using the weighted alignment mapping file, and in highly repetitive sequence regions, a detectable nucleosome enrichment peak is formed by the weighted signal of the unique supporting sequence.
3. The method according to claim 1, characterized in that, The method for calculating the nucleosome localization strength score is as follows: Each chromosome is divided into non-overlapping window segments of fixed length, and the coverage depth of sequence fragments is counted segment by segment. The whole genome was scanned using a sliding window to identify local maxima and to merge adjacent high-abundance regions to form candidate peaks. The Poisson distribution test is used to compare the fragment count of the candidate peak with the background value generated by the randomly shuffled sequence fragment or the adjacent regions on both sides, and the peaks with statistical significance reaching the preset threshold are retained. For significant peaks, expand to both sides, and then determine the precise boundary in one go using local minima or gradient descent method; The total number of sequence fragments within the final boundary is directly recorded as the nucleosome localization strength score.
4. The method according to claim 1, characterized in that, The specific implementation of the enriched region identification algorithm described in S3 is as follows: Nucleosome localization data, experimentally validated based on nucleosome sequencing data, was used as a real positive dataset. Peaks in regions on the genome that do not overlap with the real peaks were generated using a random method. Signals were extracted from the alignment mapping file and the signal peaks were standardized. Construct an input matrix, where each peak is represented as a one-dimensional vector. Stack all vectors together to form a matrix, which serves as the input to the CNN model. The trained CNN model calculates the probability value of the nucleosome localization intensity score. Set a probability threshold, retain peaks with probabilities higher than the threshold, and filter out peaks with low probabilities; use the list of filtered peaks as the final high-confidence result.
5. The method according to claim 1, characterized in that, Nucleosome localization signal intensity scores were calculated independently for long cfDNA fragments with a length greater than 120 base pairs and short cfDNA fragments with a length of no more than 120 base pairs. The sliding window width for the long cfDNA fragment is defined as 35 to 80 base pairs; the sliding window width for the short cfDNA fragment is defined as 120 to 180 base pairs.
6. The method according to claim 1, characterized in that, After calculating the nucleosome localization signal strength score, preprocessing is performed: The genome was divided into continuous 1kb intervals, and the nucleosome localization signal intensity scores within each interval were normalized to a median of 0. A Savitzky-Golay filter with a window size of 21 was used to smooth the intensity score of the normalized nucleosome localization signal.
7. The method according to claim 1, characterized in that, S3 also includes the calculation of the nucleosome occupancy peak: Within the genomic region, identify continuous windows where the nucleosome localization signal intensity score is greater than the median value within the range of 50-450 base pairs upstream and downstream of it; The center position of the area covered by the continuous window is determined as the position of the nucleosome occupancy peak. The peak value of nucleosome occupancy within a genomic region is calculated as the difference between the maximum nucleosome localization signal intensity score within the region and the average of the nucleosome localization signal intensity scores of the two nearest minimum nucleosome localization points.
8. The method according to claim 4, characterized in that, The enriched region identification algorithm calculates nucleosome localization signal intensity scores and nucleosome occupancy peaks in the following 25 preset functional genomic regions: Gene structure-related regions: promoter region, exon-intron boundary, CTCF binding site, chromatin compartments; Epigenetic marker regions: DNase hypersensitive region, H3K4me1, H3K4me3, H3K9me, H3K9me2, H3K9me3, H3K9ac, H3K27me, H3K27me2, H3K27me3, H3K27ac, H3K36me3, H3K79me, H3K79me3, H2BK5me, H2BK20ac and ATAC open region.
9. The method according to claim 1, characterized in that, The tumor risk prediction model is the LASSO penalized logistic regression model. The nucleosome localization signal intensity score and nucleosome occupancy peak value, which characterize nucleosome features, are integrated with the short fragment ratio and terminal motif, which characterize fragment omics features, to construct a feature set.
10. The method according to claim 9, characterized in that, The feature extraction stage of the data acquisition and feature extraction of the tumor risk prediction model includes the nucleosome fingerprint feature system; a unified data quality control process is implemented, and multiple imputation methods are used to fill in missing values in the data.
11. The method according to claim 1, characterized in that, The tumor risk prediction model has a structure including a feature encoder and a classifier. The feature encoder is designed with a coding network for nucleosome features, including a linear transformation layer, a ReLU activation function, and a Dropout layer, which maps the original high-dimensional features to a unified 128-dimensional semantic space. The classifier uses a two-layer fully connected network to achieve the final risk probability prediction. The output layer uses the Sigmoid function to generate tumor risk probabilities between 0 and 1; an auxiliary learning task is introduced, with the main task set as binary tumor risk prediction and the auxiliary tasks as multi-class cancer type identification and tissue origin classification.
12. A tumor risk prediction device based on cfDNA nucleosome fingerprinting, used to perform the method as described in any one of claims 1-11, characterized in that, include: The data acquisition module is used to acquire whole-genome sequencing data of plasma or serum cfDNA; The data preprocessing module is used to perform quality control processing on the whole genome sequencing data to remove non-human sequences and obtain human cfDNA sequencing fragments; the human cfDNA sequencing fragments are aligned to the human reference genome to obtain the genomic coordinates of each fragment and its coverage information on the reference genome; The nucleosome feature calculation module uses a rich region identification algorithm to determine the nucleosome localization boundary based on the whole genome coverage signal, and uses the peak signal intensity as the nucleosome localization intensity score. The prediction module is used to input the nucleosome localization intensity score into the trained tumor risk prediction model and output the predicted risk probability of the subject developing a tumor.