cfDNA extraction method and application thereof in construction of lung cancer early screening model
By combining variational autoencoder models and convolutional neural networks with single-cell ATAC-seq technology, epithelial cell-specific open chromatin regions in cfDNA are identified, and a deep learning-based cfDNA early screening model is constructed. This solves the sensitivity and cost problems of non-invasive early lung cancer detection in existing technologies and achieves high-precision lung cancer risk prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- KUNMING MEDICAL UNIVERSITY
- Filing Date
- 2025-10-11
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies lack sensitive, low-cost, and non-invasive early screening methods, making it difficult to achieve dynamic monitoring and early detection of lung nodules in high-risk individuals, especially the identification of tumor epithelial cell characteristics.
By combining variational autoencoder models and convolutional neural networks with single-cell ATAC-seq technology, epithelial cell-specific open chromatin regions in cfDNA are identified, and a deep learning-based cfDNA early screening model is constructed. This model is then combined with low-depth whole-genome sequencing for lung cancer risk prediction.
It has achieved highly sensitive, low-cost, non-invasive early lung cancer detection, improving the accuracy and accessibility of screening while reducing testing costs.
Smart Images

Figure CN120905208B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of DNA extraction, and more particularly to a cfDNA extraction method and its application in constructing an early screening model for lung cancer. Background Technology
[0002] Lung cancer is one of the leading causes of cancer-related morbidity and mortality worldwide, with non-small cell lung cancer (NSCLC) accounting for approximately 85% of all lung cancer cases. Most patients are diagnosed at an advanced stage, resulting in extremely low five-year survival rates. Studies show that effective intervention in the early stages of lung cancer can achieve a five-year survival rate of around 90%, while even with treatment, the five-year survival rate for advanced-stage patients is only 2% to 5%. Therefore, dynamic monitoring and early detection of potentially high-risk individuals are crucial for improving the clinical prognosis of NSCLC.
[0003] Currently, there is a lack of clinically effective, low-cost, and non-invasive early screening methods capable of real-time monitoring of changes in lung nodules in high-risk individuals. In recent years, cell-free plasma DNA (cfDNA) has attracted widespread attention as a non-invasive biomarker. Studies have found that cellular characteristic signals embodied in cfDNA can effectively reflect the activity state of epithelial cells and indicate potential abnormal changes in the body. In normal individuals, cfDNA mainly originates from peripheral blood mononuclear cells (PBMCs). Therefore, identifying non-PBMC-derived components of cfDNA, especially fragments from tumor epithelial cells, holds promise as an emerging molecular indicator for early cancer screening.
[0004] During tumor progression, the abnormal proliferation and apoptosis of cancer cells release cfDNA, which carries characteristics of chromatin openness, into the bloodstream. These fragments structurally reflect epithelial cell-specific chromatin states and can serve as signals for early tumor recognition. The fragmentation pattern of cfDNA is regulated by chromatin conformation and nucleosome distribution. In open chromatin regions, due to fewer nucleosomes, DNA is more easily cleaved by nucleases, resulting in significantly reduced cfDNA coverage in these regions. Conversely, regions with dense nucleosomes have more abundant cfDNA fragments due to structural protection. The literature (Snyder, MW, Kircher, M., Hill, AJ, Daza, RM, and Shendure, J. (2016). Cell-free DNA Comprises an In Vivo Nucleosome Footprint that Informs Its Tissues-Of-Origin. Cell 164, 57-68. 10.1016 / j.cell.2015.11.050) successfully used fragment distribution information of cfDNA to infer its tissue origin, verifying the feasibility of this principle.
[0005] To more accurately analyze the chromatin accessibility patterns hidden in cfDNA, single-cell level ATAC-seq (scATAC-seq) technology provides an important tool for research. This technology can reveal the chromatin accessibility of different cell subpopulations at single-cell resolution and is widely used in tumor heterogeneity analysis. By combining the fragment coverage of cfDNA with the epithelial cell chromatin atlas constructed by scATAC-seq, and supplementing it with the classification and recognition capabilities of deep learning algorithms, it is expected to efficiently extract characteristic regions of tumor epithelial origin from complex cfDNA mixed signals, thereby significantly improving the sensitivity and accuracy of early detection.
[0006] Based on the aforementioned research, establishing a strategy for early lung cancer risk prediction relying on low-depth whole-genome sequencing would significantly reduce screening costs while improving screening accessibility and accuracy. Therefore, developing a novel non-invasive early screening method combining cfDNA and epithelial cell signal recognition has significant scientific value and application prospects, and also provides a potential breakthrough for early management of high-risk groups for lung cancer. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this invention provides a method for constructing a cfDNA cancer early screening model based on deep learning and fusing epithelial cell feature signals. This invention uses a variational autoencoder model and a convolutional neural network model to identify epithelial cell-specific cfDNA region features from blood samples, providing a new analytical tool for cancer research. Based on these features, specific probes can be designed to achieve low-cost, high-precision cfDNA detection.
[0008] It should be noted that the cfDNA extraction method and the method for constructing an early cancer screening model provided by this invention aim to quantify and classify cfDNA features through bioinformatics analysis and machine learning technology. The output results are probabilistic risk assessment signals (such as confidence scores for lung cancer-related features) and are not intended for diagnostic or therapeutic purposes.
[0009] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: a cfDNA extraction method, comprising the following steps:
[0010] Step (1): Use a special blood collection tube pre-loaded with cfDNA protectant to perform whole blood extraction; place the blood collection tube containing the blood sample into a centrifuge and centrifuge initially (1500g-1600g, 10-15min) to separate the plasma.
[0011] Step (2): Add proteinase K and ACL lysis buffer to the obtained plasma, mix thoroughly, and incubate at a constant temperature to lyse the protein;
[0012] Step (3) Use magnetic beads to bind cfDNA, wash multiple times to remove impurities, and then elute cfDNA with low-salt elution buffer.
[0013] Step (4): Collect the eluent, use a fluorescence quantitative analyzer to detect the concentration, and use a bioanalyzer to detect the fragment size distribution of cfDNA.
[0014] Preferably, the cfDNA protectant in step (1) is a nuclease inhibitor; the cfDNA protectant has the ability to maintain the structural integrity of nucleated cells and can effectively inhibit cell lysis under in vitro conditions, preventing the release of genomic DNA. Simultaneously, the cfDNA protectant can inhibit the degradation of cfDNA by nucleases in plasma, thereby maintaining the structural stability and representativeness of cfDNA in the pre-analysis stage. The nuclease inhibitor is diethyl pyrocarbonate or guanidine isothiocyanate; whole blood can also be collected directly using Streck Cell-Free DNA BCT® tubes.
[0015] Preferably, in step (3), the magnetic beads are magnetic microspheres with surface-modified carboxyl (-COOH) or silanol (-SiOH) groups, with a particle size ranging from 1 to 5 μm and a carboxyl density ≥300 μmol / g. The carboxyl magnetic beads bind to cfDNA via electrostatic interaction in a buffer solution at pH 5.0–6.5, while the silanol magnetic beads require a high-salt buffer solution containing 1.5 M NaCl and 20% PEG 8000 to bind. Preferred commercial products include Dynabeads™, MyOne™, Carboxylic Acid, or AMPure XP magnetic beads to improve the adsorption rate of cfDNA.
[0016] Preferably, the low-salt elution buffer in step (3) is TE buffer, pH 7.9-pH 8.1, and contains 0.04%-0.06% Tween 20 to reduce cfDNA adsorption loss.
[0017] This invention also provides the application of cfDNA in constructing an early lung cancer screening model (a method for constructing an early cancer screening model based on cfDNA), comprising the following steps:
[0018] S1: Use single-cell chromatin open data to identify cell types and extract features of epithelial cells;
[0019] S2: Library construction and high-throughput sequencing were performed on the extracted cfDNA samples, and cfDNA and epithelial cell features were fused using a variational autoencoder model.
[0020] S3: Use a convolutional neural network to train and test the fused cfDNA epithelial cell dimensionality reduction features of the variational autoencoder model (variational autoencoder), and output the classification probability associated with the cfDNA features.
[0021] As a further description of the above scheme: the steps for identifying the single-cell chromatin open data include:
[0022] We collected publicly available single-cell ATAC-seq sequencing datasets and used the CreateChromatinAssay function of the signac R package to create objects. The input file was fragments.tsv.gz. We ensured that each feature existed in at least 10 cells and that each cell had at least 200 features. We performed uniform quality control on each dataset and selected cells with a TSS score greater than 4 and a Tn5 fragment count greater than 1000. We then used Harmony batch effect correction.
[0023] The RunTFIDF function in the signac R package was used to standardize the data, the FindTopFeatures function was used to filter out highly variable regions, the RunSVD function was used to reduce the dimensionality of the data, and the RunUMAP function was used to reduce the dimensionality using 'lsi' singular value decomposition, compressing the dimensionality while retaining key biological information.
[0024] Based on the dimensionality-reduced data, cluster analysis was performed using the FindNeighbors and FindClusters functions of the signac R package to classify cells into epithelial cells, T cells, B cells, myeloid cells, fibroblasts, and endothelial cells, constructing an expression reference map of multiple cell groups, providing a reliable epithelial cell origin background for subsequent cfDNA feature signal extraction;
[0025] Open chromatin peaks were identified in epithelial cells using MACS2 to determine the peak position; ATAC-seq signals 200 bp upstream and downstream of the peak were used as the numerator; background signals in the 1k-3k range upstream and downstream of the peak were used as the denominator; the openness of each peak was quantified by the ratio of numerator to denominator.
[0026] The 2000-2200 peaks that showed the most significant differences from other cell populations were selected as characteristics of the epithelial cell-specific open regions.
[0027] Calculate the coverage of cfDNA in epithelial cell-specific open region features.
[0028] As a further description of the above scheme: the quality control standards include: each cell promoter region score is above 4, and the number of fragments in each cell is above 1000, excluding low-quality or apoptotic cells.
[0029] As a further description of the above scheme: the variational autoencoder model is constructed in two parts: an encoder and a latent space.
[0030] The inputs are the coverage of cfDNA in the epithelial cell region and the epithelial cell opening characteristics;
[0031] The encoder's formula is:
[0032] ;
[0033] Input data;
[0034] : latent variable;
[0035] : By encoder parameters The defined approximate posterior distribution;
[0036] The average value of the encoder output;
[0037] : Variance of encoder output;
[0038] During the coding stage, cfDNA features and epithelial cell features are nonlinearly compressed and fused into a shared latent space representation.
[0039] Introduce a KL divergence regularization term into the latent space to constrain the latent variables to conform to a Gaussian distribution;
[0040] This means the encoder will input. The mapping is a Gaussian distribution, i.e.:
[0041] ;
[0042] if The KL divergence has an analytical solution;
[0043] ;
[0044] This represents the mean of the i-th latent variable;
[0045] This represents the variance of the i-th latent variable;
[0046] Dimensions representing potential space;
[0047] Represents logarithmic variance;
[0048] Represents the standard normal distribution. I Represents the identity matrix;
[0049] The dimensionality-reduced data (cfDNA epithelial cell dimensionality-reduced features) is input into the convolutional neural network for model training supervision;
[0050] During training, the reconstruction error and KL divergence are jointly optimized to obtain stable and biologically meaningful potential dimensionality reduction features.
[0051] As a further description of the above scheme: the convolutional neural network includes a model structure of an input layer, multiple convolutional layers, activation function layers, pooling layers, fully connected layers, and an output layer; the input layer is the dimensionality reduction feature of cfDNA epithelial cells output by the variational autoencoder model, the convolutional layers extract local spatial features from the dimensionality reduction feature data of cfDNA epithelial cells, and the pooling layers perform downsampling to improve the computational efficiency and robustness of the model;
[0052] During the training phase, labeled cfDNA samples (lung cancer / non-lung cancer cfDNA data) are input into the model for forward and backward propagation, and the model weights are updated using the cross-entropy loss function and the Adam optimizer; the labels are lung cancer patients or healthy controls.
[0053] A convergent model is obtained through multiple rounds of iterative training, while a Dropout layer and an early stopping mechanism are introduced to prevent overfitting.
[0054] The accuracy, recall, precision, F1 score, and ROC curve performance of the model are evaluated using cross-validation and validation sets on the training dataset.
[0055] Accuracy = ;
[0056] Recall rate = ;
[0057] Accuracy = ;
[0058] F1 = 2 * (Recall * Precision) / (Recall + Precision);
[0059] TP represents the number of true positives, FP represents the number of false positives, FN represents the number of false negatives, and TN represents the number of true negatives.
[0060] The output layer of the convolutional neural network module generates probabilistic classification results through the Softmax function.
[0061] The present invention also provides a computer system for constructing an early lung cancer screening model, comprising:
[0062] The single-cell data processing module is used to extract epithelial cell-specific open region features from single-cell chromatin openness data;
[0063] The cfDNA sequencing data processing module is used to acquire and preprocess genomic distribution data of cfDNA fragments and calculate the coverage or distribution of cfDNA in epithelial cell-specific open region features.
[0064] The variational autoencoder module is used to perform nonlinear dimensionality reduction and fusion of the cfDNA fragment distribution data and epithelial cell features to generate a latent spatial feature representation.
[0065] The convolutional neural network module is used to train a classification model based on the latent spatial feature representation and output a classification probability signal associated with the cfDNA feature.
[0066] The classification probability signal is used for one of the following non-diagnostic purposes:
[0067] Screening studies of cancer-related biomarkers;
[0068] Bioinformatics analysis of tumorigenesis mechanisms;
[0069] Dynamic monitoring of the efficacy of candidate compounds during drug development.
[0070] The features of this invention are as follows: First, cfDNA is extracted from patient plasma. Then, low-depth whole-genome sequencing is performed on the cfDNA. Combined with epithelial cell-specific open regions mined from single-cell chromatin openness data, a feature matrix for lung cancer identification is constructed. Based on this, a deep learning model, variational autoencoder, and convolutional neural network are used to train and predict the distribution patterns of cfDNA fragments in epithelial cell feature regions, achieving highly sensitive, low-cost, and non-invasive early detection of lung cancer. The detailed process is as follows:
[0071] S1 identifies epithelial cell characteristics (i.e., open chromatin regions) from single-cell ATAC-seq data, providing a reference for subsequent analysis. This includes: data collection, quality control, and screening for high-quality single-cell data. Data standardization, dimensionality reduction, and clustering were performed to classify cells into six categories (including epithelial cells). Epithelial cell-specific open chromatin regions (characteristic intervals) were extracted, and the coverage of cfDNA in these intervals was calculated.
[0072] Epithelial cell features are chromatin open regions (200 bp upstream and downstream of the peak) unique to epithelial cells. These regions are significantly open in epithelial cells but not in other cell types. The selected 2000-2200 features are the most significantly different chromatin open regions between epithelial cells and other cell types, used for subsequent cfDNA alignment and analysis.
[0073] S2 fuses cfDNA features with epithelial cell features to extract potential high-dimensional features. cfDNA features are coverage signals in sequencing data corresponding to open regions of epithelial cell chromatin. The purpose of fusion is to associate cfDNA coverage signals with specific open regions of epithelial cells to extract potential biomarkers. A variational autoencoder (VAE) is used to nonlinearly compress the cfDNA data and epithelial cell features, fusing them into a shared latent space.
[0074] S3 uses fused features to train a convolutional neural network (CNN) to distinguish between lung cancer patients and healthy controls. The CNN model is constructed with input features that are the latent space representation of VAE fusion, containing association information between cfDNA and epithelial cell features. Model performance is evaluated through training and validation.
[0075] Compared with the prior art, the present invention has the following beneficial effects:
[0076] 1) The cfDNA extraction method provided by this invention has a high yield, and the obtained cfDNA has high purity and high stability;
[0077] 2) This invention, based on single-cell chromatin openness mapping, accurately constructs the open region characteristics of lung cancer-related epithelial cells, enabling the decoding of cfDNA signals at the cellular origin level. Compared to traditional fragmentomics methods, it can combine cellular characteristic signals to enhance the inference of tumor origin, effectively overcome the interference caused by tissue heterogeneity, and improve the resolution of cfDNA analysis.
[0078] 3) This invention introduces a variational autoencoder (VAE) model to deeply fuse cfDNA fragment features with epithelial cell feature signals, automatically learn the distribution patterns in the latent space, improve the stability and generalization ability of feature representation, and provide high-quality input for subsequent discrimination modeling.
[0079] 4) This invention uses low-depth whole-genome sequencing combined with artificial intelligence deep learning algorithms, which greatly reduces detection costs compared to traditional high-depth sequencing or targeted panels. Attached Figure Description
[0080] Figure 1 This is an image showing information about the extracted cfDNA fragments.
[0081] Figure 2 For the distribution of different cell populations and the expression of tag genes: A) A total of 64 publicly available single-cell ATAC-seq sequencing datasets were collected, covering 28 cancer types, 622 patients, and 1141 tumor samples; B) After further dimensionality reduction and cluster analysis, 183,434 high-quality metacells were finally obtained, which were classified into 47 cell subpopulations, covering 6 major cell types; C) Further refined annotation of each cell type.
[0082] Figure 3 To illustrate the differences in chromatin accessibility among different types of cfDNA samples in epithelial cell types;
[0083] Figure 4 A schematic diagram for training cfDNA epithelial cell features using variational autoencoder models and convolutional neural networks;
[0084] Figure 5 The training results are for the cfDNA epithelial cell feature model.
[0085] Figure 6 The gradient boosting model (GBM) was trained after feature extraction for cfDNA TSS coverage. Detailed Implementation
[0086] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. In the following description, specific details such as specific configurations and components are provided only to help fully understand the embodiments of this application. In addition, in all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to actual scale;
[0087] It should be understood that the phrase "an embodiment" or "this embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "an embodiment" or "this embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0088] Furthermore, reference numerals and / or letters may be repeated in different examples within this application. Such repetition is for the purpose of simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or settings discussed.
[0089] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist simultaneously. The term " / and" in this article describes another type of relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " in this article generally indicates that the related objects before and after it are in an "or" relationship.
[0090] In this article, the term "at least one" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, "at least one of A and B" can mean: A exists alone, A and B exist simultaneously, or B exists alone.
[0091] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion.
[0092] The signac R package used in the following examples is a bioinformatics tool developed specifically for analyzing single-cell chromatin accessibility data (such as scATAC-seq).
[0093] Example 1
[0094] A method for extracting cfDNA includes the following steps:
[0095] Step (1) Use EDTA vessels containing anticoagulants and cfDNA protectants (using nuclease inhibitors such as diethyl pyrocarbonate or guanidine isothiocyanate); collect peripheral blood from the subject (such as a healthy individual or a cancer patient), usually in the form of 2 mL.
[0096] Step (2): Centrifuge at 1600 g for 10 minutes at room temperature, transfer the supernatant (plasma) to a new tube, centrifuge at 16,000 g for 10 minutes at room temperature to further remove residual cell debris; collect the final plasma in a tube free of DNase contamination, extract cfDNA immediately, or store at -80℃; centrifuge again, collect the final filtrate, determine the DNA concentration, and detect the distribution of cfDNA fragments, add lysis buffer (containing proteinase K (Qiagen#19131) and ACL buffer (Qiagen#939016)) to lyse plasma proteins;
[0097] Step (3) Add magnetic beads to promote cfDNA binding, wash multiple times to remove impurities; use low-salt elution buffer to elute cfDNA from magnetic beads or silica gel column, with a final elution volume of about 20-50 µL.
[0098] The magnetic beads are magnetic microspheres with surface-modified carboxyl (-COOH) or silanol (-SiOH) groups, ranging in size from 1 to 5 μm, and a carboxyl density ≥300 μmol / g. The carboxyl magnetic beads bind to cfDNA via electrostatic interaction in a buffer solution at pH 5.0–6.5, while the silanol magnetic beads require a high-salt buffer solution containing 1.5 M NaCl and 20% PEG 8000 for binding. Preferred commercial products include Dynabeads™, MyOne™, Carboxylic Acid, or AMPure XP magnetic beads to improve cfDNA adsorption rates.
[0099] The low-salt elution buffer is a TE buffer with a pH of 7.9-8.1 and contains 0.04%-0.06% Tween 20 to reduce cfDNA adsorption loss.
[0100] Step (4): Collect the eluent, detect the concentration using a fluorescence quantitative analyzer, and analyze the fragment size distribution of cfDNA using a QSEP bioanalyzer. The results are as follows: Figure 1 As shown. Figure 1 As shown, the median extracted cfDNA fragment was 182 bp, and the median concentration was 10 ng / uL.
[0101] Example 2
[0102] S1: Use single-cell chromatin open data to identify cell types and extract features of epithelial cells;
[0103] S2: Library construction and high-throughput sequencing were performed on the extracted cfDNA samples, and cfDNA and epithelial cell features were fused using a variational autoencoder model.
[0104] S3: Use convolutional neural networks to train and test cfDNA data.
[0105] In a preferred embodiment, step S1 may include:
[0106] S101. This invention first collects 64 publicly available single-cell ATAC-seq sequencing datasets, covering 28 cancer types, 622 patients, and 1,141 tumor samples;
[0107] The 64 publicly available datasets are from the GEO database, and their data numbers are shown in Table 1.
[0108] Table 1. Data IDs of 64 Public Datasets
[0109]
[0110] S102. Use the CreateChromatinAssay function of the signac R package to create objects. The input file is fragments.tsv.gz. Ensure that each feature exists in at least 10 cells and that each cell has at least 200 features. Perform uniform quality control on each dataset and select cells with a TSS score greater than 4 and a Tn5 fragment count greater than 1000. Then, use Harmony batch effect correction, setting the Harmony dimension to 1:30 and theta value to 2.
[0111] S103. Cell quality control standards include: each cell promoter region score is above 4, and each cell has more than 1000 fragments, excluding low-quality or apoptotic cells;
[0112] S104. Finally, 3,904,694 high-quality single-cell data were obtained through screening for subsequent analysis.
[0113] S105. The RunTFIDF function (default parameters) of the signac R package is used to standardize the data, the FindTopFeatures function (default parameters) is used to filter the data for highly variable regions, the RunSVD function (default parameters) is used to reduce the dimensionality of the data, and the RunUMAP function uses 'lsi' singular value decomposition to reduce the dimensionality to 2:30, compressing the dimensionality while retaining key biological information.
[0114] S106. Based on the dimensionality-reduced data, cluster analysis was performed using the FindNeighbors and FindClusters functions (default parameters) from the signac R package to divide cells into subpopulations with potential biological significance. All cells were classified into six main cell types, including: epithelial cells, T cells, B cells, myeloid cells, fibroblasts, and endothelial cells. Among them, T cells include CD4+ cells. + Regulatory T cells, CD4 + Naïve T cells, CD8 + Exhausted T cells, CD8 + Cytotoxic T cells, double-positive tissue-resident memory T cells, and NKT cells; B cells include: memory B cells, naive B cells, and plasma cells; myeloid cells include: macrophages, dendritic cells, myeloid suppressor cells, conventional dendritic cells, plasmacytoid dendritic cells, and mast cells;
[0115] S107. Through the above-mentioned fine annotation, an expression reference map covering multiple cell groups in various cancer environments is constructed, providing a reliable epithelial cell-derived background for subsequent cfDNA feature signal extraction;
[0116] S108. Align the cfDNA to the epithelial cell characteristic regions in single-cell ATAC-seq. The epithelial cell characteristic regions in single-cell ATAC-seq are defined by using 200bp upstream and downstream of the peak identified by macs2 as the numerator, and the denominator is the fragment ratio of 1k-3k upstream and downstream of the peak. Select the 2000-2200 features that are most significantly different from epithelial cells and other cells as the epithelial cell characteristic regions (the open chromatin regions unique to epithelial cells).
[0117] S109. Calculate the coverage of cfDNA in the epithelial cell characteristic region.
[0118] In a preferred embodiment, a variational autoencoder model is used to fuse cfDNA features and epithelial cell features. Detailed operations include:
[0119] S201. Construct a variational autoencoder structure, including two main modules: encoder and latent space. The input is the coverage of cfDNA in the epithelial cell characteristic region and the epithelial cell characteristic region.
[0120] The encoder has 5 coding layers, and its formula is:
[0121] ;
[0122] Input data (cfDNA data and single-cell ATAC-seq data);
[0123] : latent variable;
[0124] From encoder parameters The defined approximate posterior distribution;
[0125] The average value of the encoder output;
[0126] : Variance of encoder output.
[0127] S202. The Python package PyTorch is used for encoding. The cfDNA features and epithelial cell features are nonlinearly compressed separately. The functions torch.nn.Linear, torch.nn.ReLU, and torch.nn.Sequential (default parameters) are used for processing and then fused into the shared latent space representation.
[0128] S203. Introduce a KL divergence regularization term into the latent space and use the Python package PyTorch to calculate the KL divergence, constraining the latent variables to conform to a Gaussian distribution to enhance generalization ability.
[0129] Introduce a KL divergence regularization term into the latent space to constrain the latent variables to conform to a Gaussian distribution;
[0130] This means the encoder will input. The mapping is a Gaussian distribution, i.e.:
[0131] ;
[0132] if
[0133] ;
[0134] This represents the mean of the i-th latent variable;
[0135] This represents the variance of the i-th latent variable;
[0136] Dimensions representing potential space;
[0137] Represents logarithmic variance;
[0138] Represents the standard normal distribution. I Represents the identity matrix.
[0139] S204. During training, the reconstruction error and KL divergence are jointly optimized to obtain stable and biologically meaningful potential dimensionality reduction features.
[0140] The steps for establishing the use of convolutional neural networks to train and test cfDNA data include:
[0141] S301. Design a convolutional neural network model structure. The model should include at least: an input layer, multiple convolutional layers (Conv), an activation function layer (ReLU), a pooling layer (Pooling), a fully connected layer (FC), and an output layer. The input layer is the dimensionality reduction feature of cfDNA epithelial cells generated by the variational autoencoder model. The convolutional kernel is a 3×3 feature map with a stride of 1 and the pooling method is AvgPool.
[0142] S302: Convolutional layers extract local spatial features from cfDNA epithelial cell dimensionality reduction feature data, and pooling layers implement downsampling to improve model computational efficiency and robustness.
[0143] S303. During the training phase, labeled cfDNA samples (lung cancer patients vs. healthy controls) are input into the model for forward and backward propagation. The model weights are updated using the cross-entropy loss function and the Adam optimizer. The learning rate is 0.001, the batch size is 128, and the number of iterations is 30.
[0144] S304. A convergent model is obtained through multiple rounds of iterative training. At the same time, a Dropout layer and an early stopping mechanism are introduced to prevent overfitting. The Dropout rate is 0.2.
[0145] S305. Evaluate the model’s accuracy, recall, F1 score, and ROC curve performance using cross-validation and validation sets on the training dataset.
[0146] Accuracy = ;
[0147] Recall rate = ;
[0148] Accuracy = ;
[0149] F1 = 2 * (Recall * Precision) / (Recall + Precision);
[0150] TP represents the number of true positives, FP represents the number of false positives, FN represents the number of false negatives, and TN represents the number of true negatives.
[0151] Validation experiment of lung cancer early screening model
[0152] Sample information:
[0153] The training set consisted of 778 samples (410 lung cancer patients, 113 patients with benign nodules, and 255 healthy controls) and the validation set consisted of 259 samples (137 lung cancer patients, 37 patients with benign nodules, and 85 healthy controls), with clinical stage distribution (454 stage I lung cancer patients, 51 stage II patients, 41 stage III patients, and 1 stage IV patient).
[0154] The inclusion criteria for patients with benign pulmonary nodules were supplemented. The nodules were confirmed as benign lesions by pathology or clinical follow-up, including but not limited to the following types: pathologically diagnosed as inflammatory pseudotumor, granuloma, tuberculous nodule, histiocytic proliferative lesion, etc., and no significant growth or malignant transformation was observed after more than 2 years of imaging follow-up, and the nodules were diagnosed as stable benign nodules.
[0155] Independent validation: cfDNA samples were obtained from Peking University Cancer Hospital, Peking Union Medical College Hospital, Yunnan Provincial People's Hospital, Yunnan Cancer Hospital, and Zhejiang Cancer Hospital. Data from multiple hospitals demonstrate the model's generalization ability.
[0156] Threshold setting: This specifies the confidence threshold for the model to determine "positive for lung cancer" (e.g., probability > 0.85).
[0157] Example 3
[0158] This embodiment is based on the above embodiment 2, and the similarities with embodiment 2 will not be repeated.
[0159] This example describes the training of a cfDNA epithelial cell feature model.
[0160] Step 1: Distribution of epithelial cell groups and expression of tag genes; results are attached. Figure 2 As shown.
[0161] This embodiment collected 64 publicly available single-cell ATAC-seq sequencing datasets, covering 28 cancer types, 622 patients, and 1141 tumor samples. Figure 2 A). During data integration, quality control and batch effect correction were first performed on each dataset, and standardized Seurat objects were constructed. Cell quality control criteria included: selecting cells with a TSS score greater than 4 and a Tn5 fragment number higher than 1000 ( Figure 2A). A total of 3,904,694 cells passed quality control and were included in subsequent analyses. During the data integration phase, a cell network was constructed based on the quality-controlled cell expression matrix. The k-nearest neighbor algorithm (k=10) was used to estimate inter-cell similarity, and homogeneous cells were aggregated into metacells using γ-granularity clustering (γ=20). The average gene expression within each metacell was then calculated, generating an expression matrix containing 195,238 metacells.
[0162] Figure 2 Further dimensionality reduction and cluster analysis yielded 183,434 high-quality metacells, categorized into 47 cell subpopulations covering 6 major cell types: 56,686 epithelial cells, 53,631 T cells, 17,290 B cells, 29,164 myeloid cells, 17,459 fibroblasts, and 9,204 endothelial cells. Figure 2 B). Building on this, we further refined the annotation of various cell types, including: T cells: 6,890 CD4+ cells. + Regulatory T cells (Tregs), 16,068 CD4+ + Naïve T cells, 4,515 CD8 + Exhausted T cells, 16,481 CD8 + Cytotoxic T cells, 135 NK cells, and 8,841 double-positive tissue-resident memory T cells; B cells: 898 memory B cells, 8,954 naive B cells, and 7,438 plasma cells; Myeloid cells: 16,835 macrophages, 625 dendritic cells (DCs), 4,984 myeloid suppressor cells (MDSCs), 3,658 conventional dendritic cells (cDCs), 1,605 plasmacytoid dendritic cells (pDCs), and 1,457 mast cells, epithelial cells, endothelial cells, and fibroblasts. Figure 2 C).
[0163] In summary, this application constructed a pan-cancer single-cell atlas containing nearly four million cells based on multi-cancer data, providing a high-resolution cell atlas foundation for a non-invasive early cancer screening method and system based on cell type-specific chromatin-open cfDNA. Figure 2 A and Figure 2 C).
[0164] Step 2: Distribution of cfDNA epithelial cell characteristics in non-small cell lung cancer, benign nodules, and healthy samples. Results are attached. Figure 3 As shown.
[0165] To investigate the differences in epithelial cells among healthy peripheral blood, lung cancer patients, and benign nodule samples, this application observed that cfDNA within epithelial cell features had significantly lower mean coverage in benign nodule and lung cancer samples, while coverage was higher in healthy individuals.
[0166] Step 3: cfDNA epithelial cell feature variational autoencoder model and convolutional neural network model, as shown in the appendix. Figure 4 As shown.
[0167] The epithelial cell features of cfDNA are extracted. After extraction, a variational autoencoder (VAE) model is used to reduce the dimensionality of the data. The encoder is a three-layer fully connected network that reduces the dimensionality of the input to represent the latent variable distribution. The loss function is a weighted sum of reconstruction error (mean squared error) and KL divergence, used to guide the learning of potential fusion expression features. The latent variables learned by the VAE are represented in a format suitable for convolutional layers and input into a convolutional neural network (CNN). This is followed by convolutional layers, max pooling layers, dropout layers (to prevent overfitting), and fully connected layers that output softmax classification probabilities (lung cancer / non-lung cancer). Cross-entropy is used as the loss function, and the Adam optimizer is used for backpropagation. The VAE latent variables are further concatenated with the intermediate features of the CNN classification layer and used as input to a multilayer perceptron (MLP). A two-layer fully connected network is used to complete the final lung cancer prediction. The MLP network is used to integrate low-dimensional features and deep abstract features to improve the overall prediction performance. Finally, a confidence score is output to determine whether the cfDNA sample belongs to a lung cancer patient or a healthy individual.
[0168] Step 4: Feature extraction and model training of cfDNA TSS coverage; results are attached. Figure 5 As shown, by using cfDNA combined with epithelial cell features for variational autoencoder dimensionality reduction and then training the model with a convolutional neural network, this method can identify the risk of lung cancer with high accuracy. Its AUC reaches 0.91, accuracy 0.95, and recall 0.87, thus providing patients with precise diagnostic and treatment guidance and helping to reduce the risk of death.
[0169] Step 5: After feature extraction from cfDNA TSS coverage, a gradient boosting model (GBM) was used for training. The results are shown in the attached figure. Figure 6 As shown, the gradient boosting model has an AUC of 0.81, accuracy of 0.92, and recall of 0.7. Compared to deep learning models, its detection performance is significantly lower.
[0170] Example 4
[0171] A computer system for constructing an early lung cancer screening model, comprising:
[0172] The single-cell data processing module is used to extract epithelial cell-specific open region features from single-cell chromatin openness data;
[0173] The cfDNA sequencing data processing module is used to acquire and preprocess genomic distribution data of cfDNA fragments;
[0174] The variational autoencoder module is used to perform nonlinear dimensionality reduction and fusion of the cfDNA fragment distribution data and epithelial cell features to generate a latent spatial feature representation.
[0175] The convolutional neural network module is used to train a classification model based on the latent spatial feature representation and output a classification probability signal associated with the cfDNA feature.
[0176] The classification probability signal is used for one of the following non-diagnostic purposes:
[0177] Screening studies of cancer-related biomarkers;
[0178] Bioinformatics analysis of tumorigenesis mechanisms;
[0179] Dynamic monitoring of the efficacy of candidate compounds during drug development.
[0180] This embodiment is based on the above embodiment 2, and the similarities with embodiment 2 will not be repeated.
[0181] The epithelial cell-specific open region characteristics were obtained through the following steps: cluster analysis was performed on single-cell ATAC-seq data to screen for chromatin open regions that were significantly different from other cell types in epithelial cells; the coverage ratio of cfDNA fragments in the regions was calculated.
[0182] The variational autoencoder module generates a latent variable representation that conforms to a Gaussian distribution by jointly optimizing the reconstruction error and KL divergence.
[0183] The output layer of the convolutional neural network module generates probabilistic classification results through the Softmax function.
[0184] The above description is merely a preferred embodiment of the present invention and does not limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any changes, modifications, substitutions, integrations, and parameter alterations to these embodiments within the spirit and principles of the present invention, achieved through conventional substitutions or by achieving the same function without departing from the principles and spirit of the present invention, fall within the scope of protection of the present invention.
Claims
1. A method for constructing an early lung cancer screening model, characterized in that, Includes the following steps: S1: Use single-cell chromatin open data to identify cell types and extract features of epithelial cells; S2: Library construction and high-throughput sequencing were performed on the extracted cfDNA samples, and cfDNA and epithelial cell features were fused using a variational autoencoder model. The variational autoencoder model is constructed using two modules: an encoder and a latent space. The inputs are the coverage of cfDNA in the epithelial cell region and the epithelial cell opening characteristics. The encoder's formula is: ; x: Input data; z: Latent variable; : By encoder parameters φ The defined approximate posterior distribution; The average value of the encoder output; : Variance of encoder output; During the coding stage, the coverage of cfDNA in the epithelial cell region and the epithelial cell openness feature are nonlinearly compressed and fused into the shared latent space representation. Introduce a KL divergence regularization term into the latent space to constrain the latent variables to conform to a Gaussian distribution; This means that the encoder maps the input x to a Gaussian distribution, that is: ; if ; ; This represents the mean of the i-th latent variable; This represents the variance of the i-th latent variable; Dimensions representing potential space; Represents logarithmic variance; Represents the standard normal distribution. I The identity matrix is represented by the dimensionality-reduced data, which is then input into the convolutional neural network for training and supervision of the model. During training, the reconstruction error and KL divergence are jointly optimized to obtain stable, biologically meaningful potential dimensionality reduction features; S3: Use a convolutional neural network to train and test the dimensionality reduction features of cfDNA epithelial cells after the fusion of variational autoencoder models, and output the classification probability associated with the cfDNA features; The steps for identifying cell types using the single-cell chromatin open data include: We collected publicly available single-cell ATAC-seq sequencing datasets and used the CreateChromatinAssay function of the signac R package to create objects. The input file was fragments.tsv.gz. We ensured that each feature existed in at least 10 cells and that each cell had at least 200 features. We performed uniform quality control on each dataset and selected cells with a TSS score greater than 4 and a Tn5 fragment count greater than 1000. We then used Harmony batch effect correction. The RunTFIDF function in the signac R package was used to standardize the data, the FindTopFeatures function was used to filter out highly variable regions, the RunSVD function was used to reduce the dimensionality of the data, and the RunUMAP function was used to reduce the dimensionality using 'lsi' singular value decomposition, compressing the dimensionality while retaining key biological information. Based on the dimensionality-reduced data, cluster analysis was performed using the FindNeighbors and FindClusters functions of the signac R package to classify cells into epithelial cells, T cells, B cells, myeloid cells, fibroblasts, and endothelial cells, constructing an expression reference map of multiple cell groups, providing a reliable epithelial cell origin background for subsequent cfDNA feature signal extraction; Open chromatin peaks were identified in epithelial cells using MACS2, and the peak positions were determined. The ATAC-seq signals 200 bp upstream and downstream of the peak were used as the numerators, and the background signals in the 1k-3k range upstream and downstream of the peak were used as the denominators. The degree of openness of each peak was quantified by the ratio of numerator to denominator. The 2000-2200 peaks that showed the most significant differences from other cell populations were selected as characteristics of the epithelial cell-specific open regions. Calculate the coverage of cfDNA in epithelial cell-specific open region features.
2. The method according to claim 1, characterized in that, The cfDNA extraction method includes the following steps: Step (1): Whole blood extraction is performed using a special blood collection tube pre-loaded with cfDNA protectant; the blood collection tube containing the blood sample is placed in a centrifuge and pre-centrifuged to separate the plasma; Step (2): Add proteinase K and ACL lysis buffer to the obtained plasma, mix thoroughly, and incubate at a constant temperature to lyse the protein; Step (3): Use magnetic beads to bind cfDNA, wash repeatedly to remove impurities, and then elute cfDNA with low-salt elution buffer; the magnetic beads are magnetic microspheres with carboxyl or silanol groups on the surface to adsorb cfDNA, with a particle size range of 1-5 μm and a carboxyl density ≥300 μmol / g. Step (4): Collect the eluent, use a fluorescence quantitative analyzer to detect the concentration, and use a bioanalyzer to detect the fragment size distribution of cfDNA.
3. The method according to claim 1, characterized in that, Cell quality control standards include: each cell promoter region score must be above 4, and each cell must have more than 1000 fragments, excluding low-quality or apoptotic cells.
4. The method according to claim 1, characterized in that, The convolutional neural network includes a model structure consisting of an input layer, multiple convolutional layers, activation function layers, pooling layers, fully connected layers, and an output layer. The input layer is the dimensionality-reduced cfDNA epithelial cell feature output by the variational autoencoder model. The convolutional layers extract local spatial features from the dimensionality-reduced cfDNA epithelial cell feature data, and the pooling layers perform downsampling to improve the model's computational efficiency and robustness. During the training phase, tagged cfDNA samples were input into the model for forward and backward propagation, and the model weights were updated using the cross-entropy loss function and the Adam optimizer; the tags were lung cancer patients or healthy controls. A convergent model is obtained through multiple rounds of iterative training, while a Dropout layer and an early stopping mechanism are introduced to prevent overfitting. The accuracy, recall, precision, F1 score, and ROC curve performance of the model are evaluated using cross-validation and validation sets on the training dataset. Accuracy = ; Recall rate = ; Accuracy = ; F1 = 2 * (Recall * Precision) / (Recall + Precision); TP represents the number of true positives, FP represents the number of false positives, FN represents the number of false negatives, and TN represents the number of true negatives.
Citation Information
Patent Citations
Non-invasive cancer early screening system based on cfDNA omics characteristics
CN113160889A