CfDNA extraction method and application of cfDNA extraction method in construction of lung cancer early screening model

By combining variational autoencoder models and convolutional neural networks with single-cell chromatin openness data, we identified epithelial cell-specific cfDNA region features and constructed a low-cost, high-precision early lung cancer screening model. This solves the problem of insufficient early screening methods in existing technologies and achieves highly sensitive early detection of lung cancer.

CN120905208AActive Publication Date: 2025-11-07KUNMING MEDICAL UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511451264.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2025-11-07
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

Existing technologies lack sensitive, low-cost, and non-invasive early screening methods, making it difficult to achieve dynamic monitoring and early detection of lung nodules in high-risk populations. The fragmented pattern of cfDNA has not been fully utilized in tumor detection.

Method used

By combining variational autoencoder models and convolutional neural networks with single-cell chromatin open data, and using cfDNA extraction methods to identify epithelial cell-specific cfDNA region features from blood samples, a low-cost, high-precision early lung cancer screening model was constructed.

Benefits of technology

It has achieved a highly sensitive, low-cost, non-invasive early detection of lung cancer, improving the accessibility and accuracy of screening and reducing testing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120905208A_ABST
    Figure CN120905208A_ABST
Patent Text Reader

Abstract

The invention discloses a cfDNA extraction method and application of the cfDNA extraction method in construction of a lung cancer early screening model. The method comprises the following steps: firstly, extracting cfDNA in plasma of a patient, then carrying out whole genome sequencing on the cfDNA, and constructing a characteristic matrix for lung cancer identification by combining an epithelial cell specific open region mined by single cell chromatin open data. On the basis, a deep learning model variational self-coding and a convolutional neural network are utilized to train and predict a distribution mode of cfDNA fragments in an epithelial cell characteristic region, and high-sensitivity, low-cost and non-invasive early detection of the lung cancer is realized. According to the invention, cfDNA epithelial cell signals are creatively used as cancer recognition signals, so that the screening accuracy and coverage rate of early-stage lung cancer patients are remarkably improved. The technology is suitable for large-scale cancer population screening and high-risk population dynamic monitoring, and has wide clinical popularization and application values.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of DNA extraction, and particularly relates to a cfDNA extraction method and application thereof in constructing a lung cancer early screening model. BACKGROUND

[0002] Lung cancer is one of the highest incidence and mortality rates of cancer in the world, and non-small cell lung cancer (NSCLC) accounts for about 85% of all lung cancer cases. Most patients are in the advanced stage at the time of diagnosis, and the five-year survival rate is extremely low. According to research, if effective intervention is carried out in the early stage of lung cancer, the five-year survival rate can reach about 90%, while the five-year survival rate of patients in the advanced stage is only 2%~5% even if they receive treatment. As can be seen, dynamic monitoring and early detection of potential high-risk populations are of great significance for improving the clinical prognosis of non-small cell lung cancer.

[0003] At present, there is still a lack of a high-sensitivity, low-cost and non-invasive early screening method in the clinic, which can monitor the changes of lung nodules in high-risk populations in real time. In recent years, plasma free DNA (cfDNA) has attracted widespread attention as a non-invasive detection marker. Studies have found that the cell characteristic signals reflected in cfDNA can effectively reflect the active state of epithelial cells and indicate potential abnormalities in the body. The cfDNA of normal individuals mainly comes from peripheral blood mononuclear cells (PBMCs), therefore, identifying the non-PBMC-derived components in cfDNA, especially fragments from tumor epithelial cells, is expected to serve as a new molecular indicator for tumor early screening.

[0004] During tumor progression, the abnormal proliferation and apoptosis of cancer cells release cfDNA with features of open chromatin into the blood circulation. These fragments reflect the chromatin state of epithelial cells and can be used as an early tumor recognition signal. The fragmentation pattern of cfDNA is regulated by chromatin conformation and nucleosome distribution. In open chromatin regions, DNA is more easily cleaved by nucleases due to fewer nucleosomes, resulting in significantly lower coverage of cfDNA in these regions. Conversely, nucleosome-rich regions have more abundant cfDNA fragments due to structural protection. Literature (Snyder, M.W., Kircher, M., Hill, A.J., Daza, R.M., and Shendure, J. (2016). Cell-free DNA Comprises an In Vivo Nucleosome Footprint that Informs Its Tissues-Of-Origin. Cell 164, 57-68. 10.1016 / j.cell.2015.11.050) has successfully used the fragment distribution information of cfDNA to infer its tissue of origin, verifying the feasibility of this principle.

[0005] To more accurately analyze the chromatin open pattern hidden in cfDNA, single-cell level ATAC-seq (scATAC-seq) technology provides an important tool for research. This technology can reveal the chromatin accessibility of different cell subpopulations at single-cell resolution and is widely used in tumor heterogeneity analysis. If the fragment coverage of cfDNA is combined with the epithelial cell chromatin map constructed by scATAC-seq, supplemented by the classification and recognition ability of deep learning algorithms, it is expected to efficiently extract the characteristic regions of tumor epithelial cells from the complex cfDNA mixed signal, thereby significantly improving the sensitivity and accuracy of early detection.

[0006] Based on the above research foundation, if a strategy can be established to realize early lung cancer risk prediction relying on low-depth whole genome sequencing, it will significantly reduce the screening cost while improving the screening popularity and accuracy. Therefore, developing a non-invasive early screening method that combines cfDNA and epithelial cell signal recognition has significant scientific value and application prospects, and also provides a potential breakthrough for the early management of high-risk lung cancer populations. SUMMARY

[0007] To overcome the deficiencies of the prior art, the application provides a method for constructing a cfDNA cancer early screening model based on deep learning and fusion of epithelial cell characteristic signals. The application identifies cfDNA region characteristics specific to epithelial cells from blood samples through a variational auto-encoding model and a convolutional neural network model, and provides a new analysis tool for cancer research. Based on these characteristics, specific probes can be designed to achieve low-cost and high-precision cfDNA detection.

[0008] It should be noted that the cfDNA extraction method and the method for constructing a cancer early screening model provided by the application are intended to realize the quantification and classification of cfDNA characteristics through bioinformatics analysis and machine learning technology, and the output result is a probabilistic risk assessment signal (such as lung cancer-related feature confidence), which is not for diagnostic or therapeutic purposes.

[0009] To solve the above technical problems, the technical scheme of the application is as follows: a cfDNA extraction method comprising the following steps: Step (1): a special blood collection tube preloaded with a cfDNA protective agent is used for whole blood extraction operation; the blood collection tube containing the blood sample is placed in a centrifugal device, and preliminary centrifugation (1500g-1600g, 10-15min) is performed to separate the plasma; Step (2): proteinase K and ACL lysis buffer are added to the obtained plasma, and after being mixed thoroughly, constant temperature incubation is performed to lyse the protein; Step (3): magnetic beads are used to combine cfDNA, and impurities are removed through multiple washings, and then low-salt elution buffer is used to elute cfDNA.

[0010] Step (4): the eluate is collected, the concentration is detected using a fluorescence quantification instrument, and the fragment size distribution of cfDNA is detected by a bioanalyzer.

[0011] As a preferred, the cfDNA protective agent in step (1) is a nuclease inhibitor; the cfDNA protective agent has the ability to maintain the integrity of the structure of nucleated cells, can effectively inhibit the lysis of cells under in vitro conditions, and prevent the release of genomic DNA. At the same time, the cfDNA protective agent can inhibit the degradation of cfDNA by nucleases in the plasma, thereby maintaining the structural stability and representativeness of cfDNA in the pre-analysis stage. The nuclease inhibitor is diethyl pyrocarbonate or guanidinium isothiocyanate; a Streck Cell-Free DNA BCT tube can also be directly used to collect whole blood.

[0012] As preferred, the magnetic beads used in step (3) are magnetic microspheres with carboxyl (-COOH) or silanol (-SiOH) surface modification, with a particle size range of 1-5 μm and a carboxyl density of ≥300 μmol / g. The carboxyl magnetic beads bind cfDNA in a buffer with a pH of 5.0-6.5 through electrostatic interaction, while the silanol magnetic beads need to be bound in a high-salt buffer containing 1.5 M NaCl and 20% PEG 8000. Preferred commercial products include Dynabeads™, MyOne™, Carboxylic Acid, or AMPure XP magnetic beads to improve the adsorption rate of cfDNA.

[0013] As preferred, the low-salt elution buffer used in step (3) is a TE buffer with a pH of 7.9-pH 8.1 and containing 0.04%-0.06% Tween 20 to reduce the loss of cfDNA adsorption.

[0014] The present application also provides the use of cfDNA in constructing a lung cancer early screening model (a method for constructing a cancer early screening model based on cfDNA), comprising the following steps: S1: identifying cell types using single-cell chromatin opening data and extracting epithelial cell features; S2: library construction and high-throughput sequencing of the extracted cfDNA sample, and using a variational autoencoder model to fuse cfDNA and epithelial cell features; S3: training and testing the variational autoencoder model (variational autoencoder) fused cfDNA epithelial cell dimension reduction features using a convolutional neural network, and outputting the classification probability associated with cfDNA features.

[0015] As a further description of the above scheme: the step of identifying cell types using single-cell chromatin opening data comprises: Collect publicly available single-cell ATAC-seq sequencing data sets, use the CreateChromatinAssay of the signac R package to create objects, input the file fragments.tsv.gz, ensure that each feature exists in at least 10 cells, and each cell has at least 200 features, perform uniform quality control on each data set, select cells with a TSS score greater than 4 and a Tn5 fragment number higher than 1000; then use Harmony batch effect correction processing; Use the RunTFIDF function of the signac R package to standardize the data, the FindTopFeatures function to filter high-variation regions, the RunSVD function to perform dimension reduction processing, and the RunUMAP function to use the 'lsi' singular value decomposition method for dimension reduction processing, compressing the dimension while retaining key biological information; Based on the reduced dimension data, the signac R package FindNeighbors function and FindClusters function were used for cluster analysis to divide the cells into epithelial cells, T cells, B cells, myeloid cells, fibroblasts and endothelial cells, and to construct a multi-cell class expression reference atlas to provide a reliable epithelial cell source background for subsequent cfDNA feature extraction; MACS2 was used to identify open chromatin peaks in epithelial cells (peak calling) to determine the peak position; the ATAC-seq signal of the upstream and downstream 200bp of the peak top was used as the numerator; the background region signal of the 1k-3k range of the peak top was used as the denominator; the ratio of the numerator to the denominator was used to quantify the openness of each peak; 2000-2200 peaks with the most significant differences from other cell groups were selected as epithelial cell-specific open region features; The coverage of cfDNA in the epithelial cell-specific open region features was calculated.

[0016] As a further description of the above scheme: the quality control standards include: the promoter region score of each cell is above 4, and the number of fragments of each cell is above 1000, excluding low-quality or apoptotic cells.

[0017] As a further description of the above scheme: the variational auto-encoding model is constructed using two parts: an encoder and a latent space module; The input is the coverage of cfDNA in the epithelial cell interval and the epithelial cell open feature; The formula of the encoder is: ; : input data; : latent variable; : the approximate posterior distribution defined by the encoder parameters ; : mean of the encoder output; : variance of the encoder output; In the encoding stage, cfDNA features and epithelial cell features are nonlinearly compressed and fused into a shared latent space representation; KL divergence regularization term is introduced in the latent space to constrain the latent variable to conform to the Gaussian distribution; This means that the encoder maps the input to a Gaussian distribution, i.e. ; If , KL divergence has an analytical solution;

[0018] ; represents the mean of the i-th latent variable; represents the variance of the i-th latent variable; represents the dimension of the latent space; represents the log variance; represents the standard normal distribution, I represents the identity matrix; The dimensionality-reduced data (cfDNA epithelial cell dimensionality-reduced features) is input into a convolutional neural network for training supervision of the model; During training, the reconstruction error and KL divergence are jointly optimized to obtain stable and biologically meaningful latent dimensionality-reduced features.

[0019] As a further description of the above scheme: the convolutional neural network comprises a model structure of an input layer, multiple convolutional layers, an activation function layer, a pooling layer, a fully connected layer, and an output layer; the input layer is the cfDNA epithelial cell dimensionality-reduced features output by the variational autoencoder model, the convolutional layers extract local spatial features from the cfDNA epithelial cell dimensionality-reduced feature data, and the pooling layer realizes down-sampling to improve the computational efficiency and robustness of the model; In the training phase, labeled cfDNA samples (lung cancer / non-lung cancer cfDNA data) are input into the model for forward propagation and back propagation, and the model weight is updated using the cross-entropy loss function and the Adam optimizer; the label is a lung cancer patient or a healthy control; Through multiple rounds of iterative training, a converged model is obtained, and a Dropout layer and an early stopping mechanism are introduced to prevent overfitting; In the training data set, cross-validation and validation set are used to evaluate the accuracy, recall rate, precision, F1 score, and ROC curve performance of the model; Accuracy = ; Recall rate = ; Precision = ; F1 = 2 * (recall rate * precision) / (recall rate + precision); TP is the number of true positives, FP is the number of false positives, FN is the number of false negatives, and TN is the number of true negatives.

[0020] The output layer of the convolutional neural network module generates probabilistic classification results through the Softmax function.

[0021] The present invention also provides a computer system for constructing an early lung cancer screening model, comprising: The single-cell data processing module is used to extract epithelial cell-specific open region features from single-cell chromatin openness data; The cfDNA sequencing data processing module is used to acquire and preprocess genomic distribution data of cfDNA fragments and calculate the coverage or distribution of cfDNA in epithelial cell-specific open region features. The variational autoencoder module is used to perform nonlinear dimensionality reduction and fusion of the cfDNA fragment distribution data and epithelial cell features to generate a latent spatial feature representation. The convolutional neural network module is used to train a classification model based on the latent spatial feature representation and output a classification probability signal associated with the cfDNA feature.

[0022] The classification probability signal is used for one of the following non-diagnostic purposes: Screening studies of cancer-related biomarkers; Bioinformatics analysis of tumorigenesis mechanisms; Dynamic monitoring of the efficacy of candidate compounds during drug development.

[0023] The features of this invention are as follows: First, cfDNA is extracted from patient plasma. Then, low-depth whole-genome sequencing is performed on the cfDNA. Combined with epithelial cell-specific open regions mined from single-cell chromatin openness data, a feature matrix for lung cancer identification is constructed. Based on this, a deep learning model, variational autoencoder, and convolutional neural network are used to train and predict the distribution patterns of cfDNA fragments in epithelial cell feature regions, achieving highly sensitive, low-cost, and non-invasive early detection of lung cancer. The detailed process is as follows: S1 identifies epithelial cell characteristics (i.e., open chromatin regions) from single-cell ATAC-seq data, providing a reference for subsequent analysis. This includes: data collection, quality control, and screening for high-quality single-cell data. Data standardization, dimensionality reduction, and clustering were performed to classify cells into six categories (including epithelial cells). Epithelial cell-specific open chromatin regions (characteristic intervals) were extracted, and the coverage of cfDNA in these intervals was calculated.

[0024] Epithelial cell features are chromatin open regions (200 bp upstream and downstream of the peak) unique to epithelial cells. These regions are significantly open in epithelial cells but not in other cell types. The selected 2000-2200 features are the most significantly different chromatin open regions between epithelial cells and other cell types, used for subsequent cfDNA alignment and analysis.

[0025] S2 fuses cfDNA features with epithelial cell features to extract potential high-dimensional features. cfDNA features are coverage signals in sequencing data corresponding to open regions of epithelial cell chromatin. The purpose of fusion is to associate coverage signals of cfDNA with specific open regions of epithelial cells to extract potential biomarkers. Variational autoencoder (VAE) is used to nonlinearly compress cfDNA data and epithelial cell features into a shared latent space.

[0026] S3 trains a convolutional neural network (CNN) using the fused features to distinguish lung cancer patients from healthy controls. A CNN model is constructed, and the input features are the latent space representation after VAE fusion, which contains the association information between cfDNA and epithelial cell features. The model performance is evaluated through training and validation.

[0027] Compared with the prior art, the present application has the following beneficial effects: 1) The cfDNA extraction method provided by the present application has high yield, and the obtained cfDNA has high purity and high stability; 2) The present application is based on single-cell chromatin openness atlas, and the open region features of lung cancer related epithelial cells are accurately constructed to realize the decoding of cfDNA signals at the cell source level. Compared with traditional fragmentomics methods, the cell feature signals can be combined to strengthen the inference of tumor sources, effectively overcome the interference caused by tissue heterogeneity, and improve the resolution of cfDNA analysis.

[0028] 3) The present application introduces a variational autoencoder (VAE) model to deeply fuse cfDNA fragment features and epithelial cell feature signals, automatically learns the distribution rules in the latent space, improves the stability and generalization ability of feature representation, and provides high-quality input for subsequent discriminant modeling.

[0029] 4) The present application uses low-depth whole genome sequencing combined with artificial intelligence deep learning algorithm, which greatly reduces the detection cost compared with traditional high-depth sequencing or targeted panel. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 The extracted cfDNA fragment information map; Figure 2 Distribution and label gene expression of different cell groups: A, 64 publicly available single-cell ATAC-seq sequencing data sets are collected, covering 28 cancer types, 622 patients and 1141 tumor samples; B, after further dimension reduction and clustering analysis, 183,434 high-quality cells are obtained, which are classified into 47 cell subgroups, covering 6 main cell types; C, further fine annotation of various cells; Figure 3 the difference between the chromatin accessibility of different types of cfDNA samples in epithelial cell types; Figure 4 training the pattern map of cfDNA epithelial cell signatures using a variational autoencoder model and a convolutional neural network; Figure 5 training results of the cfDNA epithelial cell signature model; Figure 6 training using a gradient boosting model (GBM) after feature extraction of cfDNA TSS coverage. DETAILED DESCRIPTION

[0031] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some but not all of the embodiments of the present application. In the following description, specific details such as specific configurations and components are provided only to help a comprehensive understanding of the embodiments of the present application. In all the drawings, similar elements or parts are generally identified by similar reference signs. In the drawings, the elements or parts are not necessarily drawn according to the actual proportions; It should be understood that the "one embodiment" or "the embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "one embodiment" or "the embodiment" appearing throughout the specification does not necessarily mean the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner.

[0032] In addition, reference numerals and / or letters can be repeated in different examples in the present application. Such repetition is for the purpose of simplification and clarity, and does not itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0033] The term "and / or" herein is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that there are three cases of A alone, B alone, and A and B together. The term "and" herein is to describe another association relationship of the associated objects, which means that there can be two relationships, for example, A and B can mean that there are two cases of A alone and A and B together. In addition, the character " / " herein generally means that the associated objects before and after are an "or" relationship.

[0034] The term "at least one", as used herein, is merely descriptive Borel correlation, indicating that there may be three relationships, for example, A and B at least one, can represent: the existence of A alone, the existence of A and B, the existence of B alone.

[0035] It should also be noted that in this paper, such as the first and second relationship terms are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the term "includes", "contains" or any other variant thereof is intended to cover non-exclusive inclusion.

[0036] The signac R package used in the following examples is a bioinformatics tool developed specifically for analyzing single-cell chromatin accessibility data (such as scATAC-seq).

[0037] Example 1 A cfDNA extraction method comprising the following steps: Step (1) Use EDTA blood vessels containing anticoagulants and cfDNA protectants (use nuclease inhibitors such as diethyl pyrocarbonate or guanidine isothiocyanate); collect peripheral blood from subjects (such as healthy individuals or cancer patients), usually collect 2 mL; Step (2) centrifuge at 1600 g for 10 minutes, room temperature, transfer the supernatant (plasma) to a new tube, centrifuge at 16,000 g for 10 minutes, room temperature, further remove residual cell debris; collect the final plasma in a DNase-free tube, immediately perform cfDNA extraction, or store at -80°C; centrifuge again, collect the final filtrate, measure the DNA concentration, and detect the distribution of cfDNA fragments, add lysis buffer (containing proteinase K (Qiagen #19131) and ACL buffer (Qiagen #939016) to lyse plasma proteins; Step (3) Add magnetic beads to promote cfDNA binding, wash multiple times to remove impurities; use low-salt elution buffer to elute cfDNA from magnetic beads or silica gel columns, the final elution volume is about 20-50 µL; The magnetic beads use magnetic microspheres modified with carboxyl (-COOH) or silanol (-SiOH) on the surface, with a particle size range of 1-5 μm, and a carboxyl density of ≥300 μmol / g. The carboxyl magnetic beads bind cfDNA through electrostatic interaction in a buffer with a pH of 5.0-6.5, while the silanol magnetic beads need to be bound in a high-salt buffer containing 1.5 M NaCl and 20% PEG 8000. Preferred commercial products include Dynabeads™, MyOne™, Carboxylic Acid, or AMPure XP magnetic beads to improve the adsorption rate of cfDNA.

[0038] The low-salt elution buffer is a TE buffer with a pH of 7.9-8.1 and containing 0.04%-0.06% Tween 20 to reduce the loss of cfDNA adsorption.

[0039] Step (4): Collect the eluate, detect the concentration using a fluorescence quantifier, and detect the fragment size distribution of cfDNA using a bioanalyzer QSEP, with the results as shown in Figure 1 As shown in Figure 1 The median fragment size of the extracted cfDNA is 182 bp, and the median concentration is 10 ng / uL.

[0040] Example 2 S1: Identify cell types using single-cell chromatin opening data and extract features of epithelial cells; S2: Construct a library and perform high-throughput sequencing on the extracted cfDNA sample, and fuse cfDNA and epithelial cell features using a variational auto-encoding model; S3: Train and test cfDNA data using a convolutional neural network.

[0041] In a preferred embodiment, step S1 can include: S101: The present application first collects 64 publicly available single-cell ATAC-seq sequencing data sets, covering 28 cancer types, 622 patients, and 1,141 tumor samples; The 64 public data sets come from the GEO database, and the data numbers are shown in Table 1.

[0042] Table 1 Data numbers of 64 public data sets

[0043] S102, CreateChromatinAssay of signac R package is used to create objects, the input file is fragments.tsv.gz, ensure that each feature exists in at least 10 cells, the minimum number of features per cell is more than 200, uniform quality control is performed on each data set, cells with TSS score greater than 4 and Tn5 fragment number higher than 1000 are selected; then batch effect correction is processed using Harmony, the dimension of Harmony is set to 1:30, and the theta value is 2; S103, the cell quality control standard includes: the promoter region score of each cell is more than 4, and the number of fragments of each cell is more than 1000, and the low quality or apoptotic cells are excluded; S104, finally, 3,904,694 high-quality single cell data are obtained for subsequent analysis.

[0044] S105, RunTFIDF function (default parameters) of signac R package is used to standardize the data, FindTopFeatures function (default parameters) is used to filter high variation region, RunSVD function (default parameters) is used to reduce dimension processing, RunUMAP function uses 'lsi' singular value decomposition to reduce dimension processing, and the dimension is 2:30, which compresses the dimension while retaining key biological information; S106, based on the reduced dimension data, FindNeighbors function and FindClusters function (default parameters) of signac R package are used for clustering analysis, the cells are divided into subgroups with potential biological significance, and all cells are divided into 6 main cell types, including: epithelial cells, T cells, B cells, myeloid cells, fibroblasts and endothelial cells, wherein T cells include CD4 + regulatory T cells, CD4 + naive T cells, CD8 + exhausted T cells, CD8 + cytotoxic T cells, double-positive tissue-resident memory T cells and NKT cells; B cells include: memory B cells, naive B cells and plasma cells; myeloid cells include: macrophages, dendritic cells, myeloid suppressor cells, conventional dendritic cells, plasma cell-like dendritic cells and mast cells; S107, through the above fine annotation, an expression reference atlas covering multiple cell classes in multiple cancer environments is constructed, which provides a reliable epithelial cell source background for subsequent cfDNA feature signal extraction; S108, align cfDNA to the interval of epithelial cell characteristics in single-cell ATAC-seq, the interval of epithelial cell characteristics in single-cell ATAC-seq is 200 bp upstream and downstream of the peak top found by macs2 as the numerator, and the fragment ratio of 1k-3k upstream and downstream of the peak top as the denominator; select 2000-2200 characteristics with the most significant differences between epithelial cells and other cells as the epithelial cell characteristic interval (epithelial cell-specific chromatin open region); S109, calculate the coverage of cfDNA in the epithelial cell characteristic interval.

[0045] In a preferred embodiment, a variational autoencoder model is used to fuse cfDNA features and epithelial cell features, and the detailed operation includes: S201, construct a variational autoencoder structure, including two main modules of an encoder and a latent space, and the input is the coverage of cfDNA in the epithelial cell characteristic interval and the epithelial cell characteristic interval; The number of encoding layers is 5 layers, and the formula of the encoder is: ; : input data (cfDNA data and single-cell ATAC-seq data); : latent variable; The approximate posterior distribution defined by the encoder parameters ; : mean of the encoder output; : variance of the encoder output.

[0046] S202, use the python package PyTorch to perform encoding, and perform nonlinear compression on cfDNA features and epithelial cell features respectively, use functions torch.nn.Linear, torch.nn.ReLU, torch.nn.Sequential (default parameters) for processing, and fuse into a shared latent space representation; S203, introduce a KL divergence regularization term in the latent space, use the python package PyTorch to calculate the KL divergence, and constrain the latent variable to conform to a Gaussian distribution to enhance the generalization ability; Introduce a KL divergence regularization term in the latent space to constrain the latent variable to conform to a Gaussian distribution; This means that the encoder maps the input to a Gaussian distribution, that is: ; If

[0047] ; This represents the mean of the i-th latent variable; This represents the variance of the i-th latent variable; Dimensions representing potential space; Represents logarithmic variance; Represents the standard normal distribution. I Represents the identity matrix.

[0048] S204. During training, the reconstruction error and KL divergence are jointly optimized to obtain stable and biologically meaningful potential dimensionality reduction features.

[0049] The steps for establishing the use of convolutional neural networks to train and test cfDNA data include: S301. Design a convolutional neural network model structure. The model should include at least: an input layer, multiple convolutional layers (Conv), an activation function layer (ReLU), a pooling layer (Pooling), a fully connected layer (FC), and an output layer. The input layer is the dimensionality reduction feature of cfDNA epithelial cells generated by the variational autoencoder model. The convolutional kernel is a 3×3 feature map with a stride of 1 and the pooling method is AvgPool. S302: Convolutional layers extract local spatial features from cfDNA epithelial cell dimensionality reduction feature data, and pooling layers implement downsampling to improve model computational efficiency and robustness. S303. During the training phase, labeled cfDNA samples (lung cancer patients vs. healthy controls) are input into the model for forward and backward propagation. The model weights are updated using the cross-entropy loss function and the Adam optimizer. The learning rate is 0.001, the batch size is 128, and the number of iterations is 30. S304. A convergent model is obtained through multiple rounds of iterative training. At the same time, a Dropout layer and an early stopping mechanism are introduced to prevent overfitting. The Dropout rate is 0.2. S305. Evaluate the model’s accuracy, recall, F1 score, and ROC curve performance using cross-validation and validation sets on the training dataset.

[0050] Accuracy = ; Recall rate = ; Accuracy = ; F1 = 2 * (Recall * Precision) / (Recall + Precision); TP is the number of true positives, FP is the number of false positives, FN is the number of false negatives, and TN is the number of true negatives.

[0051] Verification experiment of lung cancer early screening model Sample information: The number of training set samples is 778 (410 lung cancer patients, 113 benign nodule patients, and 255 healthy controls), and the number of samples in the validation set is 259 (137 lung cancer patients, 37 benign nodule patients, and 85 healthy controls). The clinical stage distribution is as follows: 454 patients in stage I, 51 patients in stage II, 41 patients in stage III, and 1 patient in stage IV.

[0052] The inclusion criteria of benign lung nodule patients are added. The nodule is confirmed as a benign lesion by pathology or clinical follow-up, including but not limited to the following types: pathologically diagnosed as inflammatory pseudotumor, granuloma, tuberculous nodule, histiocytic hyperplastic lesion, etc. No significant growth or malignant transformation is found after 2 years of imaging follow-up, and the diagnosis is stable benign nodule.

[0053] Independent verification: cfDNA samples come from Peking University Cancer Hospital, Peking Union Medical College Hospital, Yunnan First People's Hospital, Yunnan Cancer Hospital, Zhejiang Cancer Hospital, and data from multiple hospitals prove the model's generalization ability.

[0054] Threshold setting: It explains the confidence threshold of the model to judge "lung cancer positive" (such as probability > 0.85).

[0055] Example 3 This example is based on the above-mentioned example 2, and the same parts as example 2 are not repeated.

[0056] This example introduces the training of the cfDNA epithelial cell feature model.

[0057] Step 1: Distribution of epithelial cell clusters and expression of label genes, results as shown in the following table: Figure 2

[0058] This example collects 64 publicly available single-cell ATAC-seq sequencing data sets, covering 28 cancer types, 622 patients and 1141 tumor samples (A). During data integration, first, quality control and batch effect correction are performed on each data set, and standardized Seurat objects are constructed. Cell quality control standards include: selecting cells with TSS scores greater than 4 and Tn5 fragment numbers higher than 1000 (B). Figure 2 Figure 2 ​​A). Finally, 3,904,694 cells passed quality control and were included in subsequent analysis. In the data integration stage, a cell network was constructed based on the quality-controlled cell expression matrix, the k-nearest neighbor algorithm (k = 10) was used to estimate the similarity between cells, and the same cell was aggregated into a metacell through gamma granularity clustering (γ = 20). Then the average expression of each gene in the metacell was calculated to generate an expression matrix containing 195,238 metacells.

[0059] Figure 2 B. After further dimensionality reduction and clustering analysis, 183,434 high-quality metacells were finally obtained, which were classified into 47 cell subgroups, covering 6 main cell types, including: 56,686 epithelial cells, 53,631 T cells, 17,290 B cells, 29,164 myeloid cells, 17,459 fibroblasts and 9,204 endothelial cells (Fig. 2B). Figure 2 B). On this basis, we further annotated each type of cell, including: T cells: 6,890 CD4 + regulatory T cells (Treg), 16,068 CD4 + naive T cells, 4,515 CD8 + exhausted T cells, 16,481 CD8 + cytotoxic T cells, 135 NK cells, 8,841 double-positive tissue-resident memory T cells; B cells: 898 memory B cells, 8,954 naive B cells, 7,438 plasma cells; myeloid cells: 16,835 macrophages, 625 dendritic cells (DCs), 4,984 myeloid-derived suppressor cells (MDSCs), 3,658 conventional dendritic cells (cDCs), 1,605 plasmacytoid dendritic cells (pDCs), 1,457 mast cells, epithelial cells, endothelial cells and fibroblasts (Fig. 2B). Figure 2 C).

[0060] In summary, based on multi-cancer data, the present application constructed a pan-cancer single-cell map containing nearly four million cells, which provided a high-resolution cell map basis for the cfDNA cancer non-invasive early screening method and system specific to cell type-specific chromatin openness (Fig. 2C). Figure 2 A and Figure 3 C).

[0061] Second step: Distribution of cfDNA epithelial cell features in non-small cell lung cancer, benign nodules and healthy samples, results as shown in the following figures. Figure 4

[0062] ​To investigate the differences between healthy peripheral blood, lung cancer patients and benign nodule samples in epithelial cells, the present application observes that cfDNA in the epithelial cell characteristics of benign nodules and lung cancer samples has significantly lower mean coverage, while the coverage of healthy people is higher.

[0063] Step 3: cfDNA epithelial cell characteristic variational auto-encoding model and convolutional neural network model, as shown in the accompanying Figure 5 .

[0064] The cfDNA epithelial cell characteristics are extracted, and after extraction, the variational auto-encoding (VAE) model is used to reduce the dimension of the data. The encoder is a three-layer fully connected network that reduces the input dimension and is used to represent the latent variable distribution. The loss function is the weighted sum of the reconstruction error (mean square error) and the KL divergence, which is used to guide the learning of the latent fusion expression characteristics. The latent variables learned by VAE are represented in a format suitable for convolutional layers and input into a convolutional neural network (CNN). Then, the convolutional layer, the max pooling layer, the Dropout layer (to prevent overfitting), and the fully connected layer output the softmax classification probability (lung cancer / non-lung cancer). Cross-entropy is used as the loss function, and the Adam optimizer is used for backpropagation. The VAE latent variables and the CNN classification intermediate layer features are further spliced as the input of the multi-layer perceptron (MLP), and two layers of fully connected networks are used to complete the final lung cancer prediction. The MLP network is used to integrate low-dimensional features and deep abstract features to improve the overall prediction performance. Finally, a confidence score is output to determine whether the cfDNA sample belongs to a lung cancer patient or a healthy individual.

[0065] Step 4: cfDNA TSS coverage feature extraction and model training, as shown in the accompanying Figure 6 . After using the cfDNA joint epithelial cell characteristics for variational auto-encoding dimension reduction, a convolutional neural network is used for model training. This method can accurately identify the risk of lung cancer, with an AUC of 0.91, an accuracy of 0.95, and a recall rate of 0.87. It can provide accurate diagnosis and treatment guidance for patients and help reduce the risk of death for patients.

[0066] Step 5: After the cfDNA TSS coverage feature extraction, the gradient boosting model (GBM) is trained, as shown in the accompanying ​ . The AUC of the gradient boosting model is 0.81, the accuracy is 0.92, and the recall rate is 0.7. Compared with the deep learning model, its detection performance is much lower.

[0067] Example 4 A computer system for constructing a lung cancer early screening model, comprising: The single-cell data processing module is used to extract epithelial cell-specific open region features from single-cell chromatin openness data; The cfDNA sequencing data processing module is used to acquire and preprocess genomic distribution data of cfDNA fragments; The variational autoencoder module is used to perform nonlinear dimensionality reduction and fusion of the cfDNA fragment distribution data and epithelial cell features to generate a latent spatial feature representation. The convolutional neural network module is used to train a classification model based on the latent spatial feature representation and output a classification probability signal associated with the cfDNA feature.

[0068] The classification probability signal is used for one of the following non-diagnostic purposes: Screening studies of cancer-related biomarkers; Bioinformatics analysis of tumorigenesis mechanisms; Dynamic monitoring of the efficacy of candidate compounds during drug development.

[0069] This embodiment is based on the above embodiment 2, and the similarities with embodiment 2 will not be repeated.

[0070] The epithelial cell-specific open region characteristics were obtained through the following steps: cluster analysis was performed on single-cell ATAC-seq data to screen for chromatin open regions that were significantly different from other cell types in epithelial cells; the coverage ratio of cfDNA fragments in the regions was calculated.

[0071] The variational autoencoder module generates a latent variable representation that conforms to a Gaussian distribution by jointly optimizing the reconstruction error and KL divergence.

[0072] The output layer of the convolutional neural network module generates probabilistic classification results through the Softmax function.

[0073] The above description is merely a preferred embodiment of the present invention and does not limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any changes, modifications, substitutions, integrations, and parameter alterations to these embodiments within the spirit and principles of the present invention, achieved through conventional substitutions or by achieving the same function without departing from the principles and spirit of the present invention, fall within the scope of protection of the present invention.

Claims

1. A method of cfDNA extraction, characterized by, The method comprises the following steps: Step (1): whole blood extraction is performed by using a special blood collection tube preloaded with a cfDNA protective agent; the blood collection tube containing the blood sample is placed in a centrifugal device, and preliminary centrifugation is performed to separate the plasma; Step (2): proteinase K and ACL lysis buffer are added to the obtained plasma, and after being fully mixed, constant temperature incubation is performed to lyse the protein; Step (3): magnetic beads are used to combine cfDNA, impurities are removed through multiple washing, and then the cfDNA is eluted by using a low-salt elution buffer; the magnetic beads are magnetic microspheres with carboxyl or silanol groups on the surface for adsorbing cfDNA, the particle size range is 1-5 μm, and the carboxyl density is greater than or equal to 300 μmol / g; Step (4): the eluate is collected, the concentration is detected by using a fluorescence quantitative instrument, and the fragment size distribution of the cfDNA is detected by using a bioanalyzer.

2. The cfDNA extraction method of claim 1, wherein, The cfDNA protective agent in step (1) is a nuclease inhibitor. 3.The cfDNA extraction method of claim 1, wherein, The low-salt elution buffer in step (3) is a TE buffer with a pH of 7.9-8.1 and containing 0.04%-0.06% Tween 20.

4. The cfDNA obtained by the method according to any one of claims 1-3 for use in constructing a lung cancer early screening model. The method comprises the following steps: S1: using single-cell chromatin opening data to identify cell types and extract epithelial cell characteristics; S2: library construction and high-throughput sequencing are performed on the extracted cfDNA sample, and a variational auto-encoding model is used to fuse cfDNA and epithelial cell characteristics; S3: a convolutional neural network is used to train and test the variational auto-encoding model fused cfDNA epithelial cell dimension reduction characteristics, and the classification probability of the cfDNA characteristic correlation is output.

5. Use according to claim 4, characterized in that, The step of identifying the single-cell chromatin opening data comprises: Publicly available single-cell ATAC-seq sequencing data sets are collected, and the CreateChromatinAssay of the signac R package is used to create objects, the input file is fragments.tsv.gz, it is ensured that each feature exists in at least 10 cells, and the number of minimum features for each cell is more than 200, uniform quality control is performed on each data set, and cells with a TSS score greater than 4 and a Tn5 fragment number higher than 1000 are selected; then batch effect correction processing is performed by using Harmony; The data is standardized by using the RunTFIDF function of the signac R package, the high-variation region is filtered by using the FindTopFeatures function, the data is processed by using the RunSVD function, and the data is processed by using the RunUMAP function in the 'lsi' singular value decomposition mode, so that the key biological information is retained while the dimension is compressed; Based on the dimension-reduced data, the FindNeighbors function and the FindClusters function of the signac R package are used for cluster analysis, the cells are divided into epithelial cells, T cells, B cells, myeloid cells, fibroblasts and endothelial cells, an expression reference atlas of multiple cell clusters is constructed, and a reliable epithelial cell source background is provided for subsequent cfDNA characteristic signal extraction. MACS2 was used to identify open chromatin peaks in epithelial cells, and the peak top position was determined; the ATAC-seq signal of the 200 bp upstream and downstream of the peak top was used as the numerator; the background region signal of the 1k-3k range upstream and downstream of the peak top was used as the denominator; the openness of each peak was quantified by the ratio of the numerator to the denominator; 2000-2200 peaks with the most significant differences from other cell populations were selected as epithelial cell-specific open region features; The coverage of cfDNA in the epithelial cell-specific open region features was calculated.

6. Use according to claim 5, characterized in that, The quality control standards include: the score of each cell promoter region is above 4, and the number of fragments of each cell is above 1000, excluding low-quality or apoptotic cells.

7. Use according to claim 4, characterized in that, The variational autoencoding model is constructed using an encoder and a latent space module; The input is the coverage of cfDNA in the epithelial cell interval and the epithelial cell open feature; The formula of the encoder is: ; : input data; : latent variable; : defined by the encoder parameters approximate posterior distribution; : mean of the encoder output; : variance of the encoder output; In the encoding stage, the coverage of cfDNA in the epithelial cell interval and the epithelial cell open feature are nonlinearly compressed and fused into a shared latent space representation; In the latent space, a KL divergence regularization term is introduced to constrain the latent variables to conform to a Gaussian distribution; This means that the encoder maps the input to a Gaussian distribution, i.e. ; If , KL divergence has an analytical solution; ; a mean value representing the ith latent variable; This represents the variance of the i-th latent variable; Dimensions representing potential space; Represents logarithmic variance; representing a standard normal distribution, I representing an identity matrix; The dimensionality-reduced data is input into a convolutional neural network for model training and supervision; During training, the reconstruction error and KL divergence are jointly optimized to obtain stable and biologically meaningful latent dimensionality-reduced features.

8. Use according to claim 4, characterized in that, The convolutional neural network includes a model structure of an input layer, multiple convolutional layers, an activation function layer, a pooling layer, a fully connected layer, and an output layer; the input layer is the cfDNA epithelial cell dimensionality-reduced feature output by the variational autoencoding model, the convolutional layer extracts local spatial features from the cfDNA epithelial cell dimensionality-reduced feature data, and the pooling layer realizes down-sampling to improve model calculation efficiency and robustness; In the training stage, labeled cfDNA samples are used as input to the model for forward propagation and backpropagation, and the model weight is updated using the cross-entropy loss function and the Adam optimizer; the label is a lung cancer patient or a healthy control; Through multiple rounds of iterative training, a converged model is obtained, and Dropout layers and early stopping mechanisms are introduced to prevent overfitting; In the training data set, cross-validation and validation set are used to evaluate the accuracy, recall, precision, F1 score, and ROC curve performance of the model; Accuracy = ; Recall rate = 0.5 ; Precision = ; F1=2*(recall*precision) / (recall+precision); TP is the number of true positives, FP is the number of false positives, FN is the number of false negatives, and TN is the number of true negatives.

Citation Information

Patent Citations

  • Non-invasive cancer early screening system based on cfDNA omics characteristics

    CN113160889A

  • Cancer prediction model based on cfDNA and construction method and application thereof

    CN116312774A

  • Method and system for early diagnosis and typing of esophageal squamous carcinoma based on plasma free DNA

    CN117524321A

  • Detection of lung cancer using cell-free DNA fragmentation

    WO2022140386A1

  • Machine-learning approaches to pan-cancer screening in whole genome sequencing

    WO2024238593A1