Methods, systems, and equipment for diagnosing malignant tumors or predicting the treatment efficacy of malignant tumor patients based on mature erythrocyte DNA characteristics.
By obtaining the concentration, fragment length, and coefficient of variation of erythrocyte DNA, a model for the diagnosis and treatment prediction of malignant tumors was constructed using machine learning algorithms. This solved the problem of insufficient application of erythrocyte DNA in the diagnosis and treatment prediction of malignant tumors, and achieved efficient diagnostic and prediction results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2026-03-10
AI Technical Summary
In the current technology, the application of mature red blood cell DNA in the diagnosis and efficacy prediction of malignant tumors has not been fully utilized. The number of circulating tumor cells is scarce, existing methods are inefficient, and the differences in red blood cell DNA characteristics between healthy individuals and patients with malignant tumors are still unclear.
By obtaining the concentration, fragment length, and coefficient of variation of red blood cell DNA, machine learning algorithms are used to construct diagnostic models and efficacy prediction models for malignant tumors. Diagnosis and prediction are performed based on these feature data. Development tools such as TensorFlow and Scikit-Learn are used to construct algorithm models such as linear regression and neural networks.
This approach enables the diagnosis and efficacy prediction of malignant tumors based on erythrocyte DNA characteristics, improving the accuracy of diagnosis and the reliability of efficacy assessment, and demonstrating the clinical application value of erythrocyte DNA as a potential biomarker.
Smart Images

Figure CN119650042B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics, specifically relating to methods, systems, and devices for diagnosing malignant tumors or predicting the treatment efficacy of malignant tumor patients based on the DNA characteristics of mature red blood cells. Background Technology
[0002] Because mature red blood cells lack a nucleus, red blood cell DNA has long been overlooked. In recent years, however, DNA has been discovered in mature human red blood cells. Studies have reported that red blood cells can act as "immune sentinels" in infectious diseases, activating macrophage-mediated innate immunity by uptake of pathogen-derived CpG DNA via their surface TLR9 protein. Other research indicates that red blood cells contain long DNA fragments that can drive mutations through direct contact uptake from cancer cell lines. However, circulating tumor cells (CTCs) are rare, occurring only once every 1010 times. 7 There are only about one cell-free total tumor cells (CTCs) per white blood cell. Red blood cells taking up DNA solely through direct contact is an inefficient method. Given that tumors can release DNA into the plasma, it remains unclear whether red blood cells (RBCs) can absorb cell-free DNA from tumors and its clinical significance. Therefore, it is crucial to investigate the characteristics of mature red blood cell DNA and its clinical significance in patients with malignant tumors. Summary of the Invention
[0003] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a method, system, and device for diagnosing malignant tumors or predicting the treatment efficacy of malignant tumor patients based on the characteristics of mature red blood cell DNA.
[0004] To achieve the above objectives, the present invention adopts the following technical solution:
[0005] The first aspect of this invention provides a method for diagnosing malignant tumors based on the DNA characteristics of mature erythrocytes.
[0006] Furthermore, the method is performed by a computer, and the method includes the following steps:
[0007] Data Acquisition: Acquire data on the DNA characteristics of red blood cells in the sample of the subject to be tested. The DNA characteristics include one or more of the following: DNA concentration, DNA fragment length, and DNA length variation coefficient.
[0008] Data processing: The data of the red blood cell DNA characteristics are input into the constructed malignant tumor diagnostic model, which predicts whether the subject to be tested has a malignant tumor based on the data of the red blood cell DNA characteristics;
[0009] Output results.
[0010] Furthermore, the steps for constructing the malignant tumor diagnostic model are as follows:
[0011] obtaining data of red blood cell DNA characteristics, the DNA characteristics including one or more of DNA concentration, DNA fragment length, DNA length coefficient of variation; the data of red blood cell DNA characteristics from normal healthy subjects and patients with malignant tumors; inputting the data of red blood cell DNA characteristics into a machine learning algorithm to construct a malignant tumor diagnosis model.
[0012] Preferably, the construction step of the malignant tumor diagnosis model further comprises verifying the model efficiency, and the method of verifying the model efficiency comprises ROC curve.
[0013] Further, the malignant tumor diagnosis model obtains results by the following standards: when the DNA concentration or the DNA length coefficient of variation is higher than the threshold value or more than 50% of the DNA fragment length does not fall within the reference value range, a classification result that the subject to be tested has a malignant tumor is obtained; if the DNA concentration or the DNA length coefficient of variation is lower than the threshold value or more than 50% of the DNA fragment length is within the reference value range, a classification result that the subject to be tested does not have a malignant tumor is obtained.
[0014] In some embodiments of the present application, the preset threshold value is a representative value of normal samples of a healthy subject population, including but not limited to maximum value, third quartile, average value. In some preferred embodiments of the present application, the population sample includes more than 20 samples, for example, 30, 50, 80, 100, 150, 200, 300, 500 or more. In a specific embodiment of the present application, the preset threshold value of the DNA concentration is 113 ng / ml, the preset threshold value of the DNA length coefficient of variation is 1.032, and the reference value range of the DNA fragment length is 1 kb to 8 kb.
[0015] Further, the machine learning algorithm includes algorithm models developed by various development tools.
[0016] Further, the development tools include but are not limited to TensorFlow, Scikit Learn, PyTorch, OpenNN, RapidMiner, Azure Machine Learning, Apache Mahout, Shogun, KNIME, VertexAI, H2Oai, Anaconda, Keras, Tableau, Fast.ai, Catalyst, Amazon ML, MLJAR, and Spell.
[0017] Furthermore, the algorithm models include, but are not limited to, linear regression models, logistic regression models, Lasso regression models, Ridge regression models, linear discriminant analysis models, nearest neighbor models, decision tree models, perceptron models, neural network models, support vector machine models, Naive Bayes models, AdaBoost models, GBDT models, XGBoost models, LightGBM models, CatBoost models, and random forest models.
[0018] Furthermore, the subjects to be tested include humans and / or mammals.
[0019] Furthermore, the malignant tumors include gastric cancer, pancreatic cancer, esophageal cancer, rectal cancer, colon cancer, lung cancer, nasopharyngeal cancer, laryngeal cancer, malignant lymphoma, melanoma, endometrial cancer, cervical cancer, ovarian cancer, bladder cancer, and kidney cancer.
[0020] Preferably, the malignant tumor is lung cancer.
[0021] A second aspect of the present invention provides a system for diagnosing malignant tumors based on the DNA characteristics of mature erythrocytes.
[0022] Furthermore, the system includes:
[0023] Data acquisition unit: used to acquire data on the DNA characteristics of red blood cells in the sample of the subject to be tested, wherein the DNA characteristics include one or more of the following: DNA concentration, DNA fragment length, and DNA length variation coefficient;
[0024] Data classification unit: used to classify and predict whether the data obtained by the data acquisition unit has a malignant tumor diagnosis model obtained by the construction method described in the first aspect of the present invention, and to obtain the classification result of whether the subject to be tested has a malignant tumor.
[0025] Result output unit: Used to output classification results.
[0026] Furthermore, the malignant tumors include gastric cancer, pancreatic cancer, esophageal cancer, rectal cancer, colon cancer, lung cancer, nasopharyngeal cancer, laryngeal cancer, malignant lymphoma, melanoma, endometrial cancer, cervical cancer, ovarian cancer, bladder cancer, and kidney cancer.
[0027] Preferably, the malignant tumor is lung cancer.
[0028] A third aspect of the present invention provides a method for predicting the efficacy of treatment in patients with malignant tumors based on red blood cell DNA concentration.
[0029] Furthermore, the method is performed by a computer, and the method includes the following steps:
[0030] Data Acquisition: Obtain data on the concentration of erythrocyte DNA in samples from patients with malignant tumors to be tested;
[0031] Data processing: The red blood cell DNA concentration data is input into the pre-constructed efficacy prediction model, which predicts the efficacy of treatment for patients with malignant tumors based on the red blood cell DNA concentration data;
[0032] Output the prediction results.
[0033] Furthermore, the steps for constructing the efficacy prediction model are as follows: obtaining red blood cell DNA concentration data, wherein the red blood cell DNA concentration data comes from patients with high disease control rates after treatment and patients with poor disease control rates after treatment; inputting the red blood cell DNA concentration data into a machine learning algorithm to construct the efficacy prediction model.
[0034] Preferably, the construction step of the efficacy prediction model further includes verifying the model efficacy, and the method for verifying the model efficacy includes ROC curves.
[0035] Furthermore, the efficacy prediction model obtains prediction results using the following criteria: when the red blood cell DNA concentration is higher than a threshold, a classification result of poor disease control rate after treatment is obtained for malignant tumor patients; when the red blood cell DNA concentration is lower than a threshold, a classification result of high disease control rate after treatment is obtained for malignant tumor patients.
[0036] Furthermore, the malignant tumors include gastric cancer, pancreatic cancer, esophageal cancer, rectal cancer, colon cancer, lung cancer, nasopharyngeal cancer, laryngeal cancer, malignant lymphoma, melanoma, endometrial cancer, cervical cancer, ovarian cancer, bladder cancer, and kidney cancer.
[0037] Preferably, the malignant tumor is lung cancer.
[0038] The fourth aspect of this invention provides a system for predicting the treatment efficacy of patients with malignant tumors based on red blood cell DNA concentration.
[0039] Furthermore, the system includes:
[0040] Data acquisition module: used to acquire data on the concentration of red blood cell DNA in samples from patients with malignant tumors to be tested;
[0041] Data processing module: used to classify and predict the efficacy prediction model obtained by the construction method described in the third aspect of the present invention on the data obtained by the data acquisition module, and obtain the classification results of the efficacy of the malignant tumor patients to be tested;
[0042] Output results module: Used to output classification results.
[0043] Furthermore, the malignant tumors include gastric cancer, pancreatic cancer, esophageal cancer, rectal cancer, colon cancer, lung cancer, nasopharyngeal cancer, laryngeal cancer, malignant lymphoma, melanoma, endometrial cancer, cervical cancer, ovarian cancer, bladder cancer, and kidney cancer.
[0044] Preferably, the malignant tumor is lung cancer.
[0045] The fifth aspect of the present invention provides a computer device.
[0046] Furthermore, the computer device includes: a memory and a processor, the memory being used to store program instructions; the processor being used to invoke the program instructions, and when the program instructions are executed, to implement the method for diagnosing malignant tumors as described in the first aspect of the present invention or the method for predicting the efficacy of treatment for patients with malignant tumors as described in the third aspect of the present invention.
[0047] The sixth aspect of the present invention provides a computer-readable storage medium.
[0048] Furthermore, the computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for diagnosing malignant tumors as described in the first aspect of the present invention or the method for predicting the treatment efficacy of patients with malignant tumors as described in the third aspect of the present invention.
[0049] Furthermore, the malignant tumors include gastric cancer, pancreatic cancer, esophageal cancer, rectal cancer, colon cancer, lung cancer, nasopharyngeal cancer, laryngeal cancer, malignant lymphoma, melanoma, endometrial cancer, cervical cancer, ovarian cancer, bladder cancer, and kidney cancer.
[0050] Preferably, the malignant tumor is lung cancer.
[0051] In this invention, the lung cancer includes small cell lung cancer and non-small cell lung cancer. Small cell lung cancer is a cancerous lesion occurring in the bronchial mucosa or glands, with a high degree of malignancy, including lymphocytic, intermediate cell, and mixed types. Non-small cell lung cancer includes squamous cell carcinoma, adenocarcinoma, and large cell carcinoma.
[0052] Advantages and beneficial effects of the present invention:
[0053] This invention investigated the clinical significance of red blood cell DNA (RBC DNA) in patients with malignant tumors, and found that RBC DNA characteristics differ between healthy individuals and patients with malignant tumors. Further research revealed that RBC DNA concentration in patients with malignant tumors is related to the treatment efficacy. Based on this, this invention provides a method for diagnosing malignant tumors based on the characteristics of mature red blood cell DNA, and also provides a method, system, and device for predicting the treatment efficacy of malignant tumor patients based on red blood cell DNA concentration. Attached Figure Description
[0054] Figure 1 This is a schematic flowchart of the method for diagnosing malignant tumors provided by the present invention;
[0055] Figure 2This is a schematic diagram of the system structure for diagnosing malignant tumors provided by the present invention;
[0056] Figure 3 A schematic flowchart of the method for predicting the efficacy of treatment in patients with malignant tumors provided by the present invention;
[0057] Figure 4 A schematic diagram of the system structure for predicting the treatment efficacy of patients with malignant tumors provided by the present invention;
[0058] Figure 5 A schematic diagram of the structure of the computer device provided by the present invention;
[0059] Figure 6 This image shows the differences in RBC DNA fragments between patients with malignant tumors and healthy individuals. Figure A is a confocal image showing Hoechst 33342 (blue) and CD235a (yellow) in erythrocytes, scale bar 7.5 μm; Figure B is gel electrophoresis of erythrocyte DNA from patients with malignant tumors; Figure C shows the overall fragment distribution (0–20 kb) of erythrocyte and leukocyte DNA; Figure D is a box plot showing the proportions of short (<1 kb), medium (1 kb–8 kb), and long (>8 kb) fragments; Figure E shows the distribution of ladder-labeled RBC DNA in short fragments (<1 kb) analyzed using ONT, where CR represents erythrocyte DNA from patients with malignant tumors, with vertical lines at 170, 370, 570, 770, and 970 bp; Figure F shows the distribution of erythrocyte and leukocyte DNA in short fragments (<1 kb); Figure G shows RBCs and WBCs. The distribution of DNA in the promoter region (within 3kb of the TSS); Figure H shows the distribution of short DNA fragments (<1kb) from red blood cells and white blood cells in the promoter region; Figure I shows the frequency of six terminal motifs in different ONT samples; Figure J shows the relative frequency of 4-mer motifs (left) and A / T / C / G bases (top right) in the terminal motif distribution diagram I, with the sequence marker map of distribution diagram I at the bottom right; Figure K shows the coefficient of variation for the length of fragments from malignant tumors and normal sources;
[0060] Figure 7This diagram illustrates the fragment characteristics and methylation patterns of tumor-derived DNA in erythrocyte DNA. Figure A shows the establishment and sampling of the NCG mouse CDX model; Figure B shows the tumor size (left) and volume (right) after one month of growth; Figure C shows the coverage of human DNA in mouse RBC DNA across the entire genome, with NG1 on the left, NG4 in the middle, and NG2 on the right; Figure D shows the proportion of human-derived erythrocyte DNA; Figure E shows the distribution of human DNA (left), fragments with CCCA terminal motifs (middle), and fragments with CCNN terminal motifs (right); Figure F shows the proportion of CCNN motifs in different fragment lengths (Anova: one-way ANOVA); Figure G shows the proportion of CCCA motifs in different fragment lengths; Figure H shows RBCs from samples LLC1 (left), LLC2 (middle), and LLC3 (right). The length distribution of human and murine DNA in the DNA sample; Figure I shows the CNV of short (top) and long (bottom) fragments in the CR4 sample, and the tumor purity inferred by ichorCNA; Figure J shows the methylation distribution (left) of the ONT sample and the MDS analysis (right) of methylation patterns in long (>8kb) fragments, with whole-genome methylation of ONT located at probe sites of the H450K methylation array; Figure K shows the 5mC and 5hmC methylation of long fragments in the promoter region of the CR1 sample; Figure L shows the fragment characteristics of somatic mutation-carrying fragments (left) and germline mutation-carrying fragments (right) in the RBC DNA of ID29;
[0061] Figure 8 The graph shows the relationship between erythrocyte DNA concentration, fragmentation characteristics, and clinical indicators. Figure A is a scatter plot showing the relationship between the CV (coherent corpuscular velocity) of fragment lengths of 150–340 bp (upward) and 340–400 bp (downward) and the reciprocal of tumor purity inferred from cfDNA. Figure B is a waterfall plot showing the CV for fragment lengths of 340–400 bp and the evaluation of treatment efficacy, PD, progressive disease, PR, partial remission, SD, and stable disease. Figure C shows the differences in erythrocyte DNA levels among patients with different histological types. Figure D shows the differences in erythrocyte DNA levels among patients with malignant tumors at different clinical TNM stages. Differences in erythrocyte DNA levels; Figure E shows a box plot of erythrocyte DNA concentration for different tumor sizes; Figure F shows a scatter plot of the relationship between erythrocyte DNA levels and HGB (hemoglobin); Figure G shows a scatter plot of the relationship between erythrocyte DNA levels and RDW (reactive protein level); Figure H is a scatter plot showing the relationship between erythrocyte DNA levels and the natural logarithm of CA125 (left) and NSE (right); Figure I shows the relationship between erythrocyte DNA levels and efficacy assessment; Figure J shows the changes in erythrocyte DNA concentration and tumor marker levels in serially sampled specimens. Detailed Implementation
[0062] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0063] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0064] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0065] Figure 1 This is a schematic flowchart of a method for diagnosing malignant tumors provided by the present invention. Specifically, the method includes:
[0066] 101: Acquire data, acquire data on the DNA characteristics of red blood cells in the sample of the subject to be tested, wherein the DNA characteristics include one or more of the following: DNA concentration, DNA fragment length, and DNA length variation coefficient;
[0067] In some embodiments of this invention, the applicant, through extensive and in-depth research, discovered significant differences in RBC DNA characteristics between healthy individuals and patients with malignant tumors, including DNA concentration, DNA length variation coefficient, and DNA fragment length. This suggests that DNA concentration, DNA length variation coefficient, and DNA fragment length in erythrocytes could serve as potential biomarkers for diagnosing malignant tumors.
[0068] In some implementations, the subject may be human or non-human and may include, for example, animal strains or species used as a "model system" for research purposes. Similarly, subjects may include adults or adolescents (e.g., children). Furthermore, subjects may refer to mammals (e.g., humans or non-humans). Examples of mammals include, but are not limited to, any member of the mammalian class: humans, non-human primates (e.g., chimpanzees) and other apes and monkeys; livestock such as cattle, horses, sheep, goats, and pigs; domestic animals such as rabbits, dogs, and cats; and laboratory animals including rodents such as rats, mice, and guinea pigs. Examples of non-mammals include, but are not limited to, birds, fish, etc.
[0069] In the context of this invention, the term "sample" as used refers to a composition obtained from or derived from a patient / subject that contains cells and / or other molecular entities to be characterized and / or identified based on, for example, physical, biochemical, chemical, and / or physiological characteristics. For example, a sample refers to any sample derived from a patient / subject that is expected or known to contain cells and / or molecular entities to be characterized. Samples include, but are not limited to, tissue samples, primary or cultured cells or cell lines, cell cultures, cell supernatants, cell lysates, platelets, serum, plasma, vitreous fluid, lymph, synovial fluid, follicular fluid, semen, pancreatic juice, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, mucus, tissue culture fluid, tissue extracts, homogenized tissue, cell extracts, and combinations thereof.
[0070] In a specific embodiment of the invention, the samples include red blood cells from healthy individuals and red blood cells from patients with malignant tumors. Blood samples and clinical data were obtained from the internal malignant tumor cohort of Zhongnan Hospital, Wuhan University. Genetic testing information of the primary lesions of the patients was obtained from previous clinical records. Blood samples were collected from two healthy blood donors from the applicant and their colleagues. This study was approved by the Medical Ethics Committee of Zhongnan Hospital, Wuhan University (No.: 2022032K). Informed consent was provided by all individuals.
[0071] In this invention, the DNA length variation coefficient refers to an indicator of the stability and diversity of genomic DNA during DNA replication. The DNA variation coefficient is typically detected through genomic DNA stability analysis. Specific detection methods include PCR amplification and sequencing. The DNA fragment length can be measured using base pairing techniques or molecular biology techniques (high-throughput sequencing, PCR, etc.). DNA concentration is determined by measuring the absorbance of DNA at 260 nm using a UV spectrophotometer and calculating the DNA concentration based on the absorbance value. Therefore, DNA characteristic data can be detected using methods well-known in the art.
[0072] In one embodiment of the invention, we found differences in the length, concentration, and coefficient of variation of erythrocyte DNA fragments between lung cancer patients and the normal population. To study the erythrocyte DNA of lung cancer patients, we used density gradient centrifugation to separate erythrocytes from peripheral blood samples in our internal cohort. Flow cytometry confirmed that >99% of the separated cells were CD45 negative (a leukocyte marker) and CD235a positive (an erythrocyte marker), validating the purity of our extracted erythrocytes. Using confocal microscopy, we were able to visualize the DNA within mature human erythrocytes (…). Figure 6 A) The RBC DNA of lung cancer patients is a long fragment >10kb ( Figure 6 B). Due to the low sensitivity of gel electrophoresis, the exact size of RBC DNA remains unknown. Oxford nanopore technology (ONT) is a novel method for characterizing whole-genome fragment lengths. We extracted RBC DNA from 2 healthy donors and 6 lung cancer patients for ONT sequencing. Compared to lung cancer patients, the RBC DNA distribution in healthy individuals was skewed to the right, intersecting with the RBC DNA distribution in lung cancer patients at approximately 8 kb. Figure 6 C). Compared with healthy donors (<1kb, median 20.5%; >8kb, median 16.4%), lung cancer patients had more RBC DNA in the <1kb (median 29.7%) and >8kb (median 21.3%) ranges. Figure 6 In fragment D), we found that short fragments of RBC DNA from cancer patients were mainly around 170 bp. Interestingly, in the RBC DNA of three malignant tumor patients (CR1, CR4, and CR6), the DNA fragments exhibited a periodic pattern, with peaks between 190 and 200 bp, extending to 1 kb to 2 kb. Figure 6 E), which is similar to the apoptotic ladder produced by apoptosis. Cycle signatures were not found in the DNA of healthy leukocytes (paired with CR1) and erythrocytes. Figure 6 F).
[0073] DNA was annotated to genomic regions, and most DNA samples, except for CR1 and CR6, showed no enrichment or deletion near transcription start sites (TSS). Figure 6 G). When focusing on short fragments <1kb, we observed reduced coverage of most DNA in the promoter region; however, coverage increased in CR1, CR4, and CR6 samples. Figure 6H). These samples are associated with the apoptotic ladder, indicating that ladder features in erythrocyte DNA are more frequently detached from transcriptionally active regions. Furthermore, DNA fragmentation is also influenced by DNA nucleases. Different nucleases have different preferences for cleavage and the production of 5' end base sequences. Previously, 5' 4-mer ends were used to describe the fragmentation patterns of cfDNA. Here, we calculated the frequencies of all 256 4-mer A / T / C / G themes in RBC DNA. Using the nonnegative matrix factorization (NMF) algorithm, we identified six clusters of end pattern profiles ( Figure 6 I). Compared to healthy donors, contour I is more frequently found in the RBC DNA of cancer patients, characterized by the presence of an A base at position 1. Figure 6 J). The pattern of contour I closely resembles the cleavage contour of the DNA breakage factor subunit beta (DFFB). We also found that contour V is prominent in CR1, CR4, and CR6, exhibiting characteristics of the apoptotic ladder. Contour V is predominantly T-terminal. Notably, the coefficient of variation of RBC DNA fragment length in malignant tumor patients is significantly higher than that in normal individuals. Figure 6 This indicates that RBC DNA is more disordered in cancer patients. In summary, these results suggest that the length and coefficient of variation of erythrocyte DNA fragments differ between patients with malignant tumors and the general population.
[0074] Furthermore, we demonstrated through animal experiments that tumor-derived DNA in erythrocyte DNA is concentrated in short and long fragments rather than medium-length fragments. To characterize tumor DNA in RBC DNA, we constructed a cell line-derived xenograft (CDX) model, implanting human PC9 cells into the axilla of NCG mice. One month later, we selected three mice with different tumor sizes (NG1, NG2, and NG4) and extracted their RBC DNA for WGS analysis (depth 22–26). Figure 7 (A and 7B). NG1 tumor tissue was the largest and showed rib infiltration. Human-matched RBC DNA was considered to be of tumor origin. NG1 DNA had the highest proportion of human origin, while NG2 had the lowest, consistent with tumor size. Figure 7 C and 7D).
[0075] Next, we investigated the fragmentation characteristics of erythrocyte DNA. The peak value of the NG1 fragment was between 50 and 150 bp, while the NG2 and NG4 fragments were longer, with peak values around 300 bp. Figure 7E). Previous studies have shown that circulating tumor DNA fragments are concentrated in the <150bp range. Using ichorCNA, we found that the 50–150bp fragment had higher tumor purity compared to all fragments. Furthermore, we identified human DNA mutations in RBC DNA, with the 50–150bp fragment being the most abundant. Terminal motifs CCCA and CCNN were downregulated in cfDNA from malignant tumor patients due to insufficient expression of the enzyme DNASE1L3, which produces CCNN cleavage preference in tumors. Pan-cancer data from the TCGA database showed that DNASE1L3 was generally downregulated in tumor tissues compared to normal tissues. We investigated the length distribution of CCCA motif fragments and found that they were mainly distributed around 300bp (…). Figure 7 E). CCCA and CCNN primitives have the lowest abundance in the 50–150 bp range and the highest abundance in the 200–340 bp range. Figure 7 F and 7G) indicate that the ~300bp NG2 and NG4 enriched fragments are not of tumor origin. Furthermore, after co-culturing with the LLC cell supernatant mentioned in Section II, the distribution of mouse-derived (i.e., LLC-derived) DNA in RBC DNA was shorter than that of human-derived DNA (F and 7G). Figure 7 H).
[0076] In the ONT sequencing dataset, we performed ichorCNA to infer tumor purity for short, medium, and long fragments. Short fragments showed the highest tumor purity, medium fragments the lowest, and long fragments also showed relatively high purity. Figure 7 I). Next, we present long-fragment methylation patterns. The degree of long-fragment methylation in erythrocyte DNA from lung cancer patients is lower than that in normal subjects ( Figure 7 J). Multidimensional scale (MDS) analysis distinguished between erythrocyte DNA from cancer patients, erythrocyte DNA from healthy individuals, and leukocyte DNA, while MDS analysis of medium-length fragments could not distinguish them. Figure 7 J). This indicates that long fragments of erythrocyte DNA are tumor-specific. Furthermore, we found that the 5mC of long erythrocyte DNA fragments is hypomethylated in the promoter region, while the 5hmC is hypermethylated (J). Figure 7 This feature was also found in short and medium fragments. Short fragments showed the most disordered methylation in the promoter region. Erythrocyte DNA fragments with somatic variants were defined as tumor-derived in high-depth targeted sequencing data. SVM was used to determine whether DNA fragments contained somatic variants. We found that somatic variants were concentrated in the 90–150 bp and 340–400 bp ranges. Figure 8 Surprisingly, this interval remained consistent across different patients. In contrast, the distribution of germline variation shifted more to the right. In summary, we found a preference for tumor-derived fragments in erythrocyte DNA that are either smaller than 1 kb or larger than 8 kb.
[0077] In another embodiment of the invention, we found that erythrocyte DNA levels are associated with tumor burden in lung cancer patients. To facilitate lung cancer management, we performed a clinical analysis of RBC DNA from 87 patients, of whom 16 underwent 1021 targeted sequencing sessions and 71 underwent RBC DNA concentration testing. The targeted sequencing cohort allowed us to investigate the clinical significance of fragment characteristics. The coefficient of variation for the length of 340–400 bp RBC DNA was positively correlated with tumor purity inferred from cfDNA, while 150–340 bp was negatively correlated with tumor purity. Figure 8 A). In addition, we observed that patients with low coefficient of variation in fragment length of 340–400 bp were more likely to achieve disease control rate (DCR) in their most recent efficacy assessment before blood collection (100% vs 54.5%, Fisher's test P = 0.12, cutoff value was the mean). Figure 8 B). We collected blood samples from 71 patients in our internal cohort, including 58 with non-small cell lung cancer, 9 with small cell lung cancer, 1 with pulmonary lymphoma, 1 with pneumonia, 1 with pulmonary tuberculosis, and 1 case unknown. Two patients with infectious diseases had low RBC DNA levels. Figure 8 C). Except for T1 stage, the RBC DNA concentration gradually increases in patients with T2 to T4 stage malignant tumors, with N3 stage patients showing significantly higher concentrations than N1 stage patients, and M1 stage patients showing significantly higher concentrations than M0 stage patients. Figure 8 D). From a clinical staging perspective, patients in stages III and IV have higher RBC DNA levels. Patients with tumors larger than 30 mm in maximum diameter have even higher RBC DNA concentrations. Figure 2 E), which indicates a positive correlation between RBC DNA and tumor burden, suggesting that RBC DNA concentration is a potential biomarker for the diagnosis of malignant tumors.
[0078] 102: Data Processing: The data of the red blood cell DNA characteristics are input into the constructed malignant tumor diagnostic model, which predicts whether the subject to be tested has a malignant tumor based on the data of the red blood cell DNA characteristics;
[0079] In some embodiments of the present invention, the methods for constructing malignant tumor diagnostic models are known to those skilled in the art and can be implemented and realized in different ways, including the steps of associating data of red blood cell DNA characteristics with a certain probability or risk.
[0080] In the context of this invention, the term "machine learning" refers to the use of computers to simulate or implement human learning activities. Technicians typically use various development tools to build machine learning algorithmic models. These development tools include, but are not limited to, TensorFlow, Scikit-Learn, PyTorch, OpenNN, RapidMiner, Azure Machine Learning, Apache Mahout, Shogun, KNIME, VertexAI, H2Oai, Anaconda, Keras, Tableau, Fast.ai, Catalyst, Amazon ML, MLJAR, and Spell. The algorithmic models include, but are not limited to, linear regression models, logistic regression models, Lasso regression models, Ridge regression models, linear discriminant analysis models, nearest neighbor models, decision tree models, perceptron models, neural network models, support vector machine models, Naive Bayes models, AdaBoost models, GBDT models, XGBoost models, LightGBM models, CatBoost models, or random forest models.
[0081] In one embodiment, after constructing a malignant tumor diagnostic model, the effectiveness of the model can be evaluated using ROC curve analysis.
[0082] An ROC curve is a graph of the true positive rate (sensitivity) versus the false positive rate (100% - specificity) of an experiment. It is useful for depicting the performance of a specific characteristic when distinguishing between two populations. Typically, characteristic data are selected across the entire population in ascending order based on the values of a single characteristic. Then, for each value of that characteristic, the true positive and false positive rates of the data are calculated. The true positive rate is determined by counting the number of cases with values higher than that characteristic and dividing by the total number of cases. The false positive rate is determined by counting the number of controls with values higher than that characteristic and dividing by the total number of controls. While this definition refers to cases where the characteristic is higher in cases compared to controls, it also applies to cases where the characteristic is lower in cases compared to controls (in which case samples with values lower than that characteristic are counted). ROC curves can be generated with respect to individual characteristics and can also be generated with respect to other individual outputs. For example, combinations of two or more characteristics can be mathematically combined (e.g., addition, subtraction, multiplication, etc.) to provide individual sum values that can be plotted on the ROC curve. Furthermore, any combination of multiple features derived from individual output values can be plotted on a ROC curve.
[0083] 103: Output results.
[0084] Figure 3 This is a schematic diagram of the system structure for diagnosing malignant tumors provided by the present invention.
[0085] The system is programmed or otherwise configured to include a data acquisition unit 201, a data classification unit 202, and a result output unit 203.
[0086] Data acquisition unit: used to acquire data on the DNA characteristics of red blood cells in the sample of the subject to be tested, wherein the DNA characteristics include one or more of the following: DNA concentration, DNA fragment length, and DNA length variation coefficient;
[0087] Data classification unit: used to classify and predict whether the data obtained by the data acquisition unit has a malignant tumor diagnosis model obtained by the construction method described in the first aspect of the present invention, and to obtain the classification result of whether the subject to be tested has a malignant tumor.
[0088] Result output unit: Used to output classification results.
[0089] The system may be a user's electronic device or a computer system remotely located relative to that electronic device.
[0090] Figure 8 This is a schematic flowchart of the method for predicting the treatment efficacy in patients with malignant tumors provided by the present invention. Specifically, the method includes:
[0091] 301: Data Acquisition, acquiring data on the concentration of red blood cell DNA in samples from patients with malignant tumors to be tested;
[0092] In one embodiment of the invention, we observed a correlation between erythrocyte DNA concentration and treatment efficacy in lung cancer patients. In the most recent pre-blood sampling efficacy assessment, we found that patients with lower RBC DNA levels had a higher disease control rate (DCR) (92.1% vs 76.9%, Fisher's test P = 0.16, cutoff value was the top 25%). Figure 8 I). We analyzed all 8 samples (from 4 patients), which were collected sequentially and had complete tumor marker information. Patients with altered RBC DNA levels also had altered tumor marker levels. Patients with smaller changes in RBC DNA levels also had smaller changes in tumor marker levels. Figure 4 In summary, our findings provide a comprehensive reference for understanding RBC DNA and exploring its potential as a clinical marker.
[0093] 302: Process the data by inputting the red blood cell DNA concentration data into the pre-constructed efficacy prediction model, which diagnoses the efficacy of treatment for patients with malignant tumors based on the red blood cell DNA concentration data;
[0094] 303: Output the prediction results.
[0095] Figure 5This is a schematic diagram of the system structure for predicting the treatment efficacy of patients with malignant tumors, as provided by the present invention.
[0096] The system is programmed or otherwise configured to include a data acquisition module 401, a data processing module 402, and an output module 403.
[0097] Data acquisition module 401: Used to acquire data on the concentration of red blood cell DNA in samples from patients with malignant tumors to be tested;
[0098] Data processing module 402: used to classify and predict the efficacy prediction model obtained by the construction method described in the third aspect of the present invention on the data obtained by the data acquisition module, and obtain the classification results of the efficacy of the malignant tumor patients to be tested;
[0099] Output module 403: Used to output classification results.
[0100] A schematic diagram of the structure of the computer device provided by the present invention.
[0101] The computer device 600 includes a processor 601 and a memory 602 coupled to the processor 601. The memory 602 stores program instructions, which, when executed by the processor 601, cause the processor 601 to perform the aforementioned method for diagnosing malignant tumors or predicting the treatment efficacy for patients with malignant tumors. The processor 601 may also be referred to as a CPU (Central Processing Unit). The processor 601 may be an integrated circuit chip with signal processing capabilities. The processor 601 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or any conventional processor.
[0102] Computer device 600 can be a mobile electronic device.
[0103] It should be understood that the systems, apparatuses, and methods described in this invention can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, apparatuses, or modules, and may be electrical, mechanical, or other forms.
[0104] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0105] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0106] The above are merely embodiments of this application and do not limit the scope of this patent application. Any equivalent structural or procedural changes made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of this application.
Claims
1. A method for diagnosing lung cancer based on characteristics of DNA of mature red blood cells, characterized by, The method is completed by a computer, and the method comprises the following steps: acquiring data: acquiring data of red blood cell DNA characteristics in a sample of a subject to be tested, the DNA characteristics comprising one or more of the following: DNA concentration, DNA fragment length, DNA length coefficient of variation; processing data: inputting the data of the red blood cell DNA characteristics into a constructed lung cancer diagnosis model, the lung cancer diagnosis model predicting whether the subject to be tested has lung cancer based on the data of the red blood cell DNA characteristics; outputting a result; The lung cancer diagnosis model obtains a result by the following standard: when the DNA concentration or the DNA length coefficient of variation is higher than a threshold value or more than 50% of the DNA fragment length does not fall within a reference value range, a classification result that the subject to be tested has lung cancer is obtained; if the DNA concentration or the DNA length coefficient of variation is lower than a threshold value or more than 50% of the DNA fragment length is within a reference value range, a classification result that the subject to be tested does not have lung cancer is obtained; the reference value range of the DNA fragment length is 1 kb to 8 kb.
2. The method of claim 1, wherein, The construction steps of the lung cancer diagnosis model are as follows: acquiring data of red blood cell DNA characteristics, the DNA characteristics comprising one or more of the following: DNA concentration, DNA fragment length, DNA length coefficient of variation; the data of the red blood cell DNA characteristics come from normal healthy people and lung cancer patients; inputting the data of the red blood cell DNA characteristics into a machine learning algorithm to construct a lung cancer diagnosis model.
3. The method of claim 2, wherein, The construction steps of the lung cancer diagnosis model further comprise verifying model efficiency, and the method for verifying model efficiency comprises an ROC curve.
4. A system for diagnosing lung cancer based on characteristics of mature red blood cell DNA, characterized by comprising: The system comprises: a data acquisition unit: used for acquiring data of red blood cell DNA characteristics in a sample of a subject to be tested, the DNA characteristics comprising one or more of the following: DNA concentration, DNA fragment length, DNA length coefficient of variation; a data classification unit: used for classifying and predicting, by a lung cancer diagnosis model obtained by the method of claim 2 or 3, the data obtained by the data acquisition unit, to obtain a classification result of whether the subject to be tested has lung cancer; a result output unit: used for outputting the classification result.
5. A method for predicting the therapeutic effect of lung cancer patients based on the concentration of red blood cell DNA, characterized in that, The method is completed by a computer, and the method comprises the following steps: acquiring data: acquiring data of red blood cell DNA concentration in a sample of a lung cancer patient to be tested; processing data: inputting the data of the red blood cell DNA concentration into a constructed efficacy prediction model, the efficacy prediction model diagnosing the efficacy of the lung cancer patient based on the data of the red blood cell DNA concentration; outputting a prediction result; The efficacy prediction model obtains a prediction result by the following standard: when the red blood cell DNA concentration is higher than a threshold value, a classification result that the lung cancer patient has a poor disease control rate after treatment is obtained; when the red blood cell DNA concentration is lower than a threshold value, a classification result that the lung cancer patient has a high disease control rate after treatment is obtained.
6. The method of claim 5, wherein, The construction steps of the efficacy prediction model are as follows: acquiring data of red blood cell DNA concentration, the data of the red blood cell DNA concentration coming from a population with a high disease control rate after treatment and a population with a poor disease control rate after treatment; inputting the data of the red blood cell DNA concentration into a machine learning algorithm to construct an efficacy prediction model.
7. The method of claim 6, wherein, The constructing step of the efficacy prediction model further comprises verifying model efficiency, and the method for verifying model efficiency comprises an ROC curve. 8.A system for predicting the efficacy of a lung cancer patient based on the concentration of red blood cell DNA, characterized in that, The system comprises: a data acquisition module for acquiring data of red blood cell DNA concentration in a lung cancer patient sample to be tested; a data processing module for performing classification prediction on the efficacy prediction model obtained by the method of claim 6 or 7 on the data obtained by the data acquisition module, to obtain a classification result of the efficacy of the lung cancer patient to be tested; an output result module for outputting the classification result.
9. A computer device, comprising: The computer device comprises a memory and a processor, the memory is used to store program instructions, and the processor is used to call the program instructions, and when the program instructions are executed, the method for diagnosing lung cancer according to any one of claims 1-3 or the method for predicting the efficacy of a lung cancer patient according to any one of claims 5-7 is implemented.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and when the computer program is executed by the processor, the method for diagnosing lung cancer according to any one of claims 1-3 or the method for predicting the efficacy of a lung cancer patient according to any one of claims 5-7 is implemented.
Citation Information
Patent Citations
Cancer non-invasive early screening method and system based on cfDNA fragment length distribution characteristics
CN117316278A
Method for diagnosing or predicting cancer occurrence
KR1020220131737A