Method for constructing disease risk alert model
By constructing a disease risk indication model based on cfDNA and TSS region coverage patterns, the problems of complex operation, high cost and low applicability in existing technologies are solved, and a simple, low-cost and highly accurate disease risk indication, especially cancer risk indication, is achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SHENZHEN HUADA GENE INST
- Filing Date
- 2024-12-05
- Publication Date
- 2026-06-11
Smart Images

Figure PCTCN2024137209-FTAPPB-I100001 
Figure PCTCN2024137209-FTAPPB-I100002 
Figure PCTCN2024137209-FTAPPB-I100003
Abstract
Description
Methods for constructing disease risk warning models Technical Field
[0001] This disclosure relates to the field of biotechnology, specifically to a method for constructing a disease risk warning model and a method and apparatus for warning of disease risk, electronic devices, computer-readable storage media, computer program products, and computer programs. Background Technology
[0002] Cell-free DNA (cfDNA) is fragmented DNA released into the bloodstream after cell death in various tissues of the human body. The development and treatment of many diseases, such as cancer, autoimmune diseases, and sepsis, affect cell death rates, thereby impacting the cfDNA component in the blood. Therefore, abnormal changes in plasma cfDNA composition can reveal disease-related alterations. This change can be manifested in the coverage near the cfDNA transcriptional start site (TSS).
[0003] Currently, cancer diagnostic technologies based on cell-free DNA (cfDNA) mainly rely on two types of cfDNA features: methylation-based and fragmentomics-based technologies. However, methylation-based technologies require large initial plasma volumes, and the bisulfite treatment process is complex, costly, and difficult to implement clinically. It can also cause cfDNA fragmentation, making the data unsuitable for other analytical dimensions. Furthermore, it requires obtaining large amounts of both cancerous and healthy tissue samples for methylation marker screening. Fragmentomics-based cancer diagnostic technologies, on the other hand, have low sensitivity and specificity for early-stage cancers and are primarily developed for specific cancers, with poor applicability to pan-cancer applications.
[0004] Therefore, there is an urgent need for a method to construct a disease risk indication model that has the advantages of being easy to operate, low cost, high accuracy, wide applicability, efficient analysis process and high clinical feasibility. Summary of the Invention
[0005] This disclosure aims to at least partially address one of the technical problems in the related art.
[0006] Therefore, a first aspect of this disclosure proposes a method for constructing a disease risk warning model, comprising: obtaining a first input feature based on genomic data of cell-free DNA from a sample; obtaining a second input feature based on the disease phenotype of the sample; and constructing the disease risk warning model based on the first and second input features, wherein the genomic data includes the fragment length of the cell-free DNA, the sequence information of the cell-free DNA, and the position information of the cell-free DNA aligned to a reference genome. The method provided by this disclosure, based on the coverage pattern of cell-free DNA molecules in the TSS region, can construct a disease risk warning model with advantages such as ease of operation, low cost, high accuracy, wide applicability, efficient analysis process, and high clinical feasibility.
[0007] In some embodiments, the first input feature for obtaining genomic data based on cell-free DNA from a sample includes: obtaining cell-free DNA aligned to a transcription start site (TSS) region on an autosome in the reference genome based on the location information of the cell-free DNA aligned to the reference genome; and calculating a TSS region coverage score for each TSS region as the first input feature based on the number of covered bases of the cell-free DNA corresponding to each TSS region and the total number of covered bases of each TSS region.
[0008] In some embodiments, the equation for calculating the TSS area coverage score is:
[0009] Where h is the total number of TSS regions in the whole genome, Base_count(i) is the number of bases covered by the cell-free DNA of the i-th TSS region, Base_count(j) is the total number of bases covered by the cell-free DNA of the j-th TSS region among the h TSS regions, and Cov_score(i) is the TSS region coverage score of the i-th TSS region.
[0010] In some embodiments, h is 38,865.
[0011] In some embodiments, the first input feature for obtaining genomic data based on cell-free DNA from a sample further includes: filtering cell-free DNA fragments with a length in the range of 0-600 bp based on the fragment length of the cell-free DNA for use in calculating the first input feature.
[0012] In some embodiments, the fragment length is in the range of 150 to 210 bp.
[0013] In some embodiments, constructing the disease risk warning model based on the first input feature and the second input feature includes: constructing an initial model based on the first input feature and the second input feature to obtain a third input feature; sorting the third output feature to obtain a fourth input feature; and constructing the disease risk warning model based on the fourth input feature.
[0014] In some embodiments, the construction of an initial model based on the first input feature and the second input feature to obtain a third input feature includes: based on the initial model being a machine learning model, using the first input feature as an independent variable and the second input feature as a dependent variable, obtaining the weight coefficients of the first input feature of the best-fit initial model as the third input feature.
[0015] In some embodiments, the initial model is a binary classification model, based on whether the sample has a disease, using the second input feature as the basis.
[0016] In some embodiments, the initial model is a linear regression model, based on the second input feature being a continuous disease assessment index of the sample.
[0017] In some embodiments, sorting the third input features to obtain the fourth input features includes: sorting the third input features corresponding to each of the first input features; taking the first 1-N third input features and their corresponding first input features as the fourth input features based on the third input feature being greater than 0; and taking the last 1-N third input features and their corresponding first input features as the fourth input features based on the third input feature being less than 0.
[0018] In some embodiments, N is 2000-5000, more preferably, N is 4000.
[0019] In some embodiments, the construction of the disease risk warning model based on the fourth input feature includes: constructing the disease risk warning model to obtain a disease risk index based on the fourth input feature and its corresponding second input feature, since the disease risk warning model is a machine learning model; optionally, it further includes calculating the 95th percentile based on the disease risk index of the healthy group in the sample.
[0020] In some embodiments, the TSS region is a 5Kb region upstream and downstream of the TSS.
[0021] In some embodiments, the TSS region is a 3Kb region upstream and downstream of the TSS.
[0022] In some embodiments, the TSS region is a region 1.5Kb upstream and downstream of the TSS.
[0023] In some embodiments, the TSS region is a region 1Kb upstream and downstream of the TSS.
[0024] In some embodiments, the first input feature for obtaining genomic data based on cell-free DNA from a sample further includes: calculating the sequencing read coverage depth of the cell-free DNA aligned to the TSS region of an autosome in the reference genome.
[0025] In some embodiments, the sample is derived from at least one of plasma, seminal plasma, saliva, urine, amniotic fluid, serum, pleural effusion, peritoneal fluid, cerebrospinal fluid, and synovial fluid; preferably, the sample is derived from plasma.
[0026] In some embodiments, the disease is at least one of pregnancy-related diseases, neurodegenerative diseases, cancer, and autoimmune diseases.
[0027] A second aspect of this disclosure provides a method for alerting to disease risk, comprising:
[0028] The genomic data of cell-free DNA from the test sample and the reference sample are input into a disease risk indication model obtained according to any embodiment of the first aspect of this disclosure; the disease risk indices of the test sample and the reference sample output by the disease risk indication model are compared; and based on a significant difference between the disease risk indices of the test sample and the reference sample, it is indicated that the individual from which the test sample was sampled has a disease risk.
[0029] A third aspect of this disclosure provides a method for indicating disease risk, comprising: inputting genomic data of cell-free DNA of a sample to be tested into a disease risk indication model obtained according to any embodiment of the first aspect of this disclosure; comparing a disease risk index of the sample to be tested output by the disease risk indication model with a threshold; and indicating that the individual who sampled the sample has a disease risk based on the disease risk index of the sample being higher than the threshold, optionally, the threshold being the 95th percentile calculated based on the disease risk index of the healthy population in the sample, optionally, the threshold being 0.5.
[0030] In some embodiments, the disease is at least one of pregnancy-related diseases, neurodegenerative diseases, cancer, and autoimmune diseases.
[0031] An embodiment of the fourth aspect of this disclosure provides an apparatus for constructing a disease risk warning model, comprising: a first module, the first module being used to obtain a first input feature based on genomic data of cell-free DNA of a sample, and to obtain a second input feature based on a disease phenotype of the sample; and a second module, the second module being used to construct the disease risk warning model based on the first input feature and the second input feature.
[0032] In some embodiments, the second module includes a third module, a fourth module, and a fifth module. The third module is used to construct an initial model based on the first input feature and the second input feature to obtain a third input feature. The fourth module is used to sort the third output feature to obtain a fourth input feature. The fifth module is used to construct the disease risk warning model based on the fourth input feature.
[0033] A fifth aspect of this disclosure provides a disease risk alert device, comprising: an input module for inputting genomic data of cell-free DNA from a test sample and a reference sample into a disease risk alert model obtained according to any embodiment of the first aspect of this disclosure; and an alert module for comparing disease risk indices of the test sample and the reference sample output by the disease risk alert model; and alerting the individual who sampled the test sample to have a disease risk based on a significant difference between the disease risk indices of the test sample and the reference sample.
[0034] In some embodiments, a second prompting module is further included, which is used to compare the disease risk index of the test sample output by the disease risk prompting model with a threshold; and to prompt the sampled individual of the test sample that the disease risk index of the test sample is higher than the threshold, wherein the threshold is optionally the 95th percentile calculated based on the disease risk index of the healthy group in the sample, and optionally the threshold is 0.5.
[0035] A sixth aspect of this disclosure provides an electronic device comprising: a processor; and a memory for storing executable instructions; wherein the processor is configured to read the executable instructions from the memory and execute the executable instructions to implement a method for constructing a disease risk warning model as described in any embodiment of the first aspect of this disclosure or a method for warning of disease risk as described in any embodiment of the second or third aspect of this disclosure.
[0036] A seventh aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement a method for constructing a disease risk warning model as described in any embodiment of the first aspect of this disclosure, or a method for warning of disease risk as described in any embodiment of the second or third aspect of this disclosure.
[0037] An embodiment of the eighth aspect of this disclosure provides a computer program product, wherein the computer program product includes a computer program that, when executed by a processor, implements a method for constructing a disease risk warning model as described in any embodiment of the first aspect of this disclosure, or a method for warning of disease risk as described in any embodiment of the second or third aspect of this disclosure.
[0038] An embodiment of the ninth aspect of this disclosure provides a computer program, wherein the computer program includes computer program code, which, when run on a computer, causes the computer to perform a method for constructing a disease risk warning model as described in any embodiment of the first aspect of this disclosure or a method for warning of disease risk as described in any embodiment of the second or third aspect of this disclosure.
[0039] The advantages and technical effects of this disclosure are as follows:
[0040] (1) The method provided in this disclosure does not require additional methylation conversion treatment, high starting plasma volume, or complex experimental processing; it can be achieved using only conventional cfDNA WGS sequencing, with low sequencing depth requirements, and can perform other dimension feature analyses of cfDNA while analyzing disease risk indices. Therefore, the method and model provided in this disclosure have excellent technical effects such as simple operation, high clinical feasibility, and low cost.
[0041] (2) The method provided in this disclosure can target only the upstream and downstream regions of the transcription start site (about 2-3% of WGS), so the number of sequencing samples is small (only ~0.5X is needed) and the cost is low.
[0042] (3) The method provided in this embodiment can achieve accurate disease risk indication by analyzing the coverage of the TSS region; and in practical applications, the analysis area is small, the computing resources are less, and the analysis is fast, thus ensuring the efficiency and speed of the analysis, greatly reducing the required time, and achieving a lower false positive rate.
[0043] (4) Wide applicability: The method provided in this disclosure is applicable not only to paired-end sequencing data but also to single-end sequencing data, and is not limited by sequencing read length. Therefore, it can adapt to data from various sequencing platforms and sequencing modes, and has broad application prospects.
[0044] In summary, the method provided by the embodiments of this disclosure has the advantages of simple operation, low cost, high accuracy, wide applicability, efficient analysis process and high clinical feasibility, providing a new and powerful tool for disease risk warning, especially cancer risk warning. Attached Figure Description
[0045] Figure 1 is a schematic diagram of the distribution of disease risk indices obtained by the disease risk warning model constructed by the method of this embodiment in an internal test set.
[0046] Figure 2 is a schematic diagram of the ROC analysis results of the disease risk warning model constructed by the method of this embodiment of the present disclosure in distinguishing cancer in an internal test set.
[0047] Figure 3 is a schematic diagram of the distribution of disease risk indices obtained by the disease risk warning model constructed by the method of this embodiment in validation set 1.
[0048] Figure 4 is a schematic diagram of the ROC analysis results of the disease risk warning model constructed by the method of this embodiment in validation set 1.
[0049] Figure 5 is a schematic diagram showing the distribution of disease risk indices obtained in validation set 2 by the disease risk warning model constructed by the method of this embodiment.
[0050] Figure 6 is a schematic diagram of the ROC analysis results of the disease risk warning model constructed by the method of this embodiment in validation set 2.
[0051] Figure 7 is a schematic diagram showing the distribution of disease risk indices obtained in validation set 3 by the disease risk warning model constructed by the method of this embodiment.
[0052] Figure 8 is a schematic diagram of the ROC analysis results of the disease risk warning model constructed by the method of this embodiment in validation set 3.
[0053] Figure 9 is a flowchart illustrating a method for constructing a disease risk warning model according to an embodiment of this disclosure. Detailed Implementation
[0054] Embodiments of this disclosure are described in detail below, with examples of these embodiments illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting it.
[0055] Cell-free DNA (cfDNA) in plasma is fragmented DNA released into the peripheral blood after cell death from various tissues and organs in the human body. The degree of chromatin openness at transcription start sites on the genome is related to gene expression. TSS regions of highly expressed genes are usually in an open chromatin state. Due to the lack of nucleosome protection, cfDNA is degraded and further fragmented, making it difficult to sequence. Therefore, cfDNA shows low coverage in the TSS regions of highly expressed genes. Conversely, TSS regions of genes with low expression levels are usually in a closed chromatin state. Due to the protection of nucleosomes, cfDNA is not easily degraded and sequenced, thus showing high coverage in the TSS regions of low-expression genes. Various tissues and organs in the human body can contribute cfDNA to plasma. Gene expression in cancer tissues exhibits specificity (e.g., dysregulation of proto-oncogene and tumor suppressor gene expression). Therefore, analyzing the TSS coverage of various genes in plasma holds promise for capturing tumor signals.
[0056] Currently, cancer diagnostic technologies based on cell-free DNA (cfDNA) mainly rely on two types of cfDNA features: methylation-based and fragmentomics-based technologies. However, methylation-based technologies require large initial plasma volumes, and the bisulfite treatment process is complex, costly, and difficult to implement clinically. It can also cause cfDNA fragmentation, making the data unsuitable for other analytical dimensions. Furthermore, it requires obtaining large amounts of both cancerous and healthy tissue samples for methylation marker screening. Fragmentomics-based cancer diagnostic technologies, on the other hand, have low sensitivity and specificity for early-stage cancers and are primarily developed for specific cancers, with poor applicability to pan-cancer applications.
[0057] Therefore, the first aspect of this disclosure proposes a method for constructing a disease risk indication model. By analyzing the differences in cfDNA coverage in the TSS region between healthy controls and cancer patients, cancer-related TSS regions can be identified. Subsequently, a classifier is constructed using the cfDNA coverage pattern in the TSS region as a feature for the diagnosis of various cancers. The method provided by this disclosure, based on the coverage pattern of free DNA molecules in the TSS region, can construct a disease risk indication model with advantages such as ease of operation, low cost, high accuracy, wide applicability, efficient analysis process, and high clinical feasibility.
[0058] Specifically, the method includes: obtaining a first input feature based on genomic data of cell-free DNA from a sample; obtaining a second input feature based on the disease phenotype of the sample; and constructing the disease risk warning model based on the first input feature and the second input feature, wherein the genomic data includes the fragment length of the cell-free DNA, the sequence information of the cell-free DNA, and the position information of the cell-free DNA aligned to a reference genome.
[0059] In some embodiments, the first input feature for obtaining genomic data based on cell-free DNA from a sample includes: obtaining cell-free DNA aligned to a transcription start site (TSS) region on an autosome in the reference genome based on the location information of the cell-free DNA aligned to the reference genome; and calculating a TSS region coverage score for each TSS region as the first input feature based on the number of covered bases of the cell-free DNA corresponding to each TSS region and the total number of covered bases of each TSS region.
[0060] It is understood that this method may also include the step of aligning raw sequencing data to obtain the cell-free DNA aligned to the transcription start site (TSS) region of an autosome in the reference genome. In some embodiments, the reference genome described in this application is hg38.
[0061] In some embodiments, the equation for calculating the TSS area coverage score is:
[0062] Where h is the total number of TSS regions in the whole genome, Base_count(i) is the number of bases covered by the cell-free DNA of the i-th TSS region, Base_count(j) is the total number of bases covered by the cell-free DNA of the j-th TSS region among the h TSS regions, and Cov_score(i) is the TSS region coverage score of the i-th TSS region.
[0063] In some embodiments, h is 38,865.
[0064] It is understandable that calculating the standardized TSS regional coverage score using the above formula can improve the reliability and performance of the data, and the standardized data can be applied to subsequent algorithm analysis, thereby achieving more accurate disease risk indications.
[0065] In some embodiments, the first input feature for obtaining genomic data based on cell-free DNA from a sample further includes: filtering cell-free DNA fragments with a length in the range of 0-600 bp based on the fragment length of the cell-free DNA for use in calculating the first input feature.
[0066] In some embodiments, the fragment length is in the range of 150 to 210 bp.
[0067] In this embodiment, by setting a length threshold (standard) for cfDNA aligned to the TSS region and screening the cfDNA matched to each TSS region based on this standard, cfDNA from open chromatin regions originating from lesion sites can be effectively obtained. Therefore, TSS fragments of different lengths can be selected to achieve optimal results. This disclosure is not intended to limit the scope of the invention.
[0068] In some embodiments, constructing the disease risk warning model based on the first input feature and the second input feature includes: constructing an initial model based on the first input feature and the second input feature to obtain a third input feature; sorting the third output feature to obtain a fourth input feature; and constructing the disease risk warning model based on the fourth input feature.
[0069] In some embodiments, constructing an initial model based on the first input feature and the second input feature to obtain a third input feature includes: based on the initial model being a machine learning model, using the first input feature as an independent variable and the second input feature as a dependent variable, obtaining the weight coefficients of the first input feature of the best-fitting initial model as the third input feature. The machine learning model is a model with classification capabilities, such as a logistic regression model, support vector machine, random forest, etc., and more preferably, a linear regression model.
[0070] In some embodiments, the initial model is a binary classification model, based on whether the sample has a disease, using the second input feature as the basis.
[0071] In some embodiments, based on the second input feature being the disease severity of the sample, the initial model is a multi-class logistic regression model. For example, the disease is classified into, for example, "Level I, II, III" according to the disease severity.
[0072] In some embodiments, the initial model is a linear regression model, based on the second input feature being a continuous disease assessment index of the sample.
[0073] In some embodiments, sorting the third input features to obtain the fourth input features includes: sorting the third input features corresponding to each of the first input features; taking the first 1-N third input features and their corresponding first input features as the fourth input features based on the third input feature being greater than 0; and taking the last 1-N third input features and their corresponding first input features as the fourth input features based on the third input feature being less than 0.
[0074] In some embodiments, N is 2000-5000, more preferably, N is 4000.
[0075] In some embodiments, the construction of the disease risk warning model based on the fourth input feature includes: constructing the disease risk warning model to obtain a disease risk index based on the fact that the disease risk warning model is a machine learning model, the fourth input feature and its corresponding second input feature; optionally, it further includes calculating the 95th percentile based on the disease risk index of the healthy group in the sample.
[0076] In some embodiments, the TSS region is a 5Kb region upstream and downstream of the TSS.
[0077] In some embodiments, the TSS region is a 3Kb region upstream and downstream of the TSS.
[0078] In some embodiments, the TSS region is a region 1.5Kb upstream and downstream of the TSS.
[0079] In some embodiments, the TSS region is a region 1Kb upstream and downstream of the TSS.
[0080] It is understood that the transcription start site (TSS) is the starting point of transcriptional activity, located at the 3' end of the transcription unit. The transcription start site region is the upstream and downstream region centered on the TSS, containing various promoters, transcription factors, and regulatory sequences that work together to regulate gene transcriptional expression. Based on the methods of the embodiments of this disclosure, those skilled in the art can expand or shrink the range of the TSS region in different samples to obtain optimal results. This disclosure is not intended to be limiting.
[0081] In some embodiments, the first input feature for acquiring genomic data based on cell-free DNA from the sample further includes: calculating the first input feature by comparing the sequencing read coverage depth of the cell-free DNA aligned to the TSS region of an autosome in the reference genome. It is understood that any indicator that can be used to analyze the distribution of cfDNA in the TSS region (such as coverage / depth) to reflect the chromatin open or closed state of the TSS region (e.g., sequencing depth of the TSS region, total sequencing depth, different normalized depths, or fragment length diversity) can be used as an indicator to assess disease risk. Furthermore, the method provided in this disclosure can target only the regions upstream and downstream of the transcription start site (approximately 2-3% of the WGS), thus requiring fewer sequencing reads (only ~0.5X), significantly reducing sequencing costs.
[0082] In some embodiments, the sample is derived from at least one of plasma, seminal plasma, saliva, urine, amniotic fluid, serum, pleural effusion, peritoneal fluid, cerebrospinal fluid, and synovial fluid; preferably, the sample is derived from plasma.
[0083] In some embodiments, the disease is at least one of pregnancy-related diseases, neurodegenerative diseases, cancer, and autoimmune diseases.
[0084] In some embodiments, the cancer is selected from one or more cancers including lung cancer, head and neck cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, gallbladder cancer, colorectal cancer, breast cancer, cervical cancer, ovarian cancer, bile duct cancer, endometrial cancer, prostate cancer, kidney cancer, bladder cancer, urethral cancer, thyroid cancer, skin cancer, malignant melanoma, lymphoma, leukemia, osteosarcoma, laryngeal cancer, nasopharyngeal cancer, oral cancer, pharyngeal cancer, eye cancer, vaginal cancer, and penile cancer.
[0085] In some embodiments, a flowchart illustrating the construction method of the disease risk warning model is shown in Figure 9. Disease risk warning model one is constructed by inputting the TSS region coverage score (TSS cov_score) and disease phenotype. This step can output a disease risk index. Alternatively, based on disease risk warning model one, the coefficients and rankings of each feature are further obtained for feature selection. In other embodiments, the selected TSS region coverage score and disease phenotype are input into an optimized disease risk warning model two to obtain the disease risk index.
[0086] A second aspect of this disclosure provides a method for alerting to disease risk, comprising:
[0087] The genomic data of cell-free DNA from the test sample and the reference sample are input into a disease risk indication model obtained according to any embodiment of the first aspect of this disclosure; the disease risk indices of the test sample and the reference sample output by the disease risk indication model are compared; and based on a significant difference between the disease risk indices of the test sample and the reference sample, it is indicated that the individual from which the test sample was sampled has a disease risk.
[0088] A third aspect of this disclosure provides a method for indicating disease risk, comprising: inputting genomic data of cell-free DNA from a sample to be tested into a disease risk indication model obtained according to any embodiment of the first aspect of this disclosure; comparing a disease risk index of the sample to be tested output by the disease risk indication model with a threshold; and indicating that the individual sampling the sample has a disease risk based on the disease risk index of the sample being higher than the threshold. Optionally, the threshold is the 95th percentile calculated based on the disease risk index of a healthy population in the sample, and optionally, the threshold is 0.5. It should be noted that the disease risk index may exhibit inconsistent distribution ranges due to different sequencing strategies and library preparation methods. Typically, when the training set and the sample set to be tested use the same library preparation and sequencing strategy, 0.5 can be used as the threshold; if the library preparation and sequencing strategies of the training set and the sample set to be tested are not completely consistent, the 95th percentile of the disease risk index in a known healthy population can be used as the threshold (i.e., ensuring 95% specificity).
[0089] In some embodiments, the disease is at least one of pregnancy-related diseases, neurodegenerative diseases, cancer, and autoimmune diseases.
[0090] An embodiment of the fourth aspect of this disclosure provides an apparatus for constructing a disease risk warning model, comprising: a first module, the first module being used to obtain a first input feature based on genomic data of cell-free DNA of a sample, and to obtain a second input feature based on a disease phenotype of the sample; and a second module, the second module being used to construct the disease risk warning model based on the first input feature and the second input feature.
[0091] In some embodiments, the second module includes a third module, a fourth module, and a fifth module. The third module is used to construct an initial model based on the first input feature and the second input feature to obtain a third input feature. The fourth module is used to sort the third output feature to obtain a fourth input feature. The fifth module is used to construct the disease risk warning model based on the fourth input feature.
[0092] A fifth aspect of this disclosure provides a disease risk alert device, comprising: an input module for inputting genomic data of cell-free DNA from a test sample and a reference sample into a disease risk alert model obtained according to any embodiment of the first aspect of this disclosure; and an alert module for comparing disease risk indices of the test sample and the reference sample output by the disease risk alert model; and alerting the individual who sampled the test sample to have a disease risk based on a significant difference between the disease risk indices of the test sample and the reference sample.
[0093] In some embodiments, a second prompting module is further included, which is used to compare the disease risk index of the test sample output by the disease risk prompting model with a threshold; and to prompt the sampler of the test sample that the disease risk index of the test sample is higher than the threshold based on the fact that the disease risk index of the test sample is higher than the threshold.
[0094] A sixth aspect of this disclosure provides an electronic device comprising: a processor; and a memory for storing executable instructions; wherein the processor is configured to read the executable instructions from the memory and execute the executable instructions to implement a method for constructing a disease risk warning model as described in any embodiment of the first aspect of this disclosure or a method for warning of disease risk as described in any embodiment of the second or third aspect of this disclosure.
[0095] A seventh aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement a method for constructing a disease risk warning model as described in any embodiment of the first aspect of this disclosure, or a method for warning of disease risk as described in any embodiment of the second or third aspect of this disclosure.
[0096] An embodiment of the eighth aspect of this disclosure provides a computer program product, wherein the computer program product includes a computer program that, when executed by a processor, implements a method for constructing a disease risk warning model as described in any embodiment of the first aspect of this disclosure, or a method for warning of disease risk as described in any embodiment of the second or third aspect of this disclosure.
[0097] An embodiment of the ninth aspect of this disclosure provides a computer program, wherein the computer program includes computer program code, which, when run on a computer, causes the computer to perform a method for constructing a disease risk warning model as described in any embodiment of the first aspect of this disclosure or a method for warning of disease risk as described in any embodiment of the second or third aspect of this disclosure.
[0098] It should be noted that the foregoing explanations of the method embodiments also apply to the apparatus, electronic devices, computer-readable storage media, computer program products, and computer programs described above, and will not be repeated here.
[0099] The following will explain in detail, with reference to specific embodiments, the construction method of the disease risk warning model proposed in this disclosure and its related applications.
[0100] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0101] Unless otherwise specified, the quantitative experiments in the following examples are all repeated three times, and the results are averaged.
[0102] Example 1: Construction of a Disease Risk Warning Model
[0103] This disclosure presents an embodiment of a disease risk warning model based on publicly available plasma cfDNA whole-genome sequencing data, and verifies its application effect on an internal dataset.
[0104] 1.1 Sample Grouping
[0105] Whole-genome sequencing data of plasma cfDNA stored in BED format from 523 cases were downloaded from the FinaleDB database (http: / / finaledb.research.cchmc.org / ). These included 254 healthy controls, 74 lung cancer cases, 54 breast cancer cases, 35 pancreatic cancer cases, 28 ovarian cancer cases, 27 gastric cancer cases, 26 colorectal cancer cases, and 25 bile duct cancer cases. 70% of these 523 samples were randomly selected as the training set, and the remaining 30% was used as an internal test set.
[0106] 1.2 Screening for cfDNA
[0107] Whole-genome sequencing data of plasma cfDNA from samples stored in BED format includes cfDNA fragment length, cfDNA sequence information, and cfDNA alignment to the reference genome. Based on this information, cfDNA fragments aligned within a 1000bp region upstream and downstream of the transcription start site, with a fragment length between 150 and 210bp, were selected.
[0108] 1.3 Obtaining the First Input Features
[0109] The number of cfDNA bases covered by each of the 38,865 TSS regions on autosomes and the total number of covered bases by TSS regions on autosomes were statistically analyzed. A standardized TSS coverage score was calculated for each TSS region and used as the first input feature. The equation for calculating the TSS region coverage score is as follows:
[0110] h is the total number of TSS regions in the whole genome, Base_count(i) is the number of cell-free DNA bases covered by the i-th TSS region, Base_count(j) is the total number of cell-free DNA bases covered by the j-th TSS region among the h TSS regions, and Cov_score(i) is the TSS region coverage score of the i-th TSS region.
[0111] 1.4 Obtaining the Second Input Features
[0112] The samples are labeled as either cancer samples (1) or non-cancer samples (0) based on whether the individuals from which they are derived have cancer. This binary classification label is used as the second input feature.
[0113] 1.5 Obtaining Third Input Features
[0114] Using the Cov_score(i) of the aforementioned 38,865 TSS regions as independent variables and the binary classification label of whether an individual has the disease as the dependent variable, a linear regression model was established, and the weight coefficients corresponding to each Cov_score(i) in the best-fit linear regression model were extracted. Based on the magnitude of these weight coefficients, the TSS regions were sorted, and the 4,000 features with the lowest negative coefficients and the 4,000 features with the highest positive coefficients were selected, ultimately retaining 8,000 features as the third input features.
[0115] 1.6 Building the Model
[0116] Based on the 8,000 features retained above, a disease risk warning model was reconstructed using a linear regression model. In this model, the output is defined as the "disease risk index". The higher the disease risk index, the higher the risk of the subject developing a tumor.
[0117] 1.7 Model Validation in the Internal Test Set
[0118] Based on the samples in the internal test set divided during sample grouping, the TSS coverage score is calculated as the first input feature, and the disease status of the samples is used as the second input feature. Corresponding disease risk indices are obtained for lung cancer, breast cancer, pancreatic cancer, ovarian cancer, gastric cancer, colorectal cancer, and bile duct cancer samples, respectively. A disease risk index greater than 0.5 is considered to indicate a risk of disease. Therefore, ROC analysis is performed based on the model output and the actual sample disease status to evaluate the performance of the disease risk indication model provided in this embodiment.
[0119] Figure 1 is a schematic diagram showing the distribution of disease risk indices obtained by the disease risk warning model constructed by the method of this disclosure in an internal test set. Figure 1 shows that the disease risk indices of healthy controls are significantly lower than those of cancer samples (P<0.0001). The disease risk warning model of this disclosure can effectively indicate the risk of various cancers.
[0120] Figure 2 is a schematic diagram of the ROC analysis results of the disease risk warning model constructed by the method of this disclosure embodiment in distinguishing various cancers on an internal test set. Figure 2 shows that the disease risk warning model of this disclosure embodiment can effectively distinguish between cancer and healthy controls, and its area under the curve (AUC) is higher than 0.92. The AUCs for bile duct cancer, ovarian cancer, colorectal cancer, gastric cancer, lung cancer, breast cancer, and pancreatic cancer are 0.99, 0.98, 0.97, 0.95, 0.94, 0.93, and 0.92, respectively, indicating its accuracy and reliability. Based on the fact that in practical applications, an AUC value greater than 0.5 is generally considered to be an effective model, and an AUC value close to 1 indicates that the model has extremely high predictive accuracy, the disease risk warning model provided by this application embodiment is considered to have excellent performance.
[0121] Model Validation in Validation Set 1 of Example 2
[0122] This embodiment collected 59 plasma cfDNA samples from 49 breast cancer patients and 10 healthy controls, and performed high-throughput sequencing (DNBSEQ sequencing platform, paired-end sequencing, 100bp read length, sequencing depth range: 25X-101X). The sequencing results were input into the disease risk indication model constructed in Example 1 to calculate the disease risk index.
[0123] Figure 3 is a schematic diagram of the distribution of disease risk indices obtained by the disease risk warning model constructed by the method of this embodiment in validation set 1. Figure 3 shows that the disease risk index of the sample from breast cancer patients is significantly higher than that of healthy controls (P<0.01). Figure 4 is a schematic diagram of the ROC analysis results of the disease risk warning model constructed by the method of this embodiment in validation set 1, which shows that the AUC of the disease risk index obtained by the disease risk warning model in distinguishing between breast cancer samples and healthy controls is 0.80. The disease risk warning model constructed by the method of this embodiment can effectively and accurately distinguish between breast cancer and healthy controls.
[0124] Model Validation in Validation Set 2 of Example 3
[0125] This embodiment collected 30 plasma cfDNA samples from 10 healthy controls, 10 liver cancer patients, and 10 lung cancer patients, and performed high-throughput sequencing (DNBSEQ sequencing platform, paired-end sequencing, 100bp read length, sequencing depth: 2.1-6.2X). The sequencing results were input into the disease risk indication model constructed in Example 1 to calculate the disease risk index.
[0126] Figure 5 is a schematic diagram of the distribution of disease risk indices obtained by the disease risk warning model constructed by the method of this embodiment in validation set 2. Figure 5 shows that the disease risk indices of liver cancer and lung cancer samples are higher than those of the healthy control group (P value after BH correction < 0.001). Figure 6 is a schematic diagram of the ROC analysis results of the disease risk warning model constructed by the method of this embodiment in validation set 2, which shows that the AUC of the disease risk index obtained by the disease risk warning model is 0.91 in distinguishing between liver cancer and healthy controls, and 0.96 in distinguishing between lung cancer and healthy controls. The disease risk warning model constructed by the method of this embodiment can effectively and accurately distinguish between cancer samples (lung cancer, liver cancer) and healthy controls.
[0127] Model Validation in Validation Set 3 of Example 4
[0128] This embodiment collected 106 plasma cfDNA samples from 32 healthy controls and 74 liver cancer patients, and performed high-throughput sequencing (Illumina sequencing platform, paired-end sequencing, 75bp read length, sequencing depth: 1.5-2.5X). The sequencing results were input into the disease risk indication model constructed in Example 1 to calculate the disease risk index.
[0129] Figure 7 is a schematic diagram of the distribution of disease risk indices obtained by the disease risk warning model constructed by the method of this embodiment in validation set 3. Figure 7 shows that the disease risk index of liver cancer patients is significantly higher than that of healthy controls (P<0.0001). Figure 8 is a schematic diagram of the ROC analysis results of the disease risk warning model constructed by the method of this embodiment in validation set 3, which shows that the AUC of the disease risk index obtained by the disease risk warning model in distinguishing between liver cancer and healthy controls is 0.78. The disease risk warning model constructed by the method of this embodiment can effectively and accurately distinguish between liver cancer and healthy controls.
[0130] According to the above embodiments, the disease risk warning model constructed by the method of this disclosure can be applied to multiple internal and external datasets, and is suitable for effectively and accurately warning of pan-cancer (multi-cancer) disease risks.
[0131] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0132] In this disclosure, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to a specific feature, structure, material, or characteristic described in connection with that embodiment or example, which is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0133] Although embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A method for constructing a disease risk warning model, comprising: First input features are obtained from genomic data of cell-free DNA from the sample; The second input feature is obtained based on the disease phenotype of the sample; The disease risk alert model is constructed based on the first input feature and the second input feature. The genomic data includes the fragment length of the cell-free DNA, the sequence information of the cell-free DNA, and the position information of the cell-free DNA aligned to the reference genome.
2. The method according to claim 1, wherein obtaining the first input feature from the genomic data of cell-free DNA based on the sample comprises: Based on the location information of the free DNA compared to the reference genome, the free DNA compared to the transcription start site (TSS) region of the autosome in the reference genome is obtained; and Based on the number of covered bases of the free DNA corresponding to each TSS region and the total number of covered bases of each TSS region, the TSS region coverage score of each TSS region is calculated as the first input feature.
3. The method according to claim 2, wherein the equation for calculating the TSS area coverage score is: Where h is the total number of TSS regions in the whole genome, Base_count(i) is the number of cell-free DNA bases covering the i-th TSS region, Base_count(j) is the total number of cell-free DNA bases covering the j-th TSS region among the h TSS regions, and Cov_score(i) is the TSS region coverage score of the i-th TSS region. Preferably, h is 38,865.
4. The method according to any one of claims 1 to 3, wherein the first input feature for obtaining genomic data based on cell-free DNA from the sample further includes: Based on the fragment length of the free DNA, free DNA fragments with a length in the range of 0-600 bp are selected for use in calculating the first input feature. Preferably, the fragment length is in the range of 150 to 210 bp.
5. The method according to claim 1, wherein constructing the disease risk warning model based on the first input feature and the second input feature comprises: An initial model is constructed based on the first and second input features to obtain the third input feature; The third output features are sorted to obtain the fourth input features; and Based on the fourth input feature, the disease risk warning model is constructed.
6. The method of claim 5, wherein constructing an initial model based on the first input features and the second input features to obtain the third input features comprises: Since the initial model is a machine learning model, the first input feature is used as the independent variable, the second input feature is used as the dependent variable, and the weight coefficients of the first input feature of the best-fitting initial model are used as the third input feature. Optionally, based on whether the sample suffers from a disease, the initial model is a binary classification model. Optionally, based on the second input feature as a continuous disease assessment index of the sample, the initial model is a linear regression model.
7. The method according to claim 5, wherein sorting the third input features to obtain the fourth input features comprises: The third input features corresponding to each of the first input features are sorted; Based on the fact that the third input feature is greater than 0, the third input feature of the first 1-N positions in the sorting and its corresponding first input feature are taken as the fourth input feature; and Based on the fact that the third input feature is less than 0, the third input feature in the sorted order of 1-N positions and its corresponding first input feature are used as the fourth input feature. Preferably, N is 2000-5000, more preferably, N is 4000.
8. The method according to claim 5, wherein constructing the disease risk warning model based on the fourth input feature comprises: Based on the fact that the disease risk warning model is a machine learning model, the fourth input feature and its corresponding second input feature are used to construct the disease risk warning model to obtain a disease risk index. Optionally, it also includes calculating the 95th percentile based on the disease risk index of the healthy population in the sample.
9. The method according to claim 2, wherein the TSS region is a 5Kb region upstream and downstream of the TSS, preferably a 3Kb region upstream and downstream of the TSS, more preferably a 1.5Kb region upstream and downstream of the TSS, and most preferably a 1Kb region upstream and downstream of the TSS.
10. The method according to claim 1 or 2, wherein, The first input feature for obtaining genomic data based on cell-free DNA from the sample also includes: The first input feature is calculated by comparing the sequencing read coverage depth of the free DNA aligned to the TSS region of an autosome in the reference genome.
11. The method according to any one of claims 1 to 9, wherein the sample is derived from at least one of plasma, seminal plasma, saliva, urine, amniotic fluid, serum, pleural effusion, peritoneal fluid, cerebrospinal fluid, and synovial fluid, preferably, the sample is derived from plasma.
12. A method for indicating disease risk, comprising: Genomic data of cell-free DNA from the test sample and the reference sample are input into the disease risk indication model obtained by the method according to any one of claims 1 to 11; The disease risk indices of the test sample and the reference sample output by the disease risk warning model are compared. and The significant difference in the disease risk index between the test sample and the reference sample suggests that the individuals sampled for the test sample have a disease risk.
13. A method for indicating disease risk, comprising: The genomic data of cell-free DNA from the sample to be tested are input into the disease risk indication model obtained by the method according to any one of claims 1 to 11; The disease risk index and threshold of the test sample output by the disease risk warning model are compared. and If the disease risk index of the test sample is higher than the threshold, it indicates that the individual who sampled the test sample has a disease risk. Optionally, the threshold is the 95th percentile calculated based on the disease risk index of the healthy population in the sample. Optionally, the threshold is 0.
5.
14. The method according to claim 12 or 13, wherein the disease is at least one of pregnancy-related diseases, neurodegenerative diseases, cancer, and autoimmune diseases.
15. An apparatus for constructing a disease risk warning model, comprising: The first module is used to obtain a first input feature based on the genomic data of cell-free DNA of the sample, and to obtain a second input feature based on the disease phenotype of the sample. The second module is used to construct the disease risk warning model based on the first input features and the second input features. Optionally, the second module includes a third module, a fourth module, and a fifth module. The third module is used to construct an initial model based on the first input features and the second input features to obtain the third input features; The fourth module is used to sort the third output features to obtain the fourth input features; and The fifth module is used to construct the disease risk warning model based on the fourth input feature.
16. A disease risk warning device, comprising: The input module is used to input the genomic data of cell-free DNA from the test sample and the reference sample into the disease risk indication model obtained by the method according to any one of claims 1 to 11; and Prompt module, The prompting module is used to compare the disease risk index of the test sample and the reference sample output by the disease risk prompting model; and The significant difference in disease risk indices between the test sample and the reference sample suggests that the individuals sampled in the test sample have a disease risk. Optionally, a second prompt module may also be included. The second prompting module is used to compare the disease risk index and threshold of the test sample output by the disease risk prompting model; and If the disease risk index of the test sample is higher than the threshold, it indicates that the individual sampled for the test sample has a disease risk.
17. An electronic device comprising: processor; Memory, used to store executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the method for constructing a disease risk warning model as described in any one of claims 1 to 11 or the method for warning of disease risk as described in any one of claims 12 to 14.
18. A computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement a method for constructing a disease risk warning model as claimed in any one of claims 1 to 11 or a method for warning of disease risk as claimed in any one of claims 12 to 14.
19. A computer program product, wherein the computer program product includes a computer program that, when executed by a processor, implements the method for constructing a disease risk warning model as claimed in any one of claims 1 to 11 or the method for warning of disease risk as claimed in any one of claims 12 to 14.
20. A computer program, wherein the computer program includes computer program code, which, when run on a computer, causes the computer to perform a method for constructing a disease risk warning model as claimed in any one of claims 1 to 11 or a method for warning of disease risk as claimed in any one of claims 12 to 14.
Citation Information
Patent Citations
Cancer detection model and construction method and kit thereof
CN113838533A
Method and system for multi-gene genetic risk assessment of complex diseases
CN116343902A
Fetal concentration detection method and device, electronic equipment and storage medium
CN118280439A
Methods for the analysis of gene expression and uses thereof
US20240360516A1
Diagnostic method for fibromyalgia (FMS) or chronic fatigue syndrome (CFS)
WO2008010082A2