Calculation method and application of free DNA fragment length mode
By using JS divergence to characterize the length distribution of cfDNA fragments, the problem of inconsistent definitions of long and short fragments in cfDNA fragment length distribution characterization was solved, and efficient malignant tumor detection was achieved, improving the accuracy and comprehensiveness of the detection.
Patent Information
- Application Number
- CN202510538634.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-27
AI Technical Summary
In the prior art, the characterization method of cfDNA fragment length distribution has problems such as inconsistent definitions of long and short fragments, and the incomplete variety of fragment lengths included in the analysis, resulting in low accuracy of malignant tumor classification.
Jensen-Shannon divergence (JS divergence) was used to characterize the length distribution characteristics of cfDNA fragments. By calculating the JS divergence of the sample to be tested and the normal length distribution model, we judge whether there is a malignant tumor.
The detection rate of malignant tumors is improved, the false positive rate is low, and the AUC value reaches 0.92. It can comprehensively evaluate the changes in cfDNA length distribution and retain microscopic information on fragment length.
Smart Images

Figure BDA0005378761380000031 
Figure BDA0005378761380000051 
Figure BDA0005378761380000101
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medicine and health, and relates to a method for calculating gene markers for early disease screening and diagnosis and its application, and specifically to a method for calculating free DNA fragment length patterns and its application. Background Art
[0002] Early screening and diagnosis of malignant tumors can enable patients to receive timely treatment, thereby improving quality of life and reducing mortality. Traditional methods, such as imaging and cytological smears, are generally highly invasive and have low accuracy rates. Liquid biopsy technology, based on cell-free DNA (cfDNA), offers a non-invasive diagnostic approach and is the most widely used and promising approach for early screening of malignant tumors.
[0003] Studies have shown that cfDNA is not randomly fragmented, but rather has specific fragment end sequence preferences and fragment length distributions across the genome. The cfDNA fragment length distribution is related to the cell's nucleosome structure, which varies across different tissue types.
[0004] The cfDNA fragment length distribution of healthy individuals reflects the nucleosome pattern of white blood cells, while the degraded DNA fragments during apoptosis of the patient's diseased cells will enter the blood, thereby changing the composition of cfDNA in the blood and the proportion of DNA of different types of cells. Therefore, the fragment length distribution of cfDNA in the blood of patients with malignant tumors has changed.
[0005] Based on the difference in fragment length distribution between cfDNA of malignant samples and that of healthy people, early screening and diagnosis of malignant tumor patients can be carried out.
[0006] In the existing technology, the classification of healthy and malignant cells using the fragment length distribution of cfDNA is mainly based on the ratio between the long fragment frequency and the short fragment frequency. This method has inconsistent definitions of long and short fragments. At the same time, this method classifies the fragment lengths in the long fragment range into one category and the segments in the short fragment length range into another category, ignoring the microscopic information of the fragment length, resulting in an incomplete range of fragment length types included in the analysis.
[0007] The second method is to divide the fragment length into different intervals according to a certain length (such as 4bp or 10bp), and then analyze the fragment frequencies in different length intervals as feature values. This method only alleviates the problem of "incomplete fragment length variety" in the first method to a certain extent.
[0008] Therefore, due to the inconsistent definitions of long and short fragments in the cfDNA characterization process and the incomplete types of fragment lengths included in the analysis, there is an urgent need to provide a new cfDNA characterization method. Summary of the Invention
[0009] The present invention provides a method for evaluating the length pattern of cfDNA fragments in the genome, which efficiently detects malignant tumors by using Jensen-Shannon divergence (JS divergence) to characterize the distribution characteristics of cfDNA fragment lengths.
[0010] In a first aspect of the present invention, a method for analyzing cfDNA in a sample to be tested is provided, the method comprising the following steps:
[0011] (S1) Provides the cfDNA fragment length distribution Q of healthy samples x as a normal length distribution model;
[0012] (S2) Provide the cfDNA fragment length distribution P of the sample to be tested x ;
[0013] (S3) Through Q x and P x The difference index between the fragment length distribution of the sample to be tested and the normal length distribution model is calculated.
[0014] In another preferred embodiment, the difference index of the fragment distribution length includes: KL divergence, Wasserstein distance, Hellinger distance, total variation distance, Cosine similarity, or JS divergence.
[0015] In another preferred embodiment, the sample is a blood sample, a pleural effusion sample, a peritoneal effusion sample, or a cerebrospinal fluid sample.
[0016] In another preferred embodiment, the sample is a blood sample.
[0017] In another preferred embodiment, the difference index between the length distribution of the fragments of the sample to be tested and the normal length distribution model refers to the JS divergence between the length distribution of the fragments of the sample to be tested and the normal length distribution model.
[0018] In another preferred embodiment, the cfDNA fragment is 50-500bp in length, preferably 80-300bp, more preferably 90-230bp, and optimally 100-220bp in length.
[0019] In another preferred embodiment, the cfDNA fragment is 100-220 bp in length.
[0020] In another preferred embodiment, step (S1) further comprises the following sub-steps:
[0021] (S1-1) provides the fragment length of each cfDNA in each healthy sample among n healthy samples;
[0022] (S1-2) Calculate the number of reads supporting cfDNA of all fragment lengths in each healthy sample;
[0023] (S1-3) Calculate the frequency of the number of reads supporting cfDNA of each fragment length in each healthy sample relative to the number of reads supporting cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in each healthy sample;
[0024] (S1-4) Calculate the average frequency of cfDNA of each fragment length in the n healthy samples as the cfDNA fragment length distribution model Q of the healthy samples x .
[0025] In another preferred embodiment, step (S1) further comprises the following sub-steps:
[0026] (S1-1') provides the fragment length of each cfDNA in each healthy sample among n healthy samples;
[0027] (S1-2') Calculate the number of reads supporting cfDNA of all fragment lengths in each healthy sample;
[0028] (S1-3') Calculate the frequency of the number of reads supporting cfDNA of each fragment length in each healthy sample relative to the number of reads supporting cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in each healthy sample;
[0029] (S1-4') Calculate the median value of the frequency of cfDNA of each fragment length in the n healthy samples as the cfDNA fragment length distribution model Q of the healthy samples x .
[0030] In another preferred embodiment, n is an integer ≥10, preferably an integer ≥20, and most preferably an integer ≥30.
[0031] In another preferred embodiment, step (S2) further comprises the following sub-steps:
[0032] (S2-1) providing the fragment length of each cfDNA in the sample to be tested;
[0033] (S2-2) calculating the number of reads supporting cfDNA of all fragment lengths in the sample to be tested;
[0034] (S2-3) calculating the frequency of the number of reads supporting cfDNA of each fragment length in the sample to be tested relative to the number of reads supporting cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in the sample to be tested;
[0035] (S2-4) According to the frequency of cfDNA of each fragment length in the sample to be tested, the cfDNA fragment length distribution P of the sample to be tested is obtained. x .
[0036] In another preferred embodiment, in step (S3), the calculation refers to the calculation of the Q x 、P x And formula (I), calculate the JS divergence of the sample to be tested and the normal length distribution model,
[0037]
[0038] Among them, the M x It's P x and M x The average value of .
[0039] In another preferred embodiment, providing the fragment length of each cfDNA in the sample includes the following steps:
[0040] (C1) Perform low-depth whole-genome paired-end sequencing on the sample to obtain FastQ data;
[0041] (C2) filtering the obtained fastq, aligning it with the human genome data, and sorting it to obtain a sorted alignment file;
[0042] (C3) Deduplication of the sorted alignment file, removal of alignment information with low alignment quality, and removal of alignment information on chromosomes X and Y are performed. The positional information of the cfDNA molecules in the sample aligned to the chromosomes is extracted, and the length of each cfDNA fragment is calculated.
[0043] In another preferred embodiment, the low-quality comparison information refers to low matching accuracy or poor uniqueness of the comparison.
[0044] In another preferred embodiment, the low matching accuracy refers to a low degree of matching between the base sequence aligned to the genome and the corresponding position of the reference genome, including the presence of more mismatches, insertions or deletions.
[0045] In another preferred embodiment, the more mismatches, insertions or deletions include at least 30%, preferably at least 20%, most preferably at least 10%, and most preferably at least 5% mismatches, insertions or deletions.
[0046] In another preferred embodiment, the poor uniqueness of the alignment means that a fragment can be aligned to multiple different positions, wherein the multiple refers to ≥2.
[0047] In a second aspect of the present invention, a device for analyzing cfDNA in a sample to be tested is provided, the device comprising the following modules:
[0048] (Z1) A normal length distribution model acquisition module, wherein the normal length distribution model acquisition module is configured to: provide the cfDNA fragment length distribution Q of a healthy sample x as a normal length distribution model;
[0049] (Z3) a test sample analysis module, wherein the test sample analysis module is configured to provide a cfDNA fragment length distribution P of the test sample. x ;
[0050] (Z4) a divergence analysis module, wherein the divergence analysis module is configured to: x and P x , calculate and obtain the JS divergence between the length distribution of the sample fragment to be tested and the normal length distribution model.
[0051] In another preferred embodiment, the (Z1) module further comprises the following submodules:
[0052] (Z1-1) a fragment length acquisition module, wherein the fragment length acquisition module is configured to: provide a fragment length of each cfDNA in each healthy sample among n healthy samples;
[0053] (Z1-2) a statistical calculation module, wherein the statistical calculation module is configured to: calculate the number of reads supporting cfDNA of all fragment lengths in each healthy sample; and calculate the frequency of the number of reads supporting cfDNA of each fragment length to the number of reads supporting cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in each healthy sample;
[0054] (Z1-3) A model output module, wherein the model output module is configured to calculate the average frequency of cfDNA of each fragment length in the n healthy samples as the normal length distribution model Q x .
[0055] In another preferred embodiment, the (Z1) module further comprises the following submodules:
[0056] (Z1-1') a fragment length acquisition module, wherein the fragment length acquisition module is configured to: provide a fragment length of each cfDNA in each healthy sample among n healthy samples;
[0057] (Z1-2') a statistical calculation module, wherein the statistical calculation module is configured to: calculate the number of reads supporting cfDNA of all fragment lengths in each healthy sample; and calculate the frequency of the number of reads supporting cfDNA of each fragment length to the number of reads supporting cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in each healthy sample;
[0058] (Z1-3') A model output module, wherein the model output module is configured to calculate the median value of the frequency of cfDNA of each fragment length in the n healthy samples as the cfDNA fragment length distribution model Q of the healthy samples x .
[0059] In another preferred embodiment, n is a positive integer ≥10, preferably a positive integer ≥20, and most preferably a positive integer ≥30.
[0060] In another preferred embodiment, in module (Z4), the calculation refers to the calculation according to Q x 、P x And formula (I), calculate the JS divergence of the sample to be tested and the normal length distribution model,
[0061]
[0062] Among them, the M x It's P x and M x The average value of .
[0063] In another preferred embodiment, the (Z3) module further comprises the following submodules:
[0064] (Z3-1) a fragment length acquisition module, wherein the fragment length acquisition module is configured to: provide the fragment length of each cfDNA in the sample to be tested;
[0065] (Z3-2) a statistical calculation module, the statistical calculation module being configured to: calculate the number of reads supporting cfDNA of all fragment lengths in the sample to be tested; calculate the frequency of the number of reads supporting cfDNA of each fragment length in the sample to be tested as a percentage of the number of reads supporting cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in the sample to be tested;
[0066] (Z3-3) A fragment length output module for the sample to be tested, wherein the fragment length output module for the sample to be tested is configured to obtain the cfDNA fragment length distribution P of the sample to be tested according to the frequency of cfDNA of each fragment length in the sample to be tested. x .
[0067] In another preferred embodiment, between the module (Z1) and the module (Z3), the following components are further included:
[0068] (Z2) A threshold providing module, wherein the threshold providing module is configured to provide the JS divergence of the cfDNA fragment length distribution of the first malignant tumor sample and the cfDNA fragment length distribution of the first healthy sample with the normal length distribution model, thereby providing a threshold Y1 for malignant tumor determination.
[0069] In another preferred embodiment, the threshold providing module further includes the following submodules:
[0070] (Z2-1) providing the JS divergence of the cfDNA fragment length distribution of the first malignant tumor sample and the normal length distribution model; and providing the JS divergence of the cfDNA fragment length distribution of the first healthy sample and the normal length distribution model respectively;
[0071] (Z2-2) Calculate the JS divergence when the Youden index is maximum, thereby obtaining the threshold value Y1 for determining malignant tumors.
[0072] In another preferred example, the threshold Y1 is 0.0014.
[0073] In another preferred embodiment, after the module (Z4), it further comprises:
[0074] (Z5) A result output module, wherein the result output module is configured to compare the JS divergence Y2 of the cfDNA fragment length distribution of the sample to be tested with the normal length distribution model and the threshold value Y1 for malignant tumor judgment; if Y2 ≥ Y1, it indicates that the sample to be tested is a malignant tumor sample; if Y2 < Y1, it indicates that the sample to be tested is not a malignant tumor sample.
[0075] In another preferred embodiment, the malignant tumor includes: breast cancer, gastric cancer, colon cancer, bile duct cancer, or a combination thereof.
[0076] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features described in detail below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be listed here one by one. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 The P value between the JS divergence of the cfDNA fragment length distribution of malignant tumor samples and the normal length distribution model and the JS divergence of the cfDNA fragment length distribution of healthy human samples and the normal length distribution model is shown using the Wilcoxon test.
[0078] Figure 2 The AUC values for different cancer types in the test set are shown.
[0079] Figure 3 The AUC values for malignant tumors in the test set are shown. DETAILED DESCRIPTION
[0080] After extensive and in-depth research, the inventors discovered for the first time the importance of JS divergence in characterizing the distribution of cfDNA fragment lengths. They also found for the first time that the JS divergence between the cfDNA fragment length distributions of tumor samples and healthy samples is positively correlated with the incidence of malignant tumors. After determining the threshold for determining malignant tumors through preliminary experiments, the difference in JS divergence between the cfDNA fragment length distribution of the test sample and the normal length distribution model can be compared with the threshold to determine whether a malignant tumor exists. This is the basis for the completion of the present invention.
[0081] the term
[0082] In order to more easily understand the present disclosure, some terms are first defined. As used in this application, unless otherwise expressly provided herein, each of the following terms should have the meaning given below. Other definitions are set forth throughout the application.
[0083] As used herein, the term "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0084] As used herein, the terms "comprises," "includes," and "comprising" are used interchangeably to encompass not only closed definitions but also semi-closed and open definitions. In other words, the terms encompass "consisting of," "consisting essentially of."
[0085] Where a numerical range is provided, it is understood that every intervening integer of that value, every tenth of each intervening integer of that value, between the upper and lower limits of that range, and any other intervening values in the stated range are encompassed within the invention, unless the context clearly dictates otherwise. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any expressly excluded limits in the stated range. For example, "1 to 50" includes "2 to 25," "5 to 20," "25 to 50," "1 to 10," and the like.
[0086] As used herein, the “healthy samples” and the “first healthy samples” are not the same set of healthy samples.
[0087] cfDNA analysis
[0088] cfDNA (Cell-free DNA) refers to free DNA that exists in the extracellular environment such as human blood circulation.
[0089] From a disease diagnosis perspective, it acts as an early warning system. In the early stages of cancer, cfDNA released by tumor cells carries key information such as gene mutations and methylation abnormalities. Accurate characterization of cfDNA can detect these early signs of disease, allowing tumors to be detected in their infancy and significantly improving early diagnosis rates. For genetic disease screening, cfDNA in pregnant women's blood contains fetal genetic information, enabling accurate detection of fetal chromosomal or single-gene genetic disorders, safeguarding the health of newborns.
[0090] When it comes to predicting disease prognosis, cfDNA is a powerful predictive tool. In tumors, specific characteristics correlate with malignancy and metastasis risk, enabling precise prognostic assessment and guiding subsequent treatment strategies. In cardiovascular disease, epigenetic changes in cfDNA also illuminate risk prediction. Furthermore, in scientific research, cfDNA characterization has opened new avenues for uncovering disease mechanisms and discovering novel biomarkers, fueling ongoing medical advancements.
[0091] There are two main existing methods for classifying healthy and malignant cells using the fragment length distribution of cfDNA. The first method uses the ratio of long to short fragment frequencies for classification. However, this method has inconsistent definitions of long and short fragments. Some researchers define short fragments as those in the length range [100, 150] and long fragments as those in the length range [151, 220]; others define short fragments as those in the length range [100, 166] and long fragments as those in the length range [169, 240]. Furthermore, this method groups fragment lengths within the long fragment range into one category and those within the short fragment range into another, ignoring microscopic information about fragment lengths and resulting in an incomplete analysis of the fragment lengths. The second method divides fragment lengths into different intervals based on a certain length (e.g., 4 bp or 10 bp) and then uses the fragment frequencies in these intervals as feature values for analysis. This method only partially alleviates the incompleteness of the fragment lengths in the first method.
[0092] Method of the present invention
[0093] The method of the present invention uses JS divergence to characterize the length distribution of cfDNA and to determine the presence of a malignant tumor based on the JS divergence. The method of the present invention is simple to operate, low-cost, and readily scalable. Furthermore, the false positive rate in detecting malignant tumors is low, and all malignant tumors can be detected.
[0094] Specifically, in the present invention, in order to solve the problems of inconsistent definitions of long and short fragments and incomplete types of fragment lengths included in the analysis, a method for evaluating the length pattern of cfDNA fragments in the genome was developed, namely, using Jensen-Shannon divergence (JS divergence) to characterize the distribution characteristics of cfDNA fragment lengths.
[0095] The JS divergence is a symmetric measure of two probability distributions, ranging from [0 to 1]. Smaller JS divergence values indicate smaller differences between the two distributions, and vice versa. The JS divergence is used to assess the differences in cfDNA fragment length distributions between test samples and healthy individuals. A smaller JS divergence indicates a smaller difference in fragment length distribution between the test sample and healthy individuals, and a lower likelihood of malignancy. Conversely, a higher likelihood of malignancy is observed in the test sample.
[0096] In another preferred embodiment, the cfDNA in the sample to be tested is an arbitrary random sequence with a length of 50-500bp. In another preferred embodiment, the cfDNA in the sample to be tested is an arbitrary random sequence with a length of 60-400bp. In another preferred embodiment, the cfDNA in the sample to be tested is an arbitrary random sequence with a length of 80-300bp. In another preferred embodiment, the cfDNA in the sample to be tested is an arbitrary random sequence with a length of 90-230bp. In another preferred embodiment, the cfDNA in the sample to be tested is an arbitrary random sequence with a length of 100-220bp. In another preferred embodiment, the cfDNA in the sample to be tested is an arbitrary random sequence with a length of 150-200bp.
[0097] The main advantages of the present invention include:
[0098] (a) This method uses JS divergence as a characteristic value to characterize differences in cfDNA length distribution. It does not define long and short fragment intervals, but incorporates the length of each fragment into the analysis. This allows for a comprehensive assessment of changes in cfDNA length distribution, retaining microscopic information about fragment lengths and achieving greater universality.
[0099] (b) The method of the present invention was used to detect malignant tumors, and the AUC was as high as 92%.
[0100] The present invention will be further described below in conjunction with specific examples. It should be understood that these examples are intended to illustrate the present invention only and are not intended to limit the scope of the invention. The experimental methods in the following examples, for which specific conditions are not specified, are generally based on conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and parts by weight.
[0101] Example 1: Constructing a normal length distribution model.
[0102] Several healthy human plasma cfDNA samples were selected for library construction, sequencing, and analysis to construct a normal length distribution model. The specific steps are as follows:
[0103] (N1) First, 30 healthy human plasma samples were selected for low-depth whole-genome paired-end sequencing. The obtained Fastq was filtered and aligned to the human genome using the software BWA. The alignment files were sorted using the software samtools to obtain the sorted alignment files. The sorted alignment files were deduplicated using the software Picard to obtain the deduplicated alignment files. Low-quality alignments were removed, as were alignments to chromosomes X and Y. The positional information of the cfDNA molecules aligned to the chromosomes was extracted, and the length of each cfDNA fragment was calculated.
[0104] (N2) Calculate the number of reads supporting each fragment length in each sample;
[0105] (N3) Count the number of reads supporting the fragments with a length range of [100, 220] in each sample;
[0106] (N4) Counting the ratio of each fragment length in the range of [100, 220] to the number of reads supporting the length described in step (N3) in each sample to obtain the frequency of each fragment length;
[0107] (N5) Calculate the mean frequency of each fragment length in the interval [100, 220] in all samples, and use it as a model of the fragment length distribution of healthy people, denoted by Q x , where x is the segment length, Q x is the average frequency of fragments of length x in healthy individuals.
[0108] Example 2: The method of the present invention evaluates the differences between healthy human samples and malignant tumor samples.
[0109] Another 58 healthy human cfDNA samples were selected and the JS divergence between them and the normal length distribution model obtained in Example 1 was calculated. At the same time, 52 malignant tumor cfDNA samples (13 each of breast cancer, gastric cancer, colorectal cancer, and bile duct cancer) were selected and the JS divergence between them and the normal length distribution model obtained in Example 1 was calculated. Then, the difference between the two was calculated. The specific steps are as follows:
[0110] (M1) Perform low-depth whole-genome paired-end sequencing on cfDNA samples. The resulting Fastq data is filtered and aligned to the human genome using BWA software. The alignment files are sorted using samtools software to obtain a sorted alignment file. The sorted alignment file is deduplicated using Picard software to obtain a deduplicated alignment file. Alignments with low alignment quality are removed, as are alignments to chromosomes X and Y. The positional information of the cfDNA molecules aligned to the chromosomes is extracted, and the length of each cfDNA fragment is calculated.
[0111] (M2) Calculate the number of reads supporting each fragment length in each sample;
[0112] (M3) Count the number of reads supporting the fragments with a length range of [100, 220] in each sample;
[0113] (M4) Counting the ratio of each fragment length in the range of [100, 220] to the number of reads supporting the total number of reads described in (M3) in each sample to obtain the frequency of each fragment length;
[0114] (M5) Calculate the length distribution P of each fragment described in (M4) in each sample x , where x is the length of the segment, P x is the frequency of segments of length x;
[0115] (M6) Calculate the length distribution P of each fragment in each sample described in (M5) x The model Q of each fragment length distribution of healthy people described in Example 1 x The average value between x ;
[0116] (M7) calculating the JS divergence between the length distribution of each sample fragment and the normal length distribution model in Example 1 according to formula (I);
[0117]
[0118] (M8) The Wilcoxon test was used to statistically analyze the P value between the JS divergence of the cfDNA of the malignant tumor sample and the normal length distribution model in Example 1, and the JS divergence of the cfDNA of the healthy human sample and the normal length distribution model in Example 1.
[0119] The results are as follows Figure 1 As shown, the pvalue is 1.58e-09, there is a significant difference between the two, and the JS divergence between the cfDNA fragment length distribution of malignant tumor samples and the normal length distribution model is greater than the JS divergence between the cfDNA fragment length distribution of healthy human samples and the normal length distribution model.
[0120] Example 3: Evaluate the performance of JS divergence in predicting malignant and healthy individuals in different cancer types.
[0121] The 58 healthy subjects in Example 2 were selected and combined with 12 breast cancer cases, 12 gastric cancer cases, 12 colorectal cancer cases and 12 bile duct cancer cases to form a breast cancer training set, a gastric cancer training set, a colorectal cancer training set and a bile duct cancer training set respectively.
[0122] In addition, 57 healthy subjects were selected to form the breast cancer test set, gastric cancer test set, colorectal cancer test set and bile duct cancer test set together with 12 breast cancer cases, 12 gastric cancer cases, 12 colorectal cancer cases and 12 bile duct cancer cases.
[0123] The performance of using JS divergence to predict malignant tumors and healthy subjects in different cancer types was evaluated. The specific steps are as follows:
[0124] (Q1) Following steps (M1) to (M7) in Example 2, obtain the JS divergence between the cfDNA fragment length distribution of all samples and the normal length distribution model;
[0125] (Q2) Calculate the JS divergence when the Youden index is maximized based on the JS divergence between the healthy human samples and the corresponding cancer samples in the breast cancer training set, the gastric cancer training set, the colon cancer training set, and the bile duct cancer training set and the normal length distribution model described in Example 1.
[0126] (Q3) Define the JS divergence value when the Youden index is maximum in the breast cancer training set, gastric cancer training set, colon cancer training set, and bile duct cancer training set as the threshold of the corresponding cancer type, and calculate the sensitivity and specificity of JS divergence in the corresponding cancer type test set based on the threshold.
[0127] The results are shown in Table 1 and Figure 2 Table 1 shows the performance of JS divergence in predicting malignant and healthy people in different cancer types. Figure 2 is the AUC value of different cancer types in the test set.
[0128] Table 1. Performance of JS divergence in predicting malignant and healthy people in different cancer types
[0129]
[0130] Example 4: Evaluating the performance of JS divergence in predicting malignant and healthy subjects in pan-cancer.
[0131] The 58 healthy and 52 malignant plasma samples in Example 2 were selected as the training set. In addition, 57 healthy individuals and 48 malignant plasma samples (12 each of breast cancer, gastric cancer, colorectal cancer, and bile duct cancer) were selected as the test set to evaluate the performance of using JS divergence to predict malignant tumors and health. The specific steps are as follows:
[0132] (Q1′) Following steps (M1) to (M7) in Example 2, obtain the JS divergence between the cfDNA fragment length distribution of all samples and the normal length distribution model;
[0133] (Q2') Based on the JS divergence between the cfDNA fragment length distribution of 58 healthy and 52 malignant tumor samples in the training set and the normal length distribution model, the JS divergence when the Youden index value is maximized is calculated;
[0134] (Q3') The JS divergence value when the Youden index is the largest is defined as the threshold for judging whether a sample is malignant or healthy, and the sensitivity and specificity of the JS divergence in the test set are calculated based on the threshold.
[0135] The results show that when the Youden index is the largest in the training set, the JS divergence value is 0.0014. Taking 0.0014 as the threshold for judging whether a sample is malignant or healthy, the JS divergence in the test set has a specificity of 0.91 and a sensitivity of 0.88. The AUC of JS divergence in predicting health and malignancy in the test set is 0.92. Figure 3 shown.
[0136] Example 5: Using JS divergence to predict randomly selected samples.
[0137] Ten samples were randomly selected from the 105 test samples in Example 4 as test samples, and the JS divergence values of their length distributions were calculated. A threshold of 0.0014 was used to determine whether the sample was healthy or malignant. When the JS divergence of the length distribution of the test sample was greater than or equal to 0.0014, the test sample was determined to be malignant; when the JS divergence of the length distribution of the test sample was less than 0.0014, the test sample was determined to be healthy. The specific steps are as follows:
[0138] According to steps (M1) to (M7) in Example 2, the JS divergence between the fragment length distributions of the 10 test samples and the fragment length distribution model of the healthy subjects in Example 1 is obtained;
[0139] Taking 0.0014 as the threshold for judging whether a sample is malignant or healthy, the JS divergence of the length distribution of the 10 test samples is compared with the threshold of 0.0014. When the JS divergence of the length distribution of the test sample is ≥0.0014, the test sample is judged to be malignant; when the JS divergence of the test sample is <0.0014, the test sample is judged to be healthy.
[0140] The results are shown in Table 2. Among the 10 randomly selected samples, there were 2 cases of bile duct cancer, 2 cases of gastric cancer, 1 case of breast cancer, and 5 samples from healthy individuals. Of these, 2 cases of bile duct cancer, 2 cases of gastric cancer, and 1 case of breast cancer were all judged as malignant. Of the 5 healthy individuals, 4 were judged as healthy and 1 was judged as malignant. The accuracy rate was 90%.
[0141] Among them, all subjects with malignant tumors were detected, and the malignant tumor detection rate was 100%.
[0142] Table 2 Judgment results of random samples
[0143]
[0144] All documents mentioned in this application are incorporated herein by reference, just as if each document were incorporated herein by reference individually. It should also be understood that after reading the above teachings of the present invention, those skilled in the art may make various changes or modifications to the present invention, and that such equivalents also fall within the scope of the claims appended hereto.
Claims
1. A method for analyzing cfDNA in a sample to be tested, characterized in that: The method comprises the following steps: (S1) Provides the cfDNA fragment length distribution Q of healthy samples x as a normal length distribution model; (S2) Provide the cfDNA fragment length distribution P of the sample to be tested x ; (S3) Through Q x and P x The difference index between the fragment length distribution of the sample to be tested and the normal length distribution model is calculated.
2. The method according to claim 1, wherein The difference index between the length distribution of the fragments of the sample to be tested and the normal length distribution model refers to the JS divergence between the length distribution of the fragments of the sample to be tested and the normal length distribution model.
3. The method according to claim 1, wherein The cfDNA fragments are 50-500bp in length, preferably 80-300bp, more preferably 90-230bp, and most preferably 100-220bp in length.
4. The method according to claim 1, wherein Step (S1) further comprises the following sub-steps: (S1-1) provides the fragment length of each cfDNA in each healthy sample among n healthy samples; (S1-2) Calculate the number of reads supporting cfDNA of all fragment lengths in each healthy sample; (S1-3) Calculate the frequency of the number of reads supporting cfDNA of each fragment length in each healthy sample relative to the number of reads supporting cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in each healthy sample; (S1-4) Calculate the average frequency of cfDNA of each fragment length in the n healthy samples as the cfDNA fragment length distribution model Q of the healthy samples x .
5. The method according to claim 1, wherein Step (S2) further comprises the following sub-steps: (S2-1) providing the fragment length of each cfDNA in the sample to be tested; (S2-2) calculating the number of reads supporting cfDNA of all fragment lengths in the sample to be tested; (S2-3) calculating the frequency of the number of reads supporting cfDNA of each fragment length in the sample to be tested relative to the number of reads supporting cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in the sample to be tested; (S2-4) According to the frequency of cfDNA of each fragment length in the sample to be tested, the cfDNA fragment length distribution P of the sample to be tested is obtained. x .
6. A device for analyzing cfDNA in a sample to be tested, characterized in that: The device comprises the following modules: (Z1) A normal length distribution model acquisition module, wherein the normal length distribution model acquisition module is configured to: provide the cfDNA fragment length distribution Q of a healthy sample x as a normal length distribution model; (Z3) a test sample analysis module, wherein the test sample analysis module is configured to provide a cfDNA fragment length distribution P of the test sample. x ; (Z4) a divergence analysis module, wherein the divergence analysis module is configured to: x and P x , calculate and obtain the JS divergence between the length distribution of the sample fragment to be tested and the normal length distribution model.
7. The analysis device according to claim 6, wherein Between the (Z1) module and the (Z3) module, it also includes: (Z2) A threshold providing module, wherein the threshold providing module is configured to provide the JS divergence of the cfDNA fragment length distribution of the first malignant tumor sample and the cfDNA fragment length distribution of the first healthy sample with the normal length distribution model, thereby providing a threshold Y1 for malignant tumor determination.
8. The analysis device according to claim 7, wherein The threshold providing module further includes the following submodules: (Z2-1) providing the JS divergence of the cfDNA fragment length distribution of the first malignant tumor sample and the normal length distribution model; and providing the JS divergence of the cfDNA fragment length distribution of the first healthy sample and the normal length distribution model respectively; (Z2-2) Calculate the JS divergence when the Youden index is maximum, thereby obtaining the threshold value Y1 for determining malignant tumors.
9. The analysis device according to claim 6, wherein After the (Z4) module, it also contains: (Z5) A result output module, wherein the result output module is configured to compare the JS divergence Y2 of the cfDNA fragment length distribution of the sample to be tested with the normal length distribution model and the threshold value Y1 for malignant tumor judgment; if Y2 ≥ Y1, it indicates that the sample to be tested is a malignant tumor sample; if Y2 < Y1, it indicates that the sample to be tested is not a malignant tumor sample.
10. The analysis device according to claim 7, wherein The malignant tumors include breast cancer, gastric cancer, colon cancer, bile duct cancer, or a combination thereof.
Citation Information
Patent Citations
Using cell-free DNA fragment size to detect tumor-associated variant
CN110800063A
Microsatellite instability analysis method and analysis device
CN112037859A
Cancer non-invasive early screening method and system based on cfDNA fragment length distribution characteristics
CN117316278A
Using cell-free DNA fragment size to detect tumor-associated variant
US20180307796A1