Methods and applications for calculating cell-free DNA fragment length patterns

CN120452539BActive Publication Date: 2026-08-14OMIXSCIENCE (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0008]因此,由于cfDNA表征过程中存在长片段和短片段定义不统一,纳入分析的片段长度种类不全面的问题,迫切需要提供一种新的cfDNA表征方法

Benefits of technology

[0097]本发明的主要优点包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure BDA0005378761380000031
    Figure BDA0005378761380000031
Patent Text Reader

Abstract

This invention provides a method for analyzing cfDNA, which efficiently detects malignant tumors by using JS divergence to characterize the cfDNA fragment length distribution characteristics. Specifically, firstly, a normal sample cfDNA fragment length distribution model is constructed. Then, through preliminary experiments, the JS divergence between the cfDNA fragment length distribution of malignant tumor samples and the cfDNA fragment length distribution of healthy human samples and the normal length distribution model is calculated. After determining the threshold for malignant tumor detection, the JS divergence between the cfDNA fragment length distribution of the sample to be tested and the normal length distribution model is calculated and compared with the threshold to determine whether a malignant tumor is present. This method has high accuracy, with an AUC value reaching 0.92.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medicine and health, and relates to a method for calculating gene markers for early disease screening and diagnosis, and its application, specifically involving a method and application for calculating cell-free DNA fragment length patterns. Background Technology

[0002] Early screening and diagnosis of malignant tumors can enable patients to receive timely treatment, thereby improving their quality of life and reducing mortality. Traditional methods relying on imaging examinations and cell smear techniques are generally highly invasive and have low accuracy. In contrast, liquid biopsy technology based on cell-free DNA (cfDNA) provides patients with a non-invasive diagnostic approach and is the most widely used and promising method for early screening of malignant tumors.

[0003] Studies have shown that cfDNA is not randomly fragmented, but exhibits specific terminal sequence preferences and fragment length distributions across different regions of the genome. The cfDNA fragment length distribution is related to the nucleosome structure of cells, which differs across different tissue types.

[0004] The distribution of cfDNA fragment lengths in healthy individuals reflects the nucleosome pattern of leukocytes. However, DNA fragments degraded during apoptosis of diseased cells in patients enter the bloodstream, thereby altering the composition of cfDNA in the blood and the proportion of DNA in different cell types. Therefore, the distribution of cfDNA fragment lengths in the blood of patients with malignant tumors is altered.

[0005] Based on the difference in fragment length distribution between malignant cfDNA samples and healthy individuals, early screening and diagnosis of malignant tumor patients can be achieved.

[0006] In existing technologies, the classification of healthy and malignant cells based on the fragment length distribution of cfDNA is mainly based on the ratio between the frequencies of long fragments and short fragments. However, this method does not have a unified definition of long and short fragments. Furthermore, it groups fragments within long fragment ranges into one category and fragments within short fragment ranges into another, ignoring microscopic information about fragment lengths and resulting in an incomplete range of fragment length types included in the analysis.

[0007] The second method divides the fragment length into different intervals according to a certain length (e.g., 4bp or 10bp), and then uses the fragment frequency of different length intervals as feature values ​​for analysis. This method only alleviates the problem of "incomplete fragment length types" in the first method to a certain extent.

[0008] Therefore, due to the inconsistent definitions of long and short fragments and the incomplete range of fragment lengths included in the analysis during cfDNA characterization, there is an urgent need to provide a new cfDNA characterization method. Summary of the Invention

[0009] This invention provides a method for assessing cfDNA fragment length patterns in the genome, which efficiently detects malignant tumors by characterizing cfDNA fragment length distribution features using Jensen-Shannon divergence (JS divergence).

[0010] In a first aspect of the present invention, a method for analyzing cfDNA in a sample to be tested is provided, the method comprising the following steps:

[0011] (S1) Provides the cfDNA fragment length distribution of healthy samples Q x As a normal length distribution model;

[0012] (S2) Provides the cfDNA fragment length distribution of the sample to be tested. x ;

[0013] (S3) via Q x and P x The difference index between the length distribution of the sample fragments to be tested and the normal length distribution model was calculated.

[0014] In another preferred embodiment, the difference index of the fragment distribution length includes: KL divergence, Wasserstein distance, Hellinger distance, Total variation distance, Cosine similarity, or JS divergence.

[0015] In another preferred embodiment, the sample is a blood sample, a pleural effusion sample, a peritoneal effusion sample, or a cerebrospinal fluid sample.

[0016] In another preferred embodiment, the sample is a blood sample.

[0017] In another preferred embodiment, the difference index between the length distribution of the sample fragment to be tested and the normal length distribution model refers to the JS divergence between the length distribution of the sample fragment to be tested and the normal length distribution model.

[0018] In another preferred embodiment, the cfDNA fragment is 50-500 bp in length, more preferably 80-300 bp, even more preferably 90-230 bp, and most preferably 100-220 bp.

[0019] In another preferred embodiment, the cfDNA fragment is 100-220 bp in length.

[0020] In another preferred embodiment, step (S1) further includes the following sub-steps:

[0021] (S1-1) provides the fragment length of each cfDNA in each of the n healthy samples;

[0022] (S1-2) Calculate the number of reads supporting all fragment lengths of cfDNA in each healthy sample;

[0023] (S1-3) Calculate the frequency of the number of reads supported by each fragment length of cfDNA in each healthy sample relative to the number of reads supported by all fragment lengths of cfDNA, thereby obtaining the frequency of cfDNA of each fragment length in each healthy sample.

[0024] (S1-4) Calculate the average frequency of cfDNA for each fragment length in the n healthy samples, and use it as the cfDNA fragment length distribution model Q for the healthy samples. x .

[0025] In another preferred embodiment, step (S1) further includes the following sub-steps:

[0026] (S1-1') provides the fragment length of each cfDNA in each of the n healthy samples;

[0027] (S1-2') Calculate the number of reads supporting all fragment lengths of cfDNA in each healthy sample;

[0028] (S1-3') Calculate the frequency of the number of reads supported by each fragment length of cfDNA in each healthy sample relative to the number of reads supported by all fragment lengths of cfDNA, thereby obtaining the frequency of cfDNA of each fragment length in each healthy sample.

[0029] (S1-4') Calculate the median frequency of cfDNA for each fragment length in the n healthy samples, and use it as the cfDNA fragment length distribution model Q for the healthy samples. x .

[0030] In another preferred embodiment, n is an integer ≥10, more preferably an integer ≥20, and most preferably an integer ≥30.

[0031] In another preferred embodiment, step (S2) further includes the following sub-steps:

[0032] (S2-1) Provides the fragment length of each cfDNA in the sample to be tested;

[0033] (S2-2) Calculate the number of reads supporting all fragment lengths of cfDNA in the sample to be tested;

[0034] (S2-3) Calculate the frequency of the number of reads supported by each fragment length of cfDNA in the test sample relative to the number of reads supported by all fragment lengths of cfDNA, thereby obtaining the frequency of cfDNA of each fragment length in the test sample.

[0035] (S2-4) Based on the frequency of cfDNA fragments of each fragment length in the sample to be tested, the cfDNA fragment length distribution P of the sample to be tested is obtained. x .

[0036] In another preferred embodiment, in step (S3), the calculation refers to the calculation based on Q. x P x Formula (I) is used to calculate the JS divergence between the test sample and the normal length distribution model.

[0037]

[0038] Wherein, the M x It is P x and M x The average value.

[0039] In another preferred embodiment, providing the fragment length of each cfDNA in the sample includes the following steps:

[0040] (C1) Low-depth whole-genome paired-end sequencing data were obtained from the sample to obtain fastq;

[0041] (C2) After filtering the obtained fastq data, it is compared with the human genome data and sorted to obtain the sorted alignment file;

[0042] (C3) After sorting the alignment files, remove duplicates, remove alignment information with low alignment quality, remove alignment information that is aligned to the X chromosome and Y chromosome, extract the position information of cfDNA molecules in the sample that are aligned to the chromosome, and calculate the fragment length of each cfDNA.

[0043] In another preferred embodiment, the low-quality alignment information refers to low matching accuracy or poor alignment uniqueness.

[0044] In another preferred embodiment, the low accuracy of the match refers to a low degree of matching between the base sequence aligned to the genome and the corresponding position in the reference genome, including the presence of numerous mismatches, insertions, or deletions.

[0045] In another preferred embodiment, the greater number of mismatches, insertions or omissions includes at least 30%, more preferably at least 20%, most preferably at least 10%, and most preferably at least 5%.

[0046] In another preferred embodiment, the uniqueness difference of the alignment means that a segment can be aligned to multiple different positions, wherein "multiple" means ≥2.

[0047] In a second aspect of the invention, an analytical apparatus for cfDNA in a sample to be tested is provided, the apparatus comprising the following modules:

[0048] (Z1) Normal length distribution model acquisition module, which is configured to provide the cfDNA fragment length distribution Q of healthy samples. x As a normal length distribution model;

[0049] (Z3) Sample analysis module, which is configured to provide the cfDNA fragment length distribution P of the sample to be tested. x ;

[0050] (Z4) Divergence analysis module, which is configured to: use Q x and P x The JS divergence between the length distribution of the sample fragments to be tested and the normal length distribution model is calculated.

[0051] In another preferred embodiment, the (Z1) module further comprises the following sub-modules:

[0052] (Z1-1) Fragment length acquisition module, the fragment length acquisition module is configured to: provide the fragment length of each cfDNA in each of n healthy samples;

[0053] (Z1-2) Statistical calculation module, which is configured to: calculate the number of reads supported by cfDNA of all fragment lengths in each healthy sample; and calculate the frequency of the number of reads supported by cfDNA of each fragment length relative to the number of reads supported by cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in each healthy sample.

[0054] (Z1-3) Model output module, configured to: calculate the average frequency of cfDNA for each fragment length in the n healthy samples, as the normal length distribution model Q. x .

[0055] In another preferred embodiment, the (Z1) module further comprises the following sub-modules:

[0056] (Z1-1') Fragment length acquisition module, the fragment length acquisition module is configured to: provide the fragment length of each cfDNA in each of n healthy samples;

[0057] (Z1-2') Statistical calculation module, which is configured to: calculate the number of reads supported by cfDNA of all fragment lengths in each healthy sample; and calculate the frequency of the number of reads supported by cfDNA of each fragment length relative to the number of reads supported by cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in each healthy sample.

[0058] (Z1-3') Model output module, configured to: calculate the median frequency of cfDNA for each fragment length in the n healthy samples, as the cfDNA fragment length distribution model Q of the healthy samples. x .

[0059] In another preferred embodiment, n is a positive integer ≥10, more preferably a positive integer ≥20, and most preferably a positive integer ≥30.

[0060] In another preferred embodiment, in module (Z4), the calculation refers to the calculation based on Q. x P x Formula (I) is used to calculate the JS divergence between the test sample and the normal length distribution model.

[0061]

[0062] Wherein, the M x It is P x and M x The average value.

[0063] In another preferred embodiment, the (Z3) module further includes the following sub-modules:

[0064] (Z3-1) Fragment length acquisition module, the fragment length acquisition module is configured to: provide the fragment length of each cfDNA in the sample to be tested;

[0065] (Z3-2) Statistical calculation module, the statistical calculation module is configured to: calculate the number of reads supported by cfDNA of all fragment lengths in the sample to be tested; calculate the frequency of the number of reads supported by cfDNA of each fragment length in the sample to be tested relative to the number of reads supported by cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in the sample to be tested.

[0066] (Z3-3) Sample fragment length output module, configured to: obtain the cfDNA fragment length distribution P of the sample based on the frequency of each fragment length in the sample. x .

[0067] In another preferred embodiment, between module (Z1) and module (Z3), the following is also included:

[0068] (Z2) Threshold providing module, the threshold providing module is configured to: provide the JS divergence between the cfDNA fragment length distribution of the first malignant tumor sample and the cfDNA fragment length distribution of the first healthy sample and the normal length distribution model, thereby providing the threshold Y1 for malignant tumor determination.

[0069] In another preferred embodiment, the threshold providing module further includes the following sub-modules:

[0070] (Z2-1) Provides the JS divergence between the cfDNA fragment length distribution of the first malignant tumor sample and the normal length distribution model; and provides the JS divergence between the cfDNA fragment length distribution of the first healthy sample and the normal length distribution model respectively;

[0071] (Z2-2) Calculate the JS divergence when the Youden index is at its maximum, and thus obtain the threshold Y1 for malignant tumor determination.

[0072] In another preferred embodiment, the threshold Y1 is 0.0014.

[0073] In another preferred embodiment, following module (Z4) is also included:

[0074] (Z5) Result output module, the result output module is configured to: compare the JS divergence Y2 of the cfDNA fragment length distribution of the test sample with the normal length distribution model with the threshold Y1 for malignant tumor determination; if Y2≥Y1, it indicates that the test sample is a malignant tumor sample; if Y2<Y1, it indicates that the test sample is not a malignant tumor sample.

[0075] In another preferred embodiment, the malignant tumor includes: breast cancer, gastric cancer, colon cancer, bile duct cancer, or a combination thereof.

[0076] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description

[0077] Figure 1 The p-values ​​of the JS divergence between the distribution of cfDNA fragment lengths in malignant tumor samples and the normal length distribution model, as well as the JS divergence between the distribution of cfDNA fragment lengths in healthy human samples and the normal length distribution model, are shown using the Wilcoxon test.

[0078] Figure 2 The AUC values ​​for different cancer types in the test set are shown.

[0079] Figure 3 The AUC values ​​for malignant tumors in the test set are shown. Detailed Implementation

[0080] Through extensive and in-depth research, the inventors have, for the first time, discovered the significant role of JS divergence in characterizing the distribution of cfDNA fragment lengths. They have also discovered a positive correlation between the JS divergence between the cfDNA fragment length distributions of tumor samples and those of healthy samples and the incidence of malignant tumors. After determining the threshold for malignant tumor identification through preliminary experiments, the presence of malignant tumors can be determined by comparing the difference in JS divergence between the cfDNA fragment length distribution of the test sample and a normal length distribution model with the threshold. This invention is based on these findings.

[0081] the term

[0082] To facilitate a clearer understanding of this disclosure, certain terms are first defined. As used herein, unless otherwise expressly specified herein, each of the following terms shall have the meaning given below. Other definitions are set forth throughout the application.

[0083] As used herein, the term “and / or” refers to and covers any and all possible combinations of one or more of the related listed items.

[0084] As used herein, the terms “comprising,” “including,” and “containing” are used interchangeably and include not only closed definitions but also semi-closed and open definitions. In other words, the terms include “consisting of” and “substantially consisting of”.

[0085] Where a numerical range is provided, unless the context clearly indicates otherwise, it should be understood that every intermediate integer of the value 20, every tenth of every intermediate integer of the value, any other intermediate value between the upper and lower limits of the range, and any other intermediate value within the specified range are included in this invention. The upper and lower limits of these smaller ranges may be independently included within the smaller range and also covered by this invention, but are subject to any express exclusions within the specified range. For example, "1 to 50" includes "2 to 25", "5 to 20", "25 to 50", "1 to 10", etc.

[0086] As used herein, the “health sample” and the “first health sample” are not the same set of health samples.

[0087] cfDNA analysis

[0088] cfDNA (Cell-free DNA) refers to free DNA that exists in the extracellular environment, such as the bloodstream in the human body.

[0089] From a disease diagnosis perspective, it acts as an early warning system. In the early stages of tumors, the cfDNA released by tumor cells carries crucial information such as gene mutations and abnormal methylation. Precise characterization of this DNA allows for the keen detection of these early signs of disease, making tumors impossible to miss in their nascent stages and significantly improving early diagnosis rates. For genetic disease screening, the cfDNA in a pregnant woman's blood contains fetal genetic information, helping to accurately detect chromosomal or single-gene genetic diseases in the fetus, thus safeguarding the health of the new life.

[0090] When it comes to disease prognosis, cfDNA is a powerful predictive tool. In tumors, specific characteristics are associated with malignancy and metastasis risk, allowing for precise prognosis assessment and guiding subsequent treatment strategies. In cardiovascular diseases, epigenetic changes in cfDNA also illuminate risk prediction. Furthermore, in scientific research, cfDNA characterization has opened new pathways for unraveling disease mechanisms and discovering novel biomarkers, continuously injecting momentum into medical advancements.

[0091] Existing technologies primarily employ two methods to classify healthy and malignant tumors using the fragment length distribution of cfDNA. The first method classifies fragments based on the ratio of long to short fragment frequencies. However, this method lacks a consistent definition of long and short fragments. Some researchers define short fragments as those in the range [100, 150] and long fragments as those in the range [151, 220]; others define short fragments as those in the range [100, 166] and long fragments as those in the range [169, 240]. Furthermore, this method groups fragments within long and short length ranges together, ignoring microscopic information about fragment length and resulting in an incomplete range of fragment length types included in the analysis. The second method divides fragment lengths into different intervals based on a certain length (e.g., 4 bp or 10 bp) and then uses the frequency of fragments in different length intervals as characteristic values ​​for analysis. This method only partially alleviates the problem of incomplete fragment length classification in the first method.

[0092] The method of the present invention

[0093] The method of this invention refers to using JS divergence to characterize the distribution of cfDNA length and to determine the presence of malignant tumors based on JS divergence. This method is simple to operate, low in cost, and easy to promote; moreover, it has a low probability of false positives in the detection of malignant tumors and can detect all malignant tumors.

[0094] Specifically, in order to address the problems of inconsistent definitions of long and short fragments and incomplete inclusion of fragment length types in the analysis, this invention develops a method for evaluating cfDNA fragment length patterns in the genome, namely, using Jensen-Shannon divergence (JS divergence) to characterize the cfDNA fragment length distribution features.

[0095] The JS divergence is a method for measuring the difference between two probability distributions. It is symmetrical and ranges from [0,1]. The smaller the JS divergence value, the smaller the difference between the two distributions, and vice versa. The JS divergence is used to assess the difference in cfDNA fragment length distribution between the test sample and healthy individuals. A smaller JS divergence indicates a smaller difference in fragment length distribution between the test sample and healthy individuals, suggesting a lower probability of malignancy; conversely, a larger JS divergence indicates a higher probability of malignancy in the test sample.

[0096] In another preferred embodiment, the cfDNA in the test sample is an arbitrary random sequence of 50-500 bp in length. In another preferred embodiment, the cfDNA in the test sample is an arbitrary random sequence of 60-400 bp in length. In another preferred embodiment, the cfDNA in the test sample is an arbitrary random sequence of 80-300 bp in length. In another preferred embodiment, the cfDNA in the test sample is an arbitrary random sequence of 90-230 bp in length. In another preferred embodiment, the cfDNA in the test sample is an arbitrary random sequence of 100-220 bp in length. In another preferred embodiment, the cfDNA in the test sample is an arbitrary random sequence of 150-200 bp in length.

[0097] The main advantages of this invention include:

[0098] (a) This invention uses JS divergence as a characteristic value to characterize the differences in cfDNA length distribution, without defining long and short fragment intervals, and includes the length of each fragment in the analysis. This allows for a comprehensive assessment of changes in cfDNA length distribution, preserves microscopic information about fragment length, and is more universal.

[0099] (b) The method of the present invention is used in the detection of malignant tumors with an AUC of up to 92%.

[0100] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments, unless otherwise specified, are generally performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and parts by weight.

[0101] Example 1: Constructing a normal length distribution model.

[0102] Several plasma cfDNA samples from healthy individuals were selected for library construction, sequencing, and analysis to construct a normal length distribution model. The specific steps are as follows:

[0103] (N1) First, low-depth whole-genome paired-end sequencing data were obtained from plasma samples of 30 healthy individuals. The obtained FastQ data was filtered and aligned with the human genome using BWA software. The alignment files were then sorted using Samtools software, resulting in sorted alignment files. Picard software was used to remove duplicates from the sorted alignment files, obtaining deduplicated alignment files. Alignment information with low quality was removed, as were alignments to the X and Y chromosomes. The positional information of cfDNA molecules aligned to the chromosomes was extracted, and the fragment length of each cfDNA was calculated.

[0104] (N2) Calculate the number of reads supported for each segment length in each sample;

[0105] (N3) Count the number of reads supported by all segments in each sample whose segment length ranges from [100, 220];

[0106] (N4) Calculate the ratio of the length of each segment in the range of [100, 220] in each sample to the number of read segments supported by all the read segments described in step (N3) to obtain the frequency of each segment length;

[0107] (N5) Calculate the mean frequency of each segment length in the range [100, 220] across all samples. Use this as a model for the segment length distribution in healthy individuals, denoted as Q. x Where x is the fragment length, Q x This represents the average frequency of segments of length x in healthy individuals.

[0108] Example 2: The method of the present invention assesses the differences between healthy human samples and malignant tumor samples.

[0109] Another 58 healthy individuals' cfDNA samples were selected, and their JS divergence with the normal length distribution model obtained in Example 1 was calculated. Simultaneously, 52 malignant tumor cfDNA samples (13 each of breast cancer, gastric cancer, colorectal cancer, and bile duct cancer) were selected, and their JS divergence with the normal length distribution model obtained in Example 1 was calculated. Then, the difference between the two was calculated. The specific steps are as follows:

[0110] (M1) Low-depth whole-genome paired-end sequencing data of cfDNA samples were obtained. After filtering, the obtained FastQ data was aligned with the human genome using BWA software. The alignment files were then sorted using Samtools software to obtain sorted alignment files. The sorted alignment files were then deduplicated using Picard software to obtain deduplicated alignment files, removing low-quality alignments and those aligned to the X and Y chromosomes. The chromosomal position information of cfDNA molecules was extracted, and the fragment length of each cfDNA was calculated.

[0111] (M2) Calculate the number of reads supported for each segment length in each sample;

[0112] (M3) Count the number of reads supported by each segment whose segment length ranges from [100, 220] in each sample;

[0113] (M4) Calculate the ratio of the length of each segment in the range [100, 220] in each sample to the number of reads supported by all the segments mentioned in (M3), and obtain the frequency of each segment length;

[0114] (M5) Calculate the length distribution P of each segment described in (M4) for each sample. x Where x is the segment length, P x Let x be the frequency of the segment of length x;

[0115] (M6) Calculate the length distribution P of each segment in each sample as described in (M5). x The model Q for the length distribution of each segment in healthy individuals as described in Example 1 x The average value between them is denoted as M. x ;

[0116] (M7) Calculate the JS divergence between the length distribution of each sample segment and the normal length distribution model in Example 1 according to Equation (I);

[0117]

[0118] (M8) The Wilcoxon test was used to calculate the P-value between the JS divergence of malignant tumor sample cfDNA and the normal length distribution model in Example 1, and between the JS divergence of healthy human sample cfDNA and the normal length distribution model in Example 1.

[0119] The results are as follows Figure 1 As shown, the p-value is 1.58e-09, indicating a significant difference between the two. Furthermore, the JS divergence between the cfDNA fragment length distribution model of malignant tumor samples and the normal length distribution model is greater than that between the cfDNA fragment length distribution model of healthy human samples and the normal length distribution model.

[0120] Example 3: Evaluation of the performance of JS divergence in predicting malignancy and healthy individuals in different cancer types.

[0121] Fifty-eight healthy individuals from Example 2 were selected and, together with 12 cases of breast cancer, 12 cases of gastric cancer, 12 cases of colorectal cancer, and 12 cases of bile duct cancer, formed training sets for breast cancer, gastric cancer, colorectal cancer, and bile duct cancer, respectively.

[0122] In addition, 57 healthy individuals were selected and, together with 12 individuals with breast cancer, 12 with stomach cancer, 12 with colorectal cancer, and 12 with bile duct cancer, formed the breast cancer test set, stomach cancer test set, colorectal cancer test set, and bile duct cancer test set, respectively.

[0123] The performance of JS divergence in predicting malignant tumors and healthy individuals across different cancer types was evaluated using the following steps:

[0124] (Q1) Following steps (M1) to (M7) in Example 2, obtain the JS divergence between the cfDNA fragment length distribution of all samples and the normal length distribution model;

[0125] (Q2) Based on the JS divergence between healthy human samples and corresponding cancer samples and the normal length distribution model described in Example 1 in the training sets for breast cancer, gastric cancer, colon cancer, and bile duct cancer respectively, calculate the JS divergence when the Youden index is maximized.

[0126] (Q3) Define the JS divergence value at which the Youden index is maximized in the training sets for breast cancer, gastric cancer, colon cancer, and bile duct cancer, respectively, as the threshold for the corresponding cancer type, and calculate the sensitivity and specificity of the JS divergence in the corresponding cancer type test sets based on the threshold.

[0127] The results are shown in Table 1 and... Figure 2 As shown in Table 1, the performance of JS divergence in predicting malignancy and healthy individuals across different cancer types is as follows. Figure 2 The AUC values ​​were determined for different cancer types in the test set.

[0128] Table 1. Performance of JS divergence in predicting malignancy and healthy individuals across different cancer types.

[0129]

[0130] Example 4: Evaluation of the performance of JS divergence in predicting malignancy and healthy individuals in pan-cancer.

[0131] Plasma samples from 58 healthy individuals and 52 malignant tumor patients in Example 2 were selected as the training set. Additionally, plasma samples from 57 healthy individuals and 48 malignant tumor patients (12 each of breast cancer, gastric cancer, colorectal cancer, and bile duct cancer) were selected as the test set. The performance of using JS divergence to predict malignant tumors versus healthy individuals was evaluated. The specific steps are as follows:

[0132] (Q1') Following steps (M1) to (M7) in Example 2, obtain the JS divergence between the cfDNA fragment length distribution of all samples and the normal length distribution model;

[0133] (Q2') Based on the JS divergence between the cfDNA fragment length distribution of 58 healthy and 52 malignant tumor samples in the training set and the normal length distribution model, calculate the JS divergence when the Youden index value is maximized;

[0134] (Q3') Define the JS divergence value at which the Youden index is maximized as the threshold for judging whether a sample is malignant or healthy, and calculate the sensitivity and specificity of the JS divergence in the test set based on this threshold.

[0135] The results showed that the JS divergence value was 0.0014 when the Youden index was at its maximum in the training set. Using 0.0014 as the threshold for judging whether a sample was malignant or healthy, the JS divergence showed a specificity of 0.91 and a sensitivity of 0.88 on the test set. The AUC of JS divergence in predicting healthy versus malignant tumors on the test set was 0.92. Figure 3 As shown.

[0136] Example 5: Using JS divergence to predict randomly selected samples.

[0137] Ten samples were randomly selected from the 105 test cases described in Example 4 as test samples. The JS divergence value of their length distribution was calculated, and 0.0014 was used as the threshold for judging health and malignancy. When the JS divergence of the length distribution of the test sample is greater than or equal to 0.0014, the test sample is judged to be malignant; when the JS divergence of the length distribution of the test sample is less than 0.0014, the test sample is judged to be healthy. The specific steps are as follows:

[0138] Following steps (M1)-(M7) in Example 2, the JS divergence between the fragment length distribution of 10 test samples and the fragment length distribution model of healthy individuals in Example 1 was obtained;

[0139] Using 0.0014 as the threshold for judging whether a sample is malignant or healthy, the JS divergence of the length distribution of the 10 test samples is compared with the threshold of 0.0014. When the JS divergence of the length distribution of the test sample is ≥0.0014, the test sample is judged to be malignant; when the JS divergence of the test sample is <0.0014, the test sample is judged to be healthy.

[0140] The results are shown in Table 2. The results indicate that among the 10 randomly selected samples, there were 2 cases of cholangiocarcinoma, 2 cases of gastric cancer, 1 case of breast cancer, and 5 healthy individuals. Of the 2 cases of cholangiocarcinoma, 2 cases of gastric cancer, and 1 case of breast cancer, all were determined to be malignant. Of the 5 healthy individuals, 4 were determined to be healthy and 1 to be malignant. The accuracy rate was 90%.

[0141] All malignant tumors were detected, with a malignant tumor detection rate of 100%.

[0142] Table 2. Judgment results of random samples

[0143]

[0144] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.

Claims

1. An analytical method for comprehensively assessing changes in the length distribution of cfDNA and preserving microscopic information about fragment length in a test sample, characterized in that, The method includes the following steps: (S1) Provides the distribution of cfDNA fragment lengths in the whole genome of healthy samples. x As a normal length distribution model; (S2) Provides the distribution of cfDNA fragment lengths in the whole genome of the sample to be tested. x ; (S3) via Q x and P x The difference index between the length distribution of the sample fragments under test and the normal length distribution model was calculated; in, Step (S1) also includes the following sub-steps: (S1-1) provides the fragment length of each cfDNA in the whole genome of each of the n healthy samples; (S1-2) Calculate the number of reads supporting all fragment lengths of cfDNA in each healthy sample; (S1-3) Calculate the frequency of the number of reads supported by each fragment length of cfDNA in each healthy sample relative to the number of reads supported by all fragment lengths of cfDNA, thereby obtaining the frequency of cfDNA of each fragment length in each healthy sample. (S1-4) Calculate the average or median frequency of cfDNA for each fragment length in the n healthy samples, and use it as the cfDNA fragment length distribution model Q in the whole genome of the healthy samples. x ; Step (S2) also includes the following sub-steps: (S2-1) Provides the fragment length of each cfDNA in the whole genome of the sample to be tested; (S2-2) Calculate the number of reads supporting all fragment lengths of cfDNA in the sample to be tested; (S2-3) Calculate the frequency of the number of reads supported by each fragment length of cfDNA in the test sample relative to the number of reads supported by all fragment lengths of cfDNA, thereby obtaining the frequency of cfDNA of each fragment length in the test sample. (S2-4) Based on the frequency of cfDNA fragments of each fragment length in the sample to be tested, the distribution P of cfDNA fragment lengths in the whole genome of the sample to be tested is obtained. x ; The difference index between the length distribution of the test sample fragment and the normal length distribution model refers to the JS divergence between the length distribution of the test sample fragment and the normal length distribution model. The calculation refers to the calculation based on Q... X P X Formula (I) is used to calculate the JS divergence between the test sample and the normal length distribution model. Wherein, the M x It is P x and Q x The average value.

2. The method as described in claim 1, characterized in that, The sample is a blood sample, a pleural effusion sample, a peritoneal effusion sample, or a cerebrospinal fluid sample.

3. The method as described in claim 1, characterized in that, The cfDNA fragment is 100-220 bp in length.

4. The method as described in claim 1, characterized in that, The n is an integer ≥ 10.

5. The method as described in claim 1, characterized in that, Providing the fragment length of each cfDNA in the sample includes the following steps: (C1) Low-depth whole-genome paired-end sequencing data were obtained from the sample to obtain fastq; (C2) After filtering the obtained fastq data, it is compared with the human genome data and sorted to obtain the sorted alignment file; (C3) After sorting the alignment files, remove duplicates, remove alignment information with low alignment quality, remove alignment information that is aligned to the X chromosome and Y chromosome, extract the position information of cfDNA molecules in the sample that are aligned to the chromosome, and calculate the fragment length of each cfDNA.

6. The method as described in claim 5, characterized in that, The low-quality alignment information refers to alignment information with low matching accuracy or poor uniqueness.

7. An analytical apparatus for comprehensively assessing changes in the length distribution of cfDNA and preserving microscopic information about fragment length in cfDNA samples, characterized in that, The device includes the following modules: (Z1) Normal length distribution model acquisition module, wherein the normal length distribution model acquisition module is configured to: provide the cfDNA fragment length distribution Q in the whole genome of a healthy sample. x As a normal length distribution model; (Z3) Sample analysis module, which is configured to provide the cfDNA fragment length distribution P in the whole genome of the sample to be tested. x ; (Z4) Divergence analysis module, which is configured to: use Q x and P x The JS divergence between the length distribution of the sample fragments to be tested and the normal length distribution model is calculated. The (Z1) module further includes the following sub-modules: (Z1-1) Fragment length acquisition module, wherein the fragment length acquisition module is configured to provide the fragment length of each cfDNA in the whole genome of each of n healthy samples; (Z1-2) Statistical calculation module, which is configured to: calculate the number of reads supported by cfDNA of all fragment lengths in each healthy sample; and calculate the frequency of the number of reads supported by cfDNA of each fragment length relative to the number of reads supported by cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in each healthy sample. (Z1-3) Model output module, configured to: calculate the average or median frequency of cfDNA for each fragment length in the n healthy samples, as the normal length distribution model Q. x ; The (Z3) module also includes the following sub-modules: (Z3-1) Fragment length acquisition module, the fragment length acquisition module is configured to: provide the fragment length of each cfDNA in the whole genome of the sample to be tested; (Z3-2) Statistical calculation module, the statistical calculation module is configured to: calculate the number of reads supported by cfDNA of all fragment lengths in the sample to be tested; calculate the frequency of the number of reads supported by cfDNA of each fragment length in the sample to be tested relative to the number of reads supported by cfDNA of all fragment lengths, thereby obtaining the frequency of cfDNA of each fragment length in the sample to be tested. (Z3-3) Sample fragment length output module, configured to: obtain the cfDNA fragment length distribution P in the whole genome of the sample based on the frequency of cfDNA of each fragment length in the sample. x , In module (Z4), the calculation refers to the calculation based on Q. X P X Formula (I) is used to calculate the JS divergence between the test sample and the normal length distribution model. Wherein, the M x It is P x and Q x The average value.

8. The analytical apparatus as described in claim 7, characterized in that, Between modules (Z1) and (Z3), there is also: (Z2) Threshold providing module, the threshold providing module is configured to: provide the JS divergence between the cfDNA fragment length distribution of the first malignant tumor sample and the cfDNA fragment length distribution of the first healthy sample and the normal length distribution model, thereby providing the threshold Y1 for malignant tumor determination.

9. The analytical apparatus as described in claim 8, characterized in that, The threshold providing module further includes the following sub-modules: (Z2-1) Provides the JS divergence between the cfDNA fragment length distribution of the first malignant tumor sample and the normal length distribution model; and provides the JS divergence between the cfDNA fragment length distribution of the first healthy sample and the normal length distribution model, respectively; (Z2-2) Calculate the JS divergence when the Youden index is at its maximum, and thus obtain the threshold Y1 for determining malignancy.

10. The analytical apparatus as claimed in claim 9, characterized in that, The threshold Y1 is 0.0014.

11. The analytical apparatus as claimed in claim 7, characterized in that, Following the (Z4) module, it also includes: (Z5) Result output module, the result output module is configured to: compare the JS divergence Y2 of the cfDNA fragment length distribution of the test sample with the normal length distribution model with the threshold Y1 for malignant tumor determination; if Y2≥Y1, it indicates that the test sample is a malignant tumor sample; if Y2<Y1, it indicates that the test sample is not a malignant tumor sample.

12. The analytical apparatus as claimed in claim 8, characterized in that, The malignant tumors include: breast cancer, gastric cancer, colon cancer, bile duct cancer, or combinations thereof.

Citation Information

Patent Citations

  • Microsatellite instability analysis method and analysis device

    CN112037859A

  • Cancer non-invasive early screening method and system based on cfDNA fragment length distribution characteristics

    CN117316278A