Method for identifying tumor signal in plasma sample

By integrating the fragmentation features within the repetitive region and the machine learning model through the cfRAFF method, the problem of inaccurate cfDNA fragment features was solved, and efficient identification and early detection of tumor signals were achieved.

WO2025208947A1PCT designated stage Publication Date: 2025-10-09SHANGHAI WEIHE MEDICAL LAB CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/141941
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-02
Filing Date
2024-12-24
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing technologies for tumor signal detection have problems such as inaccurate cfDNA fragment characteristics and dilution of signals by normal cells, resulting in insufficient sensitivity and specificity of tumor detection, especially in the early stages or in minimal residual states.

Method used

By analyzing low-depth whole-genome sequencing data of plasma samples, integrating fragmentation features within repetitive regions, and using the cfRAFF method, combined with the Dirichlet-polynomial model and machine learning model, tumor signals were identified, the uncertainty of sequence alignment of repetitive region sequencing fragments was reduced, and repetitive region elements were aggregated to the subfamily level for analysis.

Benefits of technology

It improves the accuracy and sensitivity of tumor detection, can identify early tumor samples and minimal residual disease, and outperforms the traditional whole genome fragment end sequence composition model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2024141941-FTAPPB-I100001
    Figure PCTCN2024141941-FTAPPB-I100001
  • Figure PCTCN2024141941-FTAPPB-I100002
    Figure PCTCN2024141941-FTAPPB-I100002
  • Figure PCTCN2024141941-FTAPPB-I100003
    Figure PCTCN2024141941-FTAPPB-I100003
Patent Text Reader

Abstract

The present invention provides a method for identifying a tumor signal in a plasma sample, in particular a method for identifying a tumor signal in a plasma sample by using sequence characteristics of genomic repetitive regions. More specifically, the present invention relates to the use of genomic repetitive regions, which have great potential as biomarkers for identifying tumor signals, for the early detection and management of tumor patients.
Need to check novelty before this filing date? Find Prior Art

Description

A method for identifying tumor signals in plasma samples

[0001] Cross-reference to related applications

[0002] This application claims priority to Chinese Patent Application No. 202410397159.5 filed on April 2, 2024, the entire contents of which are expressly incorporated herein by reference. Technical Field

[0003] The present invention belongs to the field of biological detection technology, specifically to the non-invasive detection of biological samples to achieve the purpose of disease (particularly tumor) detection. More specifically, the present invention provides a method for identifying tumor signals in plasma samples using sequence characteristics of genomic repetitive regions. By analyzing the sequence characteristics of repetitive regions in the genome, a tumor detection model is generated, which can identify potential tumor signals and achieve accurate detection of tumor samples and healthy samples. Background Art

[0004] Clinically, there has always been a huge demand for non-invasive detection, risk stratification, and treatment progress monitoring of tumors. In this regard, circulating tumor cells, extracellular vesicles, RNA, proteins, and metabolites, as potential non-invasive biomarkers, have made significant progress in recent years. However, cell-free DNA (cfDNA) has attracted the most attention. In a typical biological plasma specimen, the total cfDNA population represents a very rich source of biological and pathological information and has shown great potential as a multifunctional biomarker in oncology, non-invasive prenatal testing, and transplantation monitoring.

[0005] Over the past two decades, analysis of cfDNA integrity has become an important focus for the development of cfDNA-based tumor diagnostic and prognostic biomarkers, as it represents a factor independent of the genetic or epigenetic state of cfDNA and, in theory, is representative of all tumors. The mechanism by which cfDNA is released into the bloodstream is a major determinant of the plasma DNA fragmentation profile. While non-tumor cells undergo apoptotic cell death, cleavage of nucleosomal units results in the shedding of DNA fragments of approximately 167 bp or smaller, tumor cells undergo a variety of different death processes, including necrosis and autophagy, and can release much longer DNA fragments. Therefore, many studies have chosen to assess DNA integrity in plasma based on high-abundance DNA in repetitive regions, such as Alu or LINE-1 elements, as an alternative, for example, calculating Alu247 / Alu115, that is, the ratio of the number of Alu cfDNA long fragments (247bp) to the number of Alu cfDNA short fragments (115bp) (Park, MK et al. Alu cell-free DNA concentration, Alu index, and LINE-1 hypomethylation as a cancer predictor. Clin Biochem 94, 67-73 (2021)).

[0006] However, the above method has limitations in detecting tumor signals because: 1) Single-length cfDNA fragments such as Alu247 and Alu115 may come from both tumor and non-tumor cells. The Alu247 / Alu115 index calculated from this cannot fully reflect the characteristics of fragments derived from tumor cells and non-tumor cells; 2) Tumor cells in real plasma samples will be diluted by normal cells, which are more abundant, making the signal of cfDNA fragments derived from tumors in plasma weaker.

[0007] Whole genome sequencing (WGS) of cfDNA can provide a comprehensive understanding of the fragmentation characteristics of cfDNA in plasma. The study of cfDNA fragmentation patterns, also known as "fragmentomics," has been widely used to investigate nucleosome positioning footprints, copy number variation (e.g., enrichment of short cfDNA fragments can improve CNV detection rates), gene expression, transcription factor binding, and the composition of short sequences at the end of fragments. By studying the fragmentation characteristics of cfDNA in plasma, biomarkers for identifying tumor signals can be developed for early detection and management of cancer patients.

[0008] Therefore, there is a constant need in the art for methods of detecting and analyzing cfDNA for tumor diagnosis and prognosis. Summary of the Invention

[0009] Currently, most studies on cfDNA fragmentation characteristics are conducted unbiasedly across the whole genome using WGS methods. Tumor signals are relatively evenly distributed across the genome, and remain relatively weak in the early stages of tumors or when they are in a minimal residual state. Considering the high copy number characteristics of repetitive region elements such as Alu / LINE-1 across the genome, which naturally amplify signals, analyzing their cfDNA fragmentation characteristics can, on the one hand, further understand the fragmentation characteristics of cfDNA in repetitive region sequences across different cancer types, and on the other hand, this fragmentomic analysis integrating repetitive region elements may also enable people to further improve the performance of current tumor detection.

[0010] To address the problems and needs in the prior art, the present invention aims to propose a new method that does not rely on gene expression and methylation data, but directly analyzes the sequence characteristics of cfDNA fragments in repetitive regions of the genome in low-pass whole genome sequencing (LP-WGS) data of plasma samples to identify tumor signals and achieve the purpose of tumor detection.

[0011] In this study, we developed a method called cfRAFF (cfDNA Repeat-Aware Fragmentation Features) to assess the potential of repetitive region fragmentation features as biomarkers for identifying tumor signals. Unlike previous studies that used differences in fragment size or the frequency of motifs at the end of fragments across genomic regions, the cfRAFF method provided by this study integrates fragmentation features within repetitive regions within each sample, including length polymorphism entropy and motif frequency. For each repetitive region element within a sample, we clustered them into a set of related repeat subfamily units, quantitatively characterizing fragment length polymorphism within each subfamily unit. We also used a Bayesian model with a Dirichlet-multinomial distribution to correct for the effects of sequencing depth, GC bias, and other factors. Furthermore, we incorporated information on the frequency of motifs at the end of fragments within repetitive regions. Using over 700 samples, including samples from various cancers, we constructed a corresponding cfRAFF model and demonstrated that detecting sequence features within repetitive regions can identify potential tumor signals, thereby enabling accurate detection of tumor and healthy samples.

[0012] In this study, to reduce uncertainty in sequence alignment to repeat regions and adapt to low-depth whole-genome sequencing data, we clustered repeat region elements according to their subfamily classification. That is, repeat region elements belonging to the same subfamily were assigned to the corresponding subfamily for segment feature analysis. This method also significantly reduced the number of target regions for subsequent analysis (from over 5 million repeat region elements to hundreds of subfamily units).

[0013] In one aspect, the present invention provides a method for analyzing the fragmentation characteristics of repetitive regions of the genome as a biomarker for identifying tumor signals, comprising the following steps:

[0014] S1. Construct a reference database of genomic repetitive regions;

[0015] S2. Whole genome sequencing and processing of plasma samples;

[0016] S3. Determine the fragmentation characteristics of the repeat region;

[0017] S4. Build a machine learning detection model; and

[0018] S5. Apply the model to an independent validation set to further verify the performance of the model.

[0019] On the other hand, the present invention provides a machine learning detection model, which is constructed by the above method of the present invention.

[0020] On the other hand, the method and machine learning detection model provided by the present invention can be called cfRAFF (cfDNA Repeat-Aware Fragmentation Features) method and model, which can be used to identify tumor signals in samples, especially tumor signals in plasma samples (cfDNA), and can realize tumor detection.

[0021] In one aspect, the present invention provides a method and machine learning detection model for determining the fragmentation characteristics of repetitive regions in the genome as biomarkers for identifying tumor signals.

[0022] In a further aspect, the methods and models of the present invention and their repeat region fragmentation characteristics can be used as biomarkers to identify tumor signals, particularly for early detection of tumors or detection of minimal residual disease in tumors.

[0023] In yet another aspect, the present invention provides a computer program product comprising a hardware processor and a non-transitory computer-readable storage medium storing executable instructions that, when executed by the hardware processor, cause the processor to perform operations including the above-described method.

[0024] In other aspects, the present invention relates to methods and machine learning detection models for tumor diagnosis and prognosis based on cfDNA, particularly for early detection of tumors, and biomarkers for tumor diagnosis and prognosis.

[0025] The methods, models and / or computer program products of the present invention can be used to identify tumor signals, for example, can be used to determine the early stage of tumor samples from subjects. In some embodiments, tumor prediction is the likelihood of whether a test sample is a tumor sample.

[0026] Compared with existing technologies, especially existing analyses of cfDNA integrity and fragmentation characteristics and their applications in tumor detection, the main advantages of the method of the present invention include at least:

[0027] 1. Combining the high copy number characteristics of repetitive region elements in the genome, we targeted the repetitive regions and analyzed the characteristics of cfDNA fragments in the repetitive regions in plasma.

[0028] 2. In order to reduce the problem of sequence alignment of sequencing fragments in the repetitive interval and ensure that there are enough fragment data in the repetitive region for analysis under low-depth whole-genome sequencing, a group of related repetitive region elements are aggregated into the common repetitive subfamily layer. The analysis target is no longer a single repetitive region element, but a higher-level repetitive subfamily unit.

[0029] 3. In order to comprehensively describe the sequence length characteristics within the repetitive subfamily units, a new quantitative feature, namely fragment length polymorphism entropy, was developed, and the fragment length distribution within the sample repetitive subfamily units was normalized based on the Bayesian method of the Dirichlet-polynomial model.

[0030] 4. A new model cfRAFF was developed that combines the entropy of fragment length polymorphism (FLP) and the composition of fragment end sequences in repetitive regions. This model can accurately identify tumor signals, including early tumor samples, and its performance is better than the model based on the composition of fragment end sequences on the whole genome. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1: Whole-genome sequencing data analysis process, including preprocessing, alignment, and quality control of raw fastq data, and clustering and classification of cfDNA fragments associated with repetitive subfamily units.

[0032] Figure 2: Distribution of fragment lengths across the genome and within repetitive regions in CRC and control samples. CRC: colorectal cancer; Control: healthy controls.

[0033] Figure 3: Distribution of different cfDNA integrity indices in CRC, HCC, LUAD, LUSC, and control. Alu_244 / 83 represents the ratio of the number of cfDNA fragments with a length of 244 bp to the number of cfDNA fragments with a length of 83 bp on the Alu element. Others are similar. CRC: colorectal cancer, HCC: liver cancer, LUAD: lung adenocarcinoma, LUSC: lung squamous cell carcinoma, control: healthy control; L1: LINE-1

[0034] Figure 4: Distribution of mean RFE values ​​for repeat region sequence features in tumor samples and healthy controls. The mean RFE value is the average RFE value for the repeat subfamily unit in the relevant samples for each tumor type and corresponding stage. CRC: colorectal cancer, HCC: liver cancer, LUAD: lung adenocarcinoma, LUSC: lung squamous cell carcinoma, control: healthy control. nan indicates that for the healthy control samples, the tumor stage is not applicable and they are treated as a separate class.

[0035] Figure 5: ROC curves for machine learning detection models. Left: Average ROC curve for five cross-validation runs on the independent cohort training set. Right: ROC curve for the independent cohort validation set. gw_EDM indicates that the machine learning model is constructed using EDM features defined by genome-wide fragments. RFE indicates that the machine learning model is constructed using RFE features defined by fragments within repetitive regions. cfRAFF indicates that the machine learning model is constructed using both RFE and EDM features defined by fragments within repetitive regions.

[0036] Figure 6: The impact of random downsampling on the AUC of the cfRAFF model on the independent validation set. First, 20M, 10M, 4M, and 2M cfDNA fragments were randomly downsampled from the study cohort data, and repeated 5 times. For each repetition and each downsampling level, the RFE feature value on the repeat subfamily and the EDM feature value on the repeat region were calculated as described above. Subsequently, the features under the downsampled data were used to build a machine learning detection model on the training set, and its performance was tested on the independent validation set. The gray points in the figure represent the AUC values ​​of each downsampling level for each repetition, the black points represent the mean, and the error bars represent 1 standard deviation. All: indicates that the original data is used. DETAILED DESCRIPTION

[0037] It should be noted that, in the absence of conflict, the features of the embodiments and examples in this application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with examples.

[0038] definition

[0039] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. For the purpose of the present invention, the terms used herein are defined below.

[0040] Unless expressly indicated to the contrary, the indefinite articles "a" and "an" as used herein in the specification and in the claims should be understood to mean "at least one". It should also be understood that, unless expressly indicated to the contrary, in any method claimed herein that includes more than one step or action, the order of the steps or actions of the method is not necessarily limited to the order in which the steps or actions of the method are listed.

[0041] As used herein, the terms "about" or "approximately" refer to a quantity, level, value, number, frequency, percentage, dimension, size, amount, weight, or length that varies by up to 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% compared to a reference quantity, level, value, number, frequency, percentage, dimension, size, amount, weight, or length. Where a range of values ​​is provided, each value between and including the upper and lower ends of the range is specifically contemplated and set forth herein.

[0042] Throughout this specification, unless the context requires otherwise, the terms "comprising," "including," "containing," and "having" should be understood to imply the inclusion of the stated steps or elements or groups of steps or elements, but not the exclusion of any other steps or elements or groups of steps or elements. In certain embodiments, the terms "comprising," "including," "containing," and "having" are used synonymously. "Consisting of" means including but not limited to anything following the phrase "consisting of." Thus, the phrase "consisting of" indicates that the listed elements are required or mandatory, and no other elements may be present. As used herein, the terms "optional" and "optionally" include both selection and non-selection.

[0043] Throughout this specification, references to "one embodiment," "some embodiments," "a specific embodiment," and similar expressions mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the aforementioned phrases in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0044] As used herein, the term "individual" refers to a human individual. The term "healthy individual" refers to an individual who is presumed not to have cancer or disease. The term "subject" refers to an individual who is known to have or may have cancer or disease.

[0045] As used herein, the length units nt (nucleotide) and bp (base pair) can be used interchangeably and are only used as units to represent the length of DNA fragments or the spacing between nucleotide (or base) sites, and are not used to limit the single-stranded or double-stranded state of the DNA chain.

[0046] As used herein, the term "read" or "read" refers to any nucleotide sequence, including a nucleotide sequence derived from a sequence read obtained from an individual and / or an initial sequence read from a sample obtained from an individual.

[0047] As used herein, the term "reference genome" refers to any specific known genomic sequence, whether partial or complete, of any organism or virus that can be used to reference an identified sequence from a subject. For example, reference genomes for human subjects, as well as many other organisms, are found at the National Center for Biotechnology Information on ncbi.nlm.nih.gov. A "genome" refers to the complete genetic information of an organism or virus expressed as a nucleic acid sequence. In one embodiment, the reference sequence is a reference sequence of the full-length human genome.

[0048] As used herein, the term "cell-free DNA," "cell-free DNA," or "cfDNA" refers to a plurality of nucleic acid fragments that circulate in an individual's body (e.g., bloodstream) and originate from one or more cells. In addition, cfDNA may come from other sources, such as viruses, fetuses, etc.

[0049] In some embodiments, the isolated nucleic acid, particularly cfDNA, is extracted from a sample. In some embodiments, the sample is blood. In some embodiments, the sample is plasma, such as peripheral blood plasma.

[0050] The samples described herein can be obtained using standard techniques in the art. For example, a blood sample can be obtained using standard techniques, such as using a needle and syringe. In some embodiments, the blood sample is a peripheral blood sample. Alternatively, the blood sample can be a fraction of peripheral blood, such as a plasma sample. In some embodiments, cell-free DNA (cfDNA) is extracted from a sample, such as blood or plasma.

[0051] cfRAFF method and its model construction

[0052] In some embodiments, the present invention provides a method for analyzing the fragmentation characteristics of repetitive regions of the genome as a biomarker for identifying tumor signals, comprising the following steps:

[0053] S1. Construct a reference database of genomic repetitive regions;

[0054] S2. Whole genome sequencing and processing of plasma samples;

[0055] S3. Determine the fragmentation characteristics of the repeat region;

[0056] S4. Build a machine learning detection model; and

[0057] S5. Apply the model to an independent validation set to further verify the performance of the model.

[0058] In a preferred embodiment, the genomic repeat region reference database used in the method of the present invention is obtained from the human reference genome. In addition, the human reference genome preferably includes the hg19 reference genome, the hg38 reference genome, and the T2T reference genome.

[0059] In a further embodiment, the genomic repeat region reference database includes repeat region annotation information on the reference genome, preferably including one or more of the location of the repeat region, subfamily classification, family classification, repeat class, etc.

[0060] In some embodiments, the methods of the present invention, in step S1, comprise aggregating the repeat regions or not aggregating them. In preferred embodiments, the methods of the present invention comprise aggregating the repeat regions, including aggregating them to a subfamily level or a higher level, wherein the higher level is a repeat family level. In a more preferred embodiment, the aggregating of the repeat regions is aggregating them to a subfamily level.

[0061] In some embodiments, in step S1, the method of the present invention selects the region, subfamily layer, or family layer for downstream analysis based on the size of the repetitive region involved. In a preferred embodiment, when using the hg19 reference genome as a reference database for genomic repetitive regions, 224 repetitive subfamilies with a total length exceeding 1 Mb are selected for downstream analysis.

[0062] In some embodiments, in step S2, the method of the present invention extracts cfDNA from the plasma sample, constructs a double-stranded whole genome sequencing (WGS) library, and performs sequencing. The method of extracting cfDNA and constructing the sequencing library in the method of the present invention can adopt methods known in the art, such as the methods exemplified in the Examples of this application.

[0063] In some embodiments, the method of the present invention compares the raw data obtained by sequencing with a genomic repeat region reference database after quality control processing; and those that intersect with the repeat region subfamily unit are considered to be cfDNA sequences on the corresponding repeat region subfamily unit.

[0064] In some embodiments, the method of the present invention comprises calculating the repeat-aware fragmentation entropy (RFE) and the frequency of the end point motif (EDM) of the repeat region in step S3.

[0065] In a preferred embodiment, in step S3, the fragment length distribution within the repetitive subfamily unit of the plasma sample to be tested is directly counted to calculate its Shannon entropy as the repetitive region fragment length diversity entropy (RFE); or the fragment length distribution within the repetitive subfamily unit of the plasma sample to be tested is standardized to correct the influence of sequencing depth and GC preference, and then the repetitive region fragment length diversity entropy is calculated.

[0066] In a further preferred embodiment, the fragment length distribution within the repeated subfamily units of the plasma samples to be tested is normalized by the Dirichlet prior to correct for the effects of sequencing depth and GC bias, and is also preferably normalized according to the Bayesian method of the Dirichlet-polynomial model.

[0067] In other preferred embodiments, the fragment length distribution within the repetitive subfamily units of the plasma sample to be tested is normalized by calculating the GC content of the repetitive region sequence, using a local weighted regression algorithm (including the LOESS method), or using a sliding median method to correct the influence of GC bias.

[0068] In some embodiments, the methods of the present invention use 10-30 whole genome sequencing samples with a sequencing depth of 20X-40X to construct a Dirichlet prior distribution, for example, using 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 whole genome sequencing samples, for example, using a sequencing depth of 20X, 21X, 22X, 23X, 24X, 25X, 26X, 27X, 28X, 29X, 30X, 31X, 32X, 33X, 34X, 35X, 36X, 37X, 38X, 39X or 40X.

[0069] In a preferred embodiment, the method of the present invention uses 20 whole-genome sequencing samples with high sequencing depth (30X) to construct the Dirichlet prior distribution.

[0070] In some embodiments, the method of the present invention constructs the Dirichlet prior distribution by:

[0071] using the mean of the fragment length frequency distribution of whole-genome sequencing samples from healthy controls, following a Bayesian approach using the Dirichlet-polynomial model; or

[0072] Use its median to construct the Dirichlet prior distribution; or

[0073] The Dirichlet prior distribution is constructed by merging the sequencing data and calculating the fragment length frequency distribution on the merged data.

[0074] In a preferred embodiment, the method of the present invention constructs a Dirichlet prior distribution according to the Bayesian method of the Dirichlet-polynomial model, using the average value of the fragment length frequency distribution of the whole genome sequencing samples of healthy controls multiplied by any integer between 1.0 and 30.0; wherein, preferably multiplied by 1.0, 5.0, 10.0, 15.0, 20.0, 25.0 or 30.0; more preferably multiplied by 10.0, 20.0 or 30.0; most preferably multiplied by 20.0.

[0075] In some embodiments, in the method of the present invention, the length interval of the repeat region fragment is selected from 80-400, 100-300 or 100-220 bp; preferably, the length interval of the repeat region fragment is 100-300 bp.

[0076] In some embodiments, in the method of the present invention, the length distribution of the repeat region fragments with a step size of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bp in the length interval is calculated; preferably, the length distribution of the repeat region fragments with a step size of 1 bp in the length interval is calculated.

[0077] In some embodiments, in the methods of the present invention, the entropy of the length diversity of the repeat region fragments is calculated according to the following steps:

[0078] 1) Consider all cfDNA fragments with a length of 100-300 bp, with a step size of 1 bp, and count the number of fragments of each length within each repeat subfamily layer n i ;

[0079] 2) Calculate the frequency distribution of fragment lengths in the range of 100-300 bp within each repeat subfamily layer within the sample, i.e., p = [p1, ..., p201], where:

[0080] in represents the number of all segments in each repeat subfamily layer;

[0081] 3) Calculate the Shannon entropy of fragment length diversity:

[0082] 4) Further, assume that n i Satisfies the multinomial distribution: n i ~Multinomial(p i),i∈[1,201]

[0083] Among them, the frequency p i From a Dirichlet prior: p i ~Dirichlet(α i )

[0084] Then we can use a Bayesian method of Dirichlet-polynomial model to normalize the frequency distribution of fragment lengths within the sample;

[0085] 5) The Dirichlet prior distribution was constructed using whole-genome sequencing data from 20 healthy controls with high sequencing depth (average 30X). After calculating the frequency distribution p of fragment lengths in the 100-300 bp range within each repeat subfamily layer within the sample, the average frequency distribution of fragment lengths in the 100-300 bp range within each repeat subfamily layer was further calculated.

[0086] 6) Then, the average frequency distribution is multiplied by 20 to obtain the parameter vector α of the Dirichlet prior distribution: α = [p1, ..., p 201 ]*20

[0087] 7) For each repeat subfamily layer in each sample to be tested, the Dirichlet-polynomial model is used to update the 201 fragment length frequency distributions by calculating their Bayesian posterior probabilities based on their fragment length distributions:

[0088] 8) Randomly sample 2000 times from the Dirichlet (α*) distribution to obtain 2000 sets of fragment length frequency distribution values, and calculate the Shannon entropy e of these 2000 sets of sampling results j , j=1,…,2000;

[0089] 9) Take the average value as the fragment length diversity entropy value RFE of the subfamily layer of the repeat region:

[0090] In some embodiments, the method of the present invention uses the terminal base pair sequence (EDM) at the 5' end to calculate the frequency of the terminal base pair sequence for all cfDNA fragments belonging to the repetitive region in the sample as the frequency of the fragment end composition characteristic of the repetitive region.

[0091] In a preferred embodiment, the terminal base pair sequence used in the method of the present invention is a base sequence extending from the 5' end to the 3' direction, including 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bp, preferably 4 bp.

[0092] In a preferred embodiment, the terminal base pair sequence used in the method of the present invention is a base sequence extending from the 5' terminal position in both upstream and downstream directions, comprising 2, 4, 6, 8 or 10 bp, preferably 4 bp.

[0093] In a preferred embodiment, the terminal base pair sequence used in the method of the present invention is a base sequence extending in the 3' direction or in both the upstream and downstream directions 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bp upstream or downstream from the 5' end, including 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bp, preferably 4 bp.

[0094] In some embodiments, in step S4, the method of the present invention constructs a machine learning detection model by using a machine learning framework based on python sklearn, wherein the machine learning algorithms used include xgboost, random forest, support vector machine and logistic regression.

[0095] In some embodiments, the method of the present invention comprises the following steps in step S4:

[0096] 1) Perform a 5-fold cross-validation split on the training set and repeat it 5 times;

[0097] 2) Using the random forest-based boruta feature screening method to screen sample features or not perform feature screening;

[0098] 3) Using the Bayesian method hyperopt, we split the training set into parts for training and then perform an inner 5-fold cross-validation to optimize the hyperparameters of the corresponding machine learning algorithm;

[0099] 4) Determine the optimal machine learning algorithm and feature selection strategy by comparing the average AUC values ​​of five 5-fold cross-validation training runs on the training set;

[0100] 5) Using the entire training set to fit the optimal algorithm, a tumor detection model is constructed;

[0101] 6) Apply the model to an independent validation set to further verify the model performance, including AUC value, sensitivity and specificity, where:

[0102] AUC value = area under the receiver operating characteristic curve and the coordinate axis;

[0103] Sensitivity = true positive / (true positive + false negative);

[0104] Specificity = true negative / (true negative + false positive).

[0105] In some embodiments, the method of the present invention is in step S4, wherein in 6), a prediction score at a specified specificity level is determined from healthy control samples as a threshold,

[0106] Wherein, the specificity level is specified as any value between 90% and 99%, preferably 95%, 96%, 97%, 98% or 99%, as the threshold value,

[0107] When the prediction score of the test sample obtained by applying the model to the independent validation set is greater than or equal to the threshold, it is determined to be a tumor sample; otherwise, it is determined to be a non-tumor sample.

[0108] In some embodiments, the methods of the present invention are used for sequencing depths of 0.1X, 0.2X, 0.3X, 0.4X, 0.5X, 0.6X, 0.7X, 0.8X, 0.9X, 1.0X, 2.0X, 3.0X, 4.0X, 5.0X, 6.0X, 7.0X, 8.0X, 9.0X, 10.0X, etc. or higher, with a sequencing depth of 3.0X being preferred.

[0109] In some embodiments, the machine learning detection model constructed by the method of the present invention can be used to determine the fragmentation characteristics of the repeat region as a biomarker for identifying tumor signals, wherein preferably the biomarker is used as a biomarker for early detection of tumors or detection of minimal residual disease in tumors.

[0110] In some embodiments, the present invention provides a computer program product comprising a hardware processor and a non-transitory computer-readable storage medium storing executable instructions that, when executed by the hardware processor, cause the processor to perform operations including the methods described herein.

[0111] In some embodiments, the methods disclosed herein can be implemented in a computer system and / or a computer program product comprising a computer program mechanism embedded in a computer-readable storage medium. In addition, any method of the present invention can be implemented in one or more than one computer program product or computer system. Some embodiments of the present invention provide a computer system or a computer program product that is encoded or has instructions for executing any or all of the methods disclosed herein. These methods / instructions can be stored in a non-transitory computer-readable storage medium, such as a CD-ROM, a DVD, a disk storage product, or any other computer-readable data or program storage product. These methods encoded in the computer program product can also be distributed electronically via the Internet or other means by digital or carrier wave transmission computer data signals.

[0112] Application of the present invention and detection of tumors

[0113] In some embodiments, the method of the present invention, model and / or computer program product can be used to detect the presence of a tumor. For example, as described herein, the method of the present invention and model can be used to generate a predicted score for a test sample from a subject. In certain embodiments, the predicted score is compared with a threshold value to determine whether the subject has a tumor. In other embodiments, the predicted score can be used to make or influence clinical decisions (e.g., diagnosis of a tumor, treatment selection, assessment of therapeutic effect, etc.). For example, in one embodiment, if the predicted score exceeds a threshold value, the doctor can formulate appropriate treatment.

[0114] As described above, the methods, models and / or computer program products of the present invention can be used for early detection of tumors, for example, can be used to determine early tumor samples from subjects. In some embodiments, tumor prediction is the possibility of whether a test sample is a tumor sample.

[0115] Tumors / cancers that can be detected using the methods, models, and / or computer program products of the present invention can be malignant, benign, metastatic, or precancers, including but not limited to many solid tumors of acute and chronic leukemias, lymphomas, mesenchymal, or epithelial tissues. More specific examples of such tumors / cancers include but are not limited to squamous cell carcinoma, skin cancer, melanoma, retinoblastoma, lung cancer (including small cell lung cancer, non-small cell lung cancer, lung adenocarcinoma, and lung squamous cell carcinoma), gastric cancer, pancreatic cancer, cervical cancer, ovarian cancer, liver cancer, hepatocellular carcinoma, bladder cancer, testicular cancer, breast cancer, brain cancer, glioma, colorectal cancer, uterine cancer, salivary gland cancer, kidney cancer, prostate cancer, vulvar cancer, thyroid cancer, anal cancer, penile cancer, head and neck cancer, esophageal cancer, nasopharyngeal cancer, laryngeal cancer, hematologic malignancies, Kaposi's sarcoma, schwannoma, oligodendroglioma, neuroblastoma, rhabdomyosarcoma, osteosarcoma, leiomyosarcoma, and urinary tract cancer.

[0116] Example

[0117] This application will illustrate the beneficial effects of the present invention through the following examples. Those skilled in the art will recognize that these examples are illustrative and not restrictive. These examples do not limit the scope of the present invention in any way. The experimental techniques and experimental methods described in the following examples, unless otherwise specified, are all conventional technical methods. For example, the experimental methods in the following examples where specific conditions are not specified are generally carried out under conventional conditions or under conditions recommended by the manufacturer; unless otherwise specified, reagents and materials are all commercially available and can be obtained through regular commercial channels.

[0118] Example 1: Construction of a repeat region reference database and aggregation of repeat elements

[0119] Repeat region annotation information for the human hg19 reference genome was downloaded from the UCSC hg19 repeatmasker database (https: / / genome.ucsc.edu / cgi-bin / hgTables?db=hg19&hgta_group=rep&hgta_track=rmsk&hgta_table=rmsk&hgta_doSchema=describe+table+schema). This information includes the location of the repeat region, subfamily classification, family classification, and repeat class. Subsequently, approximately 5 million repeat elements were clustered into corresponding subfamily layers. This repeat subfamily layer contains approximately 3,800 layers. Given the low number of cfDNA fragments in shorter regions resulting from low-depth whole-genome sequencing, 224 repeat subfamilies with a total length exceeding 1 Mb (at a sequencing depth of 2X, at least 10,000 cfDNA fragments were available for subsequent analysis) were selected as target regions for repeat region analysis. Approximately half of these repeat subfamilies were Alu and LINE-1.

[0120] Example 2: Whole genome sequencing and processing of plasma samples

[0121] For the collected plasma samples, cfDNA is first extracted and a standard double-stranded whole genome sequencing (WGS) library is constructed. Sequencing is then performed on the Illumina NovaSeq 6000 sequencer, and the sequencing data is quality controlled. The specific steps include:

[0122] 1. cfDNA was extracted from plasma separated from 8-10 ml of whole blood using the QIAamp Circulating Nucleic Acid Kit (QIAGEN);

[0123] 2. Quantification was performed using the Qubit dsDNA HS assay (Thermo Fisher Scientific);

[0124] 3. Extract 5 ng of cfDNA and construct a double-stranded WGS library using the IDT xGen Prism DNA Library Prep Kit and refer to its protocol;

[0125] 4. Perform 2x150bp sequencing on a NovaSeq 6000 sequencer;

[0126] 5. The raw signal data obtained from the sequencing machine is converted into fastq format using the bcl2fastq tool provided with the sequencer;

[0127] 6. Use fastp (https: / / github.com / OpenGene / fastp) to remove the adapters and single molecule identifiers (UMIs) introduced during library construction;

[0128] 7. Use bwa-mem2 (https: / / github.com / bwa-mem2 / bwa-mem2) to align the processed fastq data to the human reference genome hg19;

[0129] 8. Use sambamba (https: / / lomereiter.github.io / sambamba / ) to mark repeat sequences. Repeat sequences here refer to two paired reads whose 5' ends are aligned at exactly the same position on the reference genome.

[0130] 9. Remove these duplicate sequences, as well as sequences that failed to align or were incorrectly paired, had poor alignment quality, were multiple aligned, or overlapped with genomic blacklist regions, to generate bed format results. Genomic blacklist regions are regions on the genome known to bias sequencing data alignment results. Here, we use the ENCODE blacklist as a reference (https: / / github.com / Boyle-Lab / Blacklist).

[0131] 10. Use bedtools (https: / / bedtools.readthedocs.io / en / latest / ) to compare the BED format results generated by the sample with the above-mentioned repetitive sequence reference database. cfDNA fragments that intersect with repetitive region subfamily units are considered to be cfDNA sequences on the corresponding repetitive region subfamily units (as shown in Figure 1).

[0132] Example 3: Defining sequence features of repetitive regions

[0133] 3.1 Dirichlet prior for the multinomial distribution of repeat subfamily unit lengths

[0134] First, plasma samples were collected from 20 healthy volunteers, and cfDNA was extracted. Standard double-stranded WGS libraries were constructed and sequenced on the NovaSeq 6000 sequencer, generating an average of approximately 100 GB of sequencing data. After data processing (as shown in Figure 1), cfDNA fragments representing the 224 predefined repeat subfamily units were obtained.

[0135] Subsequently, for each sample, the frequency distribution of cfDNA fragment lengths within the 100-300 bp range within each repeat subfamily unit within the sample was statistically calculated. Furthermore, the average frequency distribution of this fragment length range across 20 healthy control samples within each repeat subfamily unit was calculated. This distribution was multiplied by 20 and used as the Dirichlet prior for the multinomial distribution of fragment lengths within the repeat subfamily units of the plasma samples to help standardize the fragment length distribution within the repeat subfamily units of the plasma samples to remove potential interference from sequencing depth, GC bias, or other unknown factors.

[0136] 3.2 Repeat-aware fragmentation entropy (RFE)

[0137] The cfDNA fragments mapped to the repeat region are aggregated according to the repeat subfamily layer to which the repeat region elements belong, and the Shannon entropy of the fragment length diversity is calculated based on the repeat subfamily layer. The specific steps include:

[0138] 1. Consider all cfDNA fragments with a length of 100-300 bp, with a step size of 1 bp, and count the number of fragments of each length within each repeat subfamily layer n i ;

[0139] 2. Calculate the frequency distribution of fragment lengths in the range of 100-300 bp within each repeat subfamily layer within the sample, i.e., p = [p1, ..., p201], where:

[0140] in represents the number of all segments in each repeat subfamily layer;

[0141] 3. Calculate the Shannon entropy of fragment length diversity:

[0142] 4. Further, assume that n i Satisfies the multinomial distribution: n i ~Multinomial(p i ), i∈[1,201]

[0143] Among them, the frequency p i From a Dirichlet prior: p i ~Dirichlet(α i )

[0144] Then we can use a Bayesian method of Dirichlet-polynomial model to normalize the frequency distribution of fragment lengths within the sample;

[0145] 5. The Dirichlet prior distribution was constructed using the previously described high-depth whole-genome sequencing data (average 30X) of the 20 healthy controls. After calculating the frequency distribution p of fragment lengths in the 100-300 bp range within each repeat subfamily layer within the sample, the average frequency distribution of fragment lengths in the 100-300 bp range within each repeat subfamily layer was further calculated.

[0146] 6. Then, multiply the average frequency distribution by 20 to obtain the parameter vector α of the Dirichlet prior distribution: α = [p1, ..., p 201 ]*20

[0147] 7. For each repeat subfamily layer within each sample to be tested, the Dirichlet-polynomial model is used to update the 201 fragment length frequency distributions based on their fragment length distributions by calculating their Bayesian posterior probabilities:

[0148] 8. Randomly sample 2000 times from the Dirichlet (α*) distribution to obtain 2000 sets of fragment length frequency distribution values, and calculate the Shannon entropy e of these 2000 sets of sampling results j , j=1,…,2000;

[0149] 9. Take the average value as the fragment length diversity entropy value RFE of the subfamily layer of the repeat region:

[0150] 3.3 Defining the end composition characteristics of the repeat region

[0151] The frequency of the four-base sequence at the 5' end of the fragment (endpoint motif, EDM) is calculated for all cfDNA fragments within the repetitive region of the sample (each base position has four possible bases, so there are 4^4 = 256 possible sequences). Here, the four-base sequence at the 5' end of the fragment refers to the four consecutive bases in the 5' to 3' direction on the genome where the 5' end of the cfDNA fragment is located after alignment to the reference genome.

[0152] Furthermore, for the terminal base pair sequence, a base sequence of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bp can be used in addition. Such a base sequence can extend from the 5' end in the 3' direction, or extend in both upstream and downstream directions from the 5' end position, and can also extend in the 3' direction or in both upstream and downstream directions 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 bp upstream or downstream from the 5' end. In other words, the terminal base pair sequence is a base sequence that extends in the 3' direction or in both upstream and downstream directions 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 bp upstream or downstream from the 5' end.

[0153] Example 4: Building a machine learning detection model

[0154] The machine learning detection model was developed using the Python sklearn-based machine learning framework. The implemented machine learning algorithms include xgboost, random forest, support vector machine, and logistic regression. The key steps are:

[0155] 1) Perform a 5-fold cross-validation split on the training set and repeat it 5 times;

[0156] 2) Using the random forest-based boruta feature screening method to screen sample features or not perform feature screening;

[0157] 3) Using the Bayesian method hyperopt (https: / / hyperopt.github.io / hyperopt / ), we split the training set into training parts and then perform an inner 5-fold cross-validation to optimize the hyperparameters of the corresponding machine learning algorithm;

[0158] 4) Determine the optimal machine learning algorithm and feature selection strategy by comparing the average AUC (area under the curve, defined as the area under the receiver operating characteristic (ROC) curve and the coordinate axis) values ​​of five 5-fold cross-validation training runs on the training set;

[0159] 5) Using the entire training set to fit the optimal algorithm, a tumor detection model is constructed;

[0160] 6) Apply the model to an independent validation set to predict whether the sample is tumorous or non-tumorous, and further verify the model's performance, including AUC value, sensitivity, and specificity: AUC value = area under the receiver operating characteristic curve and the coordinate axis; Sensitivity = true positive / (true positive + false negative) Specificity = true negative / (true negative + false positive)

[0161] Sensitivity refers to the proportion of samples that are actually positive (tumor samples) that are judged as positive, while specificity refers to the proportion of samples that are actually negative (healthy control samples) that are judged as negative.

[0162] To calculate the sensitivity of the test at a given specificity level, we first determine a threshold for the prediction score at a given specificity level, such as 95% or 98%, based on the sample prediction scores obtained in five cross-validation runs. This threshold indicates that 95% and 98% of the prediction scores for the healthy control samples are below the threshold. The model is then applied to the validation set to obtain the prediction scores for independent test samples. Samples with scores greater than or equal to the threshold are classified as tumor samples; otherwise, they are classified as non-tumor samples. The corresponding sensitivity score is then calculated.

[0163] Example 5: Study of cfRAFF Method

[0164] 1. Sample study cohort

[0165] To validate the ability of repetitive region cfDNA sequence signatures to identify tumor signals in plasma samples, we used a low-depth whole-genome sequencing cohort. This cohort consisted of 519 training samples and 226 validation samples. The training set included 84 colorectal cancer (CRC), 87 hepatocellular carcinoma (HCC), 53 lung adenocarcinoma (LUAD), 26 lung squamous cell carcinoma (LUSC), and 269 healthy controls. The validation set included 53 CRC, 41 HCC, 27 LUAD, 12 LUSC, and 93 healthy controls. The distribution of tumor type, stage, and gender was consistent across the training and validation sets.

[0166] After extracting 5 ng of cfDNA from all plasma samples in the cohort, a double-stranded WGS library was constructed and 2x150bp low-depth whole-genome sequencing (average sequencing quantity of approximately 10G) was performed on the NovaSeq6000 sequencer. After data analysis and processing as shown in Figure 1, the cfDNA fragments within the repetitive subfamily unit of each sample were obtained.

[0167] 2. Distribution of fragment lengths in repeated regions

[0168] Due to the sequence similarity between different genomic regions, repetitive regions pose significant challenges to next-generation sequencing data analysis, particularly short-read sequencing, such as multiple alignment. Consequently, sequence lengths may exhibit different distributions compared to those observed across the entire genome. For comparison, we selected four samples from the aforementioned training cohort: three colorectal cancer (CRC) samples with varying ctDNA (cfDNA) content and one healthy control sample. We performed low-depth whole-genome sequencing on each sample and analyzed the cfDNA fragment length distribution across the different samples. As shown in Figure 2, the fragment length distribution within repetitive regions is very similar to the genome-wide fragment length distribution. For example, a major peak of approximately 167 bp and a much smaller secondary peak of approximately 334 bp are observed in both regions in the healthy control sample, corresponding to one and two nucleosome units, respectively. In addition, a small peak in subnucleosomal fragment size can be observed in CRC samples with high ctDNA content in the figure, with a period of approximately 10 bp. This may be related to the accessibility of the DNA minor groove to endonucleases when DNA wraps around the histone core, as well as the binding of transcription factors and other DNA-binding proteins. This pattern of cfDNA fragmentation within repetitive regions is similar to that within the whole genome, indicating that fragmentomics may be successfully applied to genomic repetitive regions.

[0169] 3. cfDNA integrity index

[0170] The cfDNA integrity index refers to the ratio of the number of long-fragment cfDNA to the number of short-fragment cfDNA in plasma. We calculated the cfDNA integrity indices of commonly used Alu and LINE-1 elements and their distribution in CRC, HCC, LUAD, LUSC, and control in the training set. As shown in Figure 3, relative to the control, except for HCC, which showed a relatively low trend in almost all indices, the relative changes of other cancer types were not so obvious. This shows that simply using a single indicator within the repetitive region to characterize the sample has significant limitations and cannot effectively identify tumor samples from healthy control samples.

[0171] 4. Repeat region sequence features can distinguish tumor and healthy control samples

[0172] To fully explore the length-related characteristics of genomic repetitive regions, we aggregated the repetitive region database from the most basic element unit to the repeat subfamily level. Specifically, cfDNA sequences belonging to the same repeat subfamily unit but belonging to different repeat elements were attributed to that repeat subfamily unit for analysis. Here, we considered the length range of 100-300 bp and used Shannon entropy to characterize this sequence length polymorphism within each repeat subfamily unit within a sample, defining it as RFE. To further eliminate the potential influence of sequencing depth and GC content on this characteristic, we assumed that the length of cfDNA fragments within each repeat subfamily unit satisfies a multinomial distribution and further assumed that the probability of generating fragments at each length was derived from a Dirichlet prior distribution. This prior was calculated based on 20 independent healthy control samples sequenced at high depth. Therefore, for each sample, we used the Bayesian posterior of the Dirichlet-multinomial model to update the fragment length distribution density within the repeat subfamily unit within the sample and then calculated its Shannon entropy RFE.

[0173] We calculated the RFE for each of the 224 repeat subfamily units in the training set of samples from the analysis cohort, including CRC, HCC, LUAD, LUSC, and healthy controls. As shown in Figure 4, the RFE distribution varied significantly across tumor types and stages compared with healthy controls. In particular, RFE was generally higher for samples from advanced tumors, such as those in stages III and IV. This suggests that, despite not pre-screening for tumor-specific repeat subfamily units for analysis, tumor samples still exhibited high fragment size polymorphism within genomic repeat subfamily units, indicating that specific regulation of repeat regions in the tumor genome influences the formation of cfDNA in plasma.

[0174] 5. Repeated region sequence features have excellent ability to identify tumor signals

[0175] We developed a method called cfRAFF (cfDNA Repeat-Aware Fragmentation Features) to further evaluate the hypothesis that fragmentation characteristics of repetitive regions can serve as biomarkers for identifying tumor signals. Unlike previous studies that used differences in fragment size or differences in the frequency of motifs at the ends of fragments across the genome, the cfRAFF method integrates fragmentation characteristics within repetitive regions within each sample, including length polymorphism entropy (RFE) and end motif composition frequency (EDM). For repetitive region elements within each sample, we clustered them into a set of related repeat subfamily units, quantitatively characterized the fragment length polymorphism within each subfamily unit, and used a Bayesian model with the Dirichlet-multinomial distribution to correct for the effects of sequencing depth, GC bias, and other factors. In addition, we incorporated information on the composition frequency of motifs at the ends of fragments within the repetitive regions.

[0176] We constructed a machine learning model on a cohort training set using over 700 samples, including samples from various cancers. We implemented four different machine learning algorithms, including xgboost, random forest, support vector machine, and logistic regression, and used a nested cross-validation method (5-fold cross-validation for both the inner and outer layers) to learn the corresponding hyperparameters on the inner layer and evaluate the model performance under the optimal hyperparameter combination on the outer layer. At the same time, we considered whether to use the boruta method for pre-feature screening during model training. By comparing these four machine learning algorithms and whether to perform feature screening, we determined the optimal algorithm and feature screening strategy, and constructed the gw_EDM model characterized by the whole genome terminal motif, the RTE model characterized by the repeat region sequence, and the cfRAFF model (as shown in Figure 5), and applied them to the research cohort validation set for independent prediction.

[0177] Results showed that the RTE feature demonstrated excellent tumor signal recognition, with AUCs of 0.795 and 0.782 on the training and validation sets, respectively. Incorporating the EDM feature at the end of the repeat region further improved the cfRAFF model's average AUC to 0.913 on the training set and 0.920 on the validation set. These metrics also outperformed conventional models based on genome-wide EDM features. Furthermore, Figure 5 shows that the cfRAFF model exhibits enhanced sensitivity at high specificity levels. At a 95% specificity level, the cfRAFF model achieved a sensitivity of 69.9% on the validation set, with sensitivity of 54.3% and 60.0% for stages I and II, respectively (Table 1). This demonstrates that by detecting sequence features within repeat regions, potential tumor signals can be identified, enabling accurate detection of both tumor and healthy samples.

[0178] Table 1. Tumor detection sensitivity of the cfRAFF model in the validation set of the study cohort

[0179] *: The numerator in the brackets represents the number of samples correctly detected at the specified specificity, and the denominator represents the total number of samples tested.

[0180] 6. Evaluating the impact of sequencing depth on cfRAFF model performance

[0181] To evaluate the impact of sequencing depth on cfRAFF model performance, we randomly downsampled the data from the aforementioned cohort to 20M (20 million), 10M, 4M, and 2M cfDNA fragments, corresponding to effective sequencing depths of approximately 1X, 0.5X, 0.2X, and 0.1X, respectively. To reduce the impact of random fluctuations in the random downsampling operation, we repeated this downsampling process five times. For each repetition and each downsampling level, we trained and constructed the corresponding cfRAFF model on the training set as described above and validated its performance on an independent validation set. As shown in Figure 6, when the sequencing data was downsampled to 10M, the model performance remained relatively stable compared to using the full original data (AUC 0.920 vs. 0.912). Even when downsampled to 2M, the AUC still reached 0.862. This demonstrates that the cfRAFF model does not require high-depth sequencing data to identify tumor signals and performs remarkably well with less sequencing data.

[0182] Alternative implementation plans

[0183] 1. In addition to being aggregated into the subfamily level, repeat regions can also be not aggregated or aggregated into a higher level, such as the repeat family level;

[0184] 2. Repeat region subfamily units: In addition to the 224 units selected based on the size of the repeat regions involved, more or fewer regions, subfamily layers, or family layers can be selected for downstream analysis based on other principles;

[0185] 3. When analyzing human plasma samples, in addition to using the hg19 reference genome, the definition of repeat regions can also refer to the hg38 reference genome, T2T reference genome, etc.;

[0186] 4. Calculation of entropy of sequence length polymorphism of repeat region fragments. In addition to the aforementioned normalization using the Bayesian method of the Dirichlet-polynomial model, this process can also be omitted. That is, the Shannon entropy can be calculated as its feature by directly counting the fragment length distribution in the repeat region within the sample;

[0187] 5. Calculation of entropy of length polymorphism of repeat region fragments. In addition to the aforementioned normalization using the Bayesian method of the Dirichlet-polynomial model, other methods can be used to reduce the influence of GC bias, such as calculating the GC content of the repeat region sequence, using a locally weighted regression algorithm such as LOESS, or using the sliding median method to correct its fragment length distribution.

[0188] 6. Calculation of the entropy of sequence length polymorphisms in repeat regions. In addition to constructing the Dirichlet prior distribution using 20 high-depth whole-genome sequencing samples (30X) according to the aforementioned Bayesian method of the Dirichlet-polynomial model, more or fewer whole-genome sequencing samples can also be used to construct the prior distribution.

[0189] 7. Calculation of entropy of sequence length polymorphisms in repeat regions. In addition to constructing the Dirichlet prior distribution using 20 high-depth whole-genome sequencing samples (30X) according to the aforementioned Bayesian method of the Dirichlet-polynomial model, whole-genome sequencing samples with higher or lower sequencing depths can also be used to construct the prior distribution.

[0190] 8. Calculation of entropy of repeat region fragment length polymorphism. In addition to using the mean value of the fragment length frequency distribution of healthy control whole genome sequencing samples according to the aforementioned Bayesian method of the Dirichlet-polynomial model to construct the Dirichlet prior distribution, the median value can also be used for construction.

[0191] 9. Calculation of entropy of sequence length polymorphisms in repeat regions. In addition to constructing the Dirichlet prior distribution using the average value of the fragment length frequency distribution of the whole genome sequencing samples of healthy controls according to the aforementioned Bayesian method of the Dirichlet-polynomial model, the sequencing data can also be merged together and the fragment length frequency distribution of the merged data is calculated to construct the entropy.

[0192] 10. Calculation of entropy of repeat region fragment length polymorphism: In addition to the aforementioned Bayesian method based on the Dirichlet-polynomial model, which uses the average value of the fragment length frequency distribution of the whole genome sequencing samples of healthy controls multiplied by 20 to construct the Dirichlet prior distribution, other coefficients can also be used to construct the distribution, such as 1.0, 2.0, 10.0, 30.0, etc.

[0193] 11. Fragment length interval: In addition to the aforementioned 100-300 bp, other length intervals can also be used, such as 80-400, 100-220, etc.

[0194] 12. Calculation of fragment length distribution: In addition to the length distribution with a step size of 1 bp within the aforementioned calculation interval, the length distribution with other step sizes can also be calculated, such as 2, 3, 4, 5, 6, 7, 8, 9, 10 bp, etc.

[0195] 13. The sequence end composition characteristics of the cfRAFF model can be combined with other length sequences, such as 1, 2, 3, 5, 6, 7, 8, 9, 10 bp, in addition to the aforementioned 4 bp base sequence from the terminal 5' to 3' direction;

[0196] 14. The cfRAFF model combines the sequence end composition features. In addition to the aforementioned 4 bp base sequence from the 5' to 3' direction of the terminal, the base composition in the upstream and downstream directions of the terminal position can also be used. The length can be 2, 4, 6, 8, 10 bp, etc.

[0197] 15. The terminal sequence composition characteristics of the cfRAFF model can be combined. In addition to the aforementioned base sequence of 0 bp from the terminal position and 4 bp from the 5' to 3' direction, different base sequences of 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 bp upstream and downstream from the terminal position can also be used. The sequence length can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 bp, etc.

[0198] 16. In addition to its aforementioned applications in early tumor detection, the cfRAFF model can also be used in the detection of minimal residual disease.

[0199] 17. The whole-genome sequencing depth applicable to the cfRAFF model, in addition to the approximately 3X depth used above, can also be other sequencing depths, such as 0.1X, 0.2X, 0.3X, 0.4X, 0.5X, 0.6X, 0.7X, 0.8X, 0.9X, 1.0X, 2.0X, 4.0X, 5.0X, 6.0X, 7.0X, 8.0X, 9.0X, 10.0X, etc., or depths in between or higher sequencing depths.

[0200] The foregoing detailed description refers to the examples and drawings, which illustrate specific embodiments of the present disclosure. Other embodiments with different structures and operations do not depart from the scope of the present disclosure, and reference may be made to certain specific examples of the many alternative aspects or embodiments of the applicant's invention set forth in this specification, and none of them is intended to limit the scope of the applicant's invention or the claims. It will be understood that the various details of the present disclosure may be changed without departing from the scope of the present disclosure. In addition, the foregoing description is for illustrative purposes only and not for limiting purposes.

Claims

1. A method for analyzing the fragmentation characteristics of repetitive regions of the genome as a biomarker for identifying tumor signals, comprising the following steps: S1. Build a reference database of genomic repeat regions; S2. Whole genome sequencing and processing of plasma samples; S3. Determine the fragmentation characteristics of the repeat region; S4. Build a machine learning detection model; and S5. Apply the model to an independent validation set to further verify the performance of the model.

2. The method according to claim 1, wherein the genomic repeat region reference database is obtained from the human reference genome, preferably including the hg19 reference genome, the hg38 reference genome, and the T2T reference genome; The genomic repeat region reference database includes repeat region annotation information on the reference genome, preferably including one or more of the position, subfamily classification, family classification and repeat class of the repeat region.

3. The method according to claim 1, wherein in step S1, the repeat region is polymerized or not polymerized; Preferably, clustering the repeat regions comprises clustering to a subfamily level or a higher level, wherein the higher level is a repeat family level; in, Repeat regions are preferably clustered into subfamily levels.

4. The method of claim 3, wherein the regions, subfamily layers or family layers are selected for downstream analysis according to the size of the repetitive regions involved, When using the hg19 reference genome as the reference database for genomic repeat regions, 224 repeat subfamilies with a total length of more than 1 Mb were selected for downstream analysis.

5. The method according to claim 1, wherein in step S2, cfDNA is extracted from the plasma sample to construct a double-stranded whole genome sequencing library for sequencing; After the raw data obtained from sequencing has been quality controlled, it is compared with the genomic repeat region reference database; Those that intersect with the repeat region subfamily units are considered to be cfDNA sequences on the corresponding repeat region subfamily units.

6. The method according to claim 1, wherein step S3 comprises calculating the repeat-aware fragmentation entropy (RFE) and the frequency of the end point motif (EDM) of the repeat region.

7. The method according to claim 6, wherein the Shannon entropy is calculated by directly counting the fragment length distribution within the repetitive subfamily unit of the plasma sample to be tested, and the Shannon entropy is used as the repetitive region fragment length diversity entropy (RFE); or the fragment length distribution within the repetitive subfamily unit of the plasma sample to be tested is normalized to correct the effects of sequencing depth and GC bias, and then the repetitive region fragment length diversity entropy is calculated.

8. The method according to claim 7, wherein the fragment length distribution within the repeated subfamily unit of the plasma sample to be tested is normalized by Dirichlet prior to correct the influence of sequencing depth and GC bias, in, Normalization is preferably performed according to the Bayesian method of the Dirichlet-polynomial model.

9. The method according to claim 7, wherein the fragment length distribution within the repeat subfamily unit of the plasma sample to be tested is normalized by calculating the GC content of the repeat region sequence and using a local weighted regression algorithm, including the LOESS method, or using a sliding median method to correct the influence of GC bias.

10. The method according to claim 8, wherein 10-30 whole genome sequencing samples with a sequencing depth of 20X-40X are used to construct the Dirichlet prior distribution. Preferably, 20 whole genome sequencing samples with a sequencing depth of 30X are used to construct the Dirichlet prior distribution.

11. The method according to claim 8, wherein the Dirichlet prior distribution is constructed by: Following the Bayesian approach of the Dirichlet-multinomial model, the Dirichlet prior distribution is constructed using the mean of the fragment length frequency distribution of the whole-genome sequencing samples of healthy controls; or Use its median to construct the Dirichlet prior distribution; or The Dirichlet prior distribution is constructed by merging the sequencing data and calculating the fragment length frequency distribution on the merged data.

12. The method according to claim 11, wherein the Dirichlet prior distribution is constructed using the Bayesian method of the Dirichlet-polynomial model by multiplying the mean of the fragment length frequency distribution of the whole genome sequencing samples of the healthy control by any integer between 1.0 and 30.0; in, Preferably, it is multiplied by 1.0, 5.0, 10.0, 15.0, 20.0, 25.0 or 30.0; more preferably, it is multiplied by 10.0, 20.0 or 30.0; and most preferably, it is multiplied by 20.

0.

13. The method according to claim 7, wherein the length interval of the repeat region fragment is selected from 80-400, 100-300 or 100-220 bp; preferably, the length interval of the repeat region fragment is 100-300 bp.

14. The method according to claim 13, wherein the length distribution of the length interval of the repeat region fragments is calculated with a step size of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bp; preferably, the length distribution of the length interval of the repeat region fragments is calculated with a step size of 1 bp.

15. The method according to claim 14, wherein the entropy of the length diversity of the repeat region fragments is calculated according to the following steps: 1) Consider all cfDNA fragments with a length of 100-300 bp, with a step size of 1 bp, and count the number of fragments of each length within each repeat subfamily layer n i ; 2) Calculate the frequency distribution of fragment lengths in the range of 100-300 bp within each repeat subfamily layer within the sample, i.e., p = [p1, ..., p201], where: in represents the number of all segments in each repeat subfamily layer; 3) Calculate the Shannon entropy of fragment length diversity: 4) Further, assume that n i Satisfies the multinomial distribution: n i ~Multinomial(p i ),i∈[1,201] Among them, the frequency p i From a Dirichlet prior: p i ~Dirichlet(a i ) Then we can use a Bayesian method of Dirichlet-polynomial model to normalize the frequency distribution of fragment lengths within the sample; 5) The Dirichlet prior distribution was constructed using whole-genome sequencing data from 20 healthy controls with an average sequencing depth of 30X. After calculating the frequency distribution p of fragment lengths in the 100-300 bp range within each repeat subfamily layer within the sample, the average frequency distribution of fragment lengths in the 100-300 bp range within each repeat subfamily layer was further calculated. 6) Then, multiply the mean frequency distribution by 20 to obtain the parameter vector α of the Dirichlet prior distribution: α=[p1,...,p 201 ]*20 7) For each repeat subfamily layer in each sample to be tested, the Dirichlet-polynomial model is used to update the 201 fragment length frequency distributions by calculating their Bayesian posterior probabilities based on their fragment length distributions: 8) Randomly sample 2000 times from the Dirichlet (α*) distribution to obtain 2000 sets of fragment length frequency distribution values, and calculate the Shannon entropy e of these 2000 sets of sampling results j , j=1,…,2000; 9) Take the average value as the fragment length diversity entropy value RFE of the subfamily layer of the repeat region:

16. The method according to claim 6, wherein the terminal base pair sequence (EDM) at the 5' end is used to calculate the frequency of the terminal base pair sequence for all cfDNA fragments belonging to the repetitive region in the sample, as the frequency of the fragment end composition characteristic of the repetitive region.

17. The method according to claim 16, wherein the terminal base pair sequence is a base sequence extending from the 5' end to the 3' direction, including 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bp, preferably 4 bp.

18. The method according to claim 16, wherein the terminal base pair sequence is a base sequence extending from the 5' terminal position in both upstream and downstream directions, comprising 2, 4, 6, 8 or 10 bp, preferably 4 bp.

19. The method according to claim 16, wherein the terminal base pair sequence is a base sequence extending in the 3' direction or in both the upstream and downstream directions 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bp upstream or downstream from the 5' end, including 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bp, preferably 4 bp.

20. The method according to claim 1, wherein in step S4, a machine learning detection model is constructed by using a machine learning framework based on python sklearn, The machine learning algorithms used include xgboost, random forest, support vector machine, and logistic regression.

21. The method according to claim 20, comprising the steps of: 1) Perform a 5-fold cross-validation split on the training set and repeat it 5 times; 2) Using the random forest-based boruta feature screening method to screen sample features or not perform feature screening; 3) Using the Bayesian method hyperopt, we split the training set into parts for training and then perform an inner 5-fold cross-validation to optimize the hyperparameters of the corresponding machine learning algorithm; 4) Determine the optimal machine learning algorithm and feature selection strategy by comparing the average AUC values ​​of five 5-fold cross-validation training runs on the training set; 5) Using the entire training set to fit the optimal algorithm, a tumor detection model is constructed; 6) Apply the model to an independent validation set to further verify the model performance, including AUC value, sensitivity and specificity, where: AUC value = area under the receiver operating characteristic curve and the coordinate axis; Sensitivity = true positive / (true positive + false negative); Specificity = true negative / (true negative + false positive).

22. The method according to claim 21, wherein in step 6), a prediction score at a specified specificity level is determined from healthy control samples as a threshold, in, Assigning a specificity level to any value between 90% and 99%, preferably 95%, 96%, 97%, 98% or 99%, as a threshold value, When the prediction score of the test sample obtained by applying the model to the independent validation set is greater than or equal to the threshold, it is determined to be a tumor sample; otherwise, it is determined to be a non-tumor sample.

23. The method of any one of claims 1-22, for a sequencing depth of 0.1X, 0.2X, 0.3X, 0.4X, 0.5X, 0.6X, 0.7X, 0.8X, 0.9X, 1.0X, 2.0X, 3.0X, 4.0X, 5.0X, 6.0X, 7.0X, 8.0X, 9.0X, 10.0X or more, with a sequencing depth of 3.0X being preferred.

24. The method according to any one of claims 1 to 22, wherein the machine learning detection model constructed by the method is used to determine the repeat region fragmentation feature as a biomarker for identifying tumor signals, in, Preferably, the biomarker is used as a biomarker for early detection of tumors or detection of minimal residual disease in tumors.

25. A computer program product comprising a hardware processor and a non-transitory computer-readable storage medium storing executable instructions that, when executed by the hardware processor, cause the processor to perform operations including the method of any one of claims 1-24.

Citation Information

Patent Citations

  • DNA methylation markers for diagnosing early liver cancer by using peripheral blood and application thereof

    CN109825584A

  • Colorectal cancer gene mutation identification method, equipment and application

    CN115807083A

  • Method for identifying tumor signal in plasma sample

    CN118471325A

  • Epigenetics analysis of cell-free DNA

    US20240043935A1

  • Repeat-aware profiling of cell-free RNA

    WO2024010875A1