Method for predicting association relationship between sample and tumor, device, medium and program
By acquiring methylation detection data of the sample to be tested, generating regional methylation features and using a classification model, the problem of accuracy and comprehensiveness in predicting the association between the sample to be tested and tumors in existing technologies has been solved, realizing accurate prediction of tumor risk and early cancer detection of the sample to be tested.
Patent Information
- Application Number
- PCT/CN2024/133122
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-22
- Filing Date
- 2024-11-20
- Publication Date
- 2025-10-30
AI Technical Summary
Existing methods are insufficient to accurately and comprehensively predict the association between a sample and a tumor, especially in early cancer detection, where DNA methylation mutation detection schemes suffer from limited information and false positives.
By acquiring methylation detection data of the sample to be tested, the target methylation level template associated with the cfDNA fragment is determined, regional methylation features are generated, and multiple methylation state models and classification models are used to generate prediction results of the association between the sample to be tested and the tumor, simplifying the algorithm and improving the information richness.
It enables accurate and comprehensive prediction of the risk of tumor formation in the tested samples, reduces harm to the human body, and improves the accuracy and comprehensiveness of early cancer detection.
Smart Images

Figure CN2024133122_30102025_PF_FP_ABST
Abstract
Description
Methods, equipment, media, and procedures for predicting the association between samples and tumors
[0001] This application claims priority to Chinese Patent Application No. 202410487905.X, filed on April 22, 2024, entitled “Method, apparatus, medium and procedure for predicting the association between a sample and a tumor”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This invention relates generally to the processing of biological information, and more specifically, to methods for predicting the association between a test sample and a tumor, methods for determining the tumor tissue origin of a test sample, methods for constructing a classification model for tumor prediction, computing devices, computer storage media, and computer program products. Background Technology
[0003] Previous studies have shown that DNA methylation variations are closely related to the occurrence of early-stage cancer. Compared with gene mutations, DNA methylation variations have the characteristics of wider coverage, higher stability, and earlier occurrence, making them more suitable for the detection of early-stage cancer.
[0004] Traditional methods for predicting the association between a test sample and a tumor include identifying gene mutations based on sequencing data or predicting the likelihood of tumor development based on DNA methylation levels. However, these traditional methods struggle to accurately and comprehensively predict tumor occurrence.
[0005] In summary, the shortcomings of traditional methods for predicting the association between test samples and tumors are that they are difficult to accurately and comprehensively predict the occurrence of tumors. Summary of the Invention
[0006] This invention provides a method for predicting the association between a test sample and a tumor, a method for determining the tumor tissue origin of the test sample, a method for constructing a classification model for tumor prediction, a computing device, a computer storage medium, and a computer program product. The method provided by this invention for predicting the association between a test sample and a tumor can conveniently and accurately predict the tumor formation risk of the test sample, while causing almost no harm to the human body.
[0007] According to a first aspect of the present invention, a method for predicting the association between a test sample and a tumor is provided. The method includes: acquiring methylation detection data of the test sample, said methylation detection data including at least CpG site information indicating a fragment of cell-free blood DNA (cfDNA) and methylation state data of each CpG site; determining a target methylation level template associated with the cfDNA fragment among a plurality of methylation level templates based on the CpG site information; generating regional methylation features based on a plurality of methylation state models associated with the target methylation level template and the methylation state data of the CpG sites covered by the cfDNA fragment, each methylation state model being associated with a tissue or blood; and generating a prediction result indicating the association between the test sample and one or more tumors via a first classification model trained on the sample data, based on the generated regional methylation features.
[0008] In some embodiments, generating a prediction result indicating the association between a test sample and one or more tumors includes: screening for regional methylation features based on screened differentially methylated regions to generate differentially methylated region features; and extracting features of the differentially methylated region features via a first classification model to generate a prediction result indicating the association between a test sample and one or more tumors.
[0009] In some embodiments, the differential methylation features are one-dimensional feature vectors, and the prediction results of the first classification model indicate the association between the sample to be tested and a type of tumor.
[0010] In some embodiments, the differential methylation features are multidimensional feature vectors, and the prediction results of the first classification model indicate the association between the sample under test and various tumors.
[0011] In some embodiments, the sample data for training the first classification model is generated by: sorting the distances between each region and the methylation distribution data of a predetermined type of cancer tissue and healthy blood, so as to determine the regions with the largest distances as the differentially methylated regions corresponding to the predetermined type of cancer tissue; and generating the sample data for training the first classification model based on the feature vectors of the methylation features on the differentially methylated regions that correspond to the dimension of the predetermined type of cancer tissue.
[0012] In some embodiments, the sample data for training the first classification model is generated by: sorting the distances between each region and the methylation distribution data of each cancer tissue and healthy blood to determine a predetermined number of regions with the largest distances as the differentially methylated regions corresponding to each cancer tissue; and merging the feature vectors of the methylation features on the differentially methylated regions with the dimension corresponding to each cancer tissue to generate the sample data for training the first classification model.
[0013] According to a second aspect of the present invention, a method for determining the tumor tissue origin of a test sample is provided. This method enables convenient and accurate comprehensive prediction of the tumor tissue origin of a test sample. The method includes: acquiring methylation detection data of the test sample, said methylation detection data including at least CpG site information indicating the coverage of cell-free blood DNA (cfDNA) fragments and methylation status data of each CpG site; based on the CpG site information, determining a target methylation level template associated with the cfDNA fragment from multiple methylation level templates, so as to generate regional methylation features based at least on the target methylation level template and the methylation status data of the CpG sites covered by the cfDNA fragment; screening the regional methylation features based on the generated regional methylation features and differentially methylated regions to generate differentially methylated regional features; and generating a prediction result indicating the tumor tissue origin of the test sample via a second classification model trained on the sample data based on the differentially methylated regional features.
[0014] In some embodiments, generating regional methylation features includes: substituting methylation state data of CpG sites covered by the cfDNA fragment into multiple methylation state models associated with a determined target methylation level template to calculate a posterior probability for each methylation state model; determining a methylation feature value associated with the target methylation level template based on the calculated posterior probability of each methylation state model; and generating regional methylation features based on the determined methylation feature value associated with the target methylation level template.
[0015] In some embodiments, generating regional methylation features based on the determined methylation feature values associated with the target methylation level template includes: converting the posterior probability of each methylation state model associated with the target methylation level template into a probability ratio; and logarithmically summing the converted probability ratios of all methylation state models associated with the target methylation level template within the region to generate regional methylation features.
[0016] In some embodiments, determining a target methylation level template associated with a cfDNA fragment among multiple methylation level templates includes: determining whether the CpG site information contained in each cfDNA fragment data in the methylation detection data of the test sample matches any one of the multiple methylation level templates; and, in response to determining that the CpG site information contained in the current cfDNA fragment data in the test result matches the current methylation level template among the multiple methylation level templates, determining the current methylation level template as the target methylation level template associated with the current cfDNA fragment data.
[0017] In some embodiments, the method for predicting the association between a test sample and a tumor or the method for predicting the tumor tissue origin of a test sample further includes any one of the following: determining multiple methylation level templates based on a plurality of consecutive CpG sites within the probe region; and screening for differentially methylated regions associated with significant methylation differences for each tissue or blood.
[0018] In some embodiments, the sample data is generated by: sorting the distances between methylation distribution data of each of two different cancer tissues to select a predetermined number of regions with the largest distances as methylation regions corresponding to the differences between the two different cancer tissues; and merging the feature vectors of the corresponding two dimensions of methylation features on the methylation regions corresponding to the differences between the two different cancer tissues to generate sample data for training a second classification model.
[0019] According to a third aspect of the present invention, a method for constructing a classification model for tumor prediction is also provided. This method enables the invention to obtain a predictive model capable of accurately and comprehensively predicting the risk of formation of various tumor tissues or the origin of tumor tissues. The method includes: acquiring methylation detection data for various tissue and blood samples, wherein the methylation detection data at least indicates CpG site information covered by cell-free DNA (cfDNA) fragments in blood and methylation status data for each CpG site; within a probe region, determining multiple methylation level templates based on a plurality of consecutive CpG sites to determine multiple associated methylation status models and corresponding model parameters for each methylation level template, each methylation status model being associated with a tissue or blood; for each tissue or blood, screening for differentially methylated regions associated with significant methylation differences; generating sample data based on the methylation status data of CpG sites covered by the cfDNA fragments of the sample, the multiple methylation status models associated with methylation level templates matching the cfDNA fragments of the sample, and the differentially methylated regions, so as to train a classification model based on the sample data.
[0020] In some embodiments, the classification model is any one of the following: a first classification model for predicting the association between a test sample and one or more tumors; or a second classification model for predicting and determining the tumor tissue origin of the test sample. In some embodiments, determining multiple methylation level templates based on a plurality of consecutive CpG sites includes: within the probe region, determining whether a current group of consecutive CpG sites meets predetermined template conditions, the predetermined template conditions including: the number of consecutive CpG sites is greater than or equal to a predetermined number threshold; the distance between the first CpG site and the last CpG site is less than or equal to a predetermined distance threshold; and the consecutive CpG sites are different from those included in a determined methylation level template; in response to determining that a current group of consecutive CpG sites meets the predetermined template conditions, determining that the current group of consecutive CpG sites is a methylation level template; and in response to determining that the current group of consecutive CpG sites does not meet the predetermined template conditions, determining whether the next group of consecutive CpG sites meets the predetermined template conditions.
[0021] In some embodiments, the predetermined bit threshold is 3 and the predetermined distance threshold is 200bp.
[0022] In some embodiments, for each tissue or blood, screening for differentially methylated regions includes: determining the distance between the methylation distributions of each of a plurality of methylation level templates for different tissues or blood; calculating the distance between the methylation distribution data of each region for different tissues or blood based on the distance between the methylation distributions of each methylation level template for different tissues or blood; sorting the plurality of regions based on the calculated distance between the methylation distribution data of each region for different tissues or blood; and determining, based on the sorting results, differentially methylated regions that exhibit significant methylation differences for different tissues or blood.
[0023] In some embodiments, determining multiple associated methylation state models and corresponding model parameters for each methylation level template includes: for each methylation level template, establishing multiple methylation state models associated with each methylation level template based on a Bernoulli mixture model; using a maximum likelihood estimation algorithm to estimate the model parameters of each methylation state model associated with the methylation level template; and based on the estimated model parameters of each methylation state model, determining a baseline model of cfDNA fragment methylation state associated with each methylation level template, for use in determining multiple methylation state models associated with a target methylation level template.
[0024] According to a fourth aspect of the invention, a computing device is also provided, the device comprising: a memory configured to store one or more computer programs; and a processor coupled to the memory and configured to execute one or more programs to cause the device to perform the methods of the first, second, or third aspects of the invention.
[0025] According to a fifth aspect of the present invention, a non-transient computer-readable storage medium is also provided. The non-transient computer-readable storage medium stores machine-executable instructions that, when executed, cause a machine to perform the methods of the first, second, or third aspect of the present invention.
[0026] According to a sixth aspect of the present invention, a computer program product is also provided. The computer program product includes instructions that, when executed by a machine, implement the methods of the first, second, or third aspect of the present invention.
[0027] The summary section is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of the invention, nor is it intended to limit the scope of the invention. Attached Figure Description
[0028] Figure 1 shows a schematic diagram of a system for implementing a method for predicting the association between a test sample and a tumor according to an embodiment of the present invention.
[0029] Figure 2 shows a flowchart of a method for predicting the association between a test sample and a tumor according to an embodiment of the present invention.
[0030] Figure 3A shows a schematic diagram of the ROC curve of the training set according to an embodiment of the present invention.
[0031] Figure 3B shows a schematic diagram of the ROC curves of the test set according to an embodiment of the present invention.
[0032] Figure 4 illustrates a schematic diagram of the distinguishing effect of methylation features constructed according to an embodiment of the present invention and those constructed according to prior art methods on cancer / health.
[0033] Figure 5 shows a flowchart of a method for screening methylated regions that differ for each type of tissue or blood according to an embodiment of the present invention.
[0034] Figure 6 shows a flowchart of a method for determining multiple associated methylation state models and corresponding model parameters for each methylation level template according to an embodiment of the present invention.
[0035] Figure 7 shows a flowchart of a method for predicting and determining the tumor tissue origin of a sample to be tested according to an embodiment of the present invention.
[0036] Figure 8 shows a flowchart of a method 800 for constructing a classification model for tumor prediction.
[0037] Figure 9 schematically illustrates the content gradient results of simulated tissue sample data according to an embodiment of the present invention.
[0038] Figure 10 schematically illustrates a block diagram of an electronic device suitable for implementing embodiments of the present invention.
[0039] In the various figures, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation
[0040] Preferred embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0041] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects.
[0042] As described above, the traditional methods for predicting the association between a sample and a tumor have the following shortcomings: they are difficult to predict the occurrence of tumors accurately and comprehensively.
[0043] To at least partially address one or more of the aforementioned problems and other potential issues, exemplary embodiments of the present invention propose a scheme for predicting the association between a test sample and a tumor. In this scheme, a target methylation level template associated with a cfDNA fragment is determined based on CpG site information of the cfDNA fragment indicated in the methylation detection data of the test sample. Regional methylation features are generated based on multiple methylation state models associated with the target methylation level template and methylation state data of CpG sites covered by the cfDNA fragment. The present invention selects methylation data of CpG sites covered by cfDNA fragments that are closely related to tumor occurrence and development and exhibit high consistency with the DNA methylation features of the genomic DNA of its tissue origin to construct regional methylation features, which facilitates a more convenient and efficient expression of tumor-related features. Furthermore, the present invention clusters and matches the target methylation template and its associated model based on differences in CpG site information, which helps to correlate the construction method of regional methylation features with the differences in CpG site information. Therefore, it simplifies related algorithms and saves computational resources required for prediction while comprehensively reflecting the influence of differentiated CpG sites on methylation features. Furthermore, by summarizing the methylation status information of fragments to generate regional methylation features, this invention can obtain richer information, avoiding the disadvantages of limited information and false positives provided by a single cfDNA fragment. Moreover, based on the generated regional methylation features, and through a first classification model trained on multiple samples, this invention generates predictive results indicating the association between the test sample and one or more tumors. Therefore, this invention can conveniently predict the association between the test sample and one or more tumors. Thus, this invention can conveniently and accurately predict the tumor formation risk of the test sample with almost no harm to the human body.
[0044] Figure 1 illustrates a schematic diagram of a system 100 for implementing a method for predicting the association between a test sample and a tumor according to an embodiment of the present invention. As shown in Figure 1, the system 100 includes a computing device 110, a sequencing device 130, a server 140, and a network 150. In some embodiments, the computing device 110, the sequencing device 130, and the server 140 interact with each other via the network 150.
[0045] Regarding sequencing device 130, it is used, for example, for sample library preparation and sequencing of a target library to generate methylation detection data for a test sample or multiple tissue samples or blood samples. Sample library preparation can be performed using various methods, and in some embodiments, it is used, for example but not limited to, the brELSA™ method (Burning Rock Biotech, Guangzhou, China). The sample library preparation method includes, for example, the following steps: first, DNA extraction and purification; then, sodium bisulfite treatment; amplification of single-stranded DNA using DNA polymerase; subsequently, capturing specific cancer-related methylation variant high-incidence regions using oligonucleotide probes; and quantification of the target library using real-time polymerase chain reaction (PCR). Sequencing of the target library can be performed using various sequencers, with an average sequencing depth of, for example, 1000X. Sequencing can be performed, for example but not limited to, using the Illumina NovaSeq 6000 sequencer. It should be understood that the above methods for preparing sample libraries and sequencing target libraries are merely exemplary, and other methods can also be used for preparing sample libraries and sequencing target libraries.
[0046] Regarding the computing device 110, it can be used to construct a classification model for tumor prediction, and to predict the association between a test sample and one or more tumors based on methylation detection data of the acquired test sample. In some embodiments, the computing device 110 can also be used to predict the tumor tissue origin of the test sample based on the methylation detection data of the acquired test sample.
[0047] In some embodiments, the computing device 110 may have one or more processing units, including dedicated processing units such as GPUs, FPGAs, and ASICs, and general-purpose processing units such as CPUs. One or more virtual machines may also run on each computing device.
[0048] Taking an embodiment where the computing device 110 is used to predict the association between a sample to be tested and a tumor as an example, the computing device 110 includes, for example, a methylation detection data acquisition unit 112, a target methylation level template determination unit 114, a regional methylation feature generation unit 116, and a prediction result generation unit 118. The methylation detection data acquisition unit 112, the target methylation level template determination unit 114, the regional methylation feature generation unit 116, and the prediction result generation unit 118 can be configured on one or more computing devices 110.
[0049] Regarding the methylation detection data acquisition unit 112, it is used to acquire methylation detection data about the sample to be tested, wherein the methylation detection data at least indicates the CpG site information covered by the cell-free DNA (cfDNA) fragment and the methylation status data of each CpG site.
[0050] Regarding the target methylation level template determination unit 114, it is used to determine the target methylation level template associated with the cfDNA fragment among multiple methylation level templates based on CpG site information.
[0051] Regarding the regional methylation feature generation unit 116, it is used to generate regional methylation features based on multiple methylation state models associated with a target methylation level template and methylation state data of CpG sites covered by cfDNA fragments, each methylation state model being associated with a tissue or blood.
[0052] Regarding the prediction result generation unit 118, it is used to generate prediction results indicating the association between the test sample and one or more tumors based on the generated regional methylation features and via a first classification model trained on the sample data.
[0053] The following description, in conjunction with Figures 2, 3A, 3B, and 4, describes a method for predicting the association between a test sample and a tumor according to an embodiment of the present invention. Figure 2 shows a flowchart of a method 200 for predicting the association between a test sample and a tumor according to an embodiment of the present invention. It should be understood that method 200 can be performed, for example, at the electronic device 1000 described in Figure 10. It can also be performed at the computing device 110 described in Figure 1. It should be understood that method 200 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the invention is not limited in this respect.
[0054] At step 202, computing device 110 acquires methylation detection data about the sample to be tested, the methylation detection data including at least information on CpG sites covered by cell-free DNA (cfDNA) fragments and methylation status data of each CpG site.
[0055] Regarding the sample to be tested, it is, for example, a tissue sample slice of the object to be tested.
[0056] The methylation detection data for the test sample is, for example, acquired by computing device 110 from sequencing device 130. The methylation detection data for the test sample is generated, for example, through the following steps: first, DNA extraction and purification of the test sample; then, sodium bisulfite treatment of the extracted and purified DNA; amplification of single-stranded DNA using DNA polymerase; subsequently, capture of specific cancer-related high-incidence regions of methylation variants using oligonucleotide probes; quantification of the target library by real-time PCR; and sequencing by a sequencer to generate the methylation detection data for the test sample. It should be understood that DNA methylation is characterized by the addition of a methyl group (-CH3) to the 5th carbon atom of the cytosine ring (5-methylcytosine; 5mC) under the action of DNA methyltransferases, and the covalent addition of the methyl group typically occurs in the cytosine of a CpG dinucleotide. Given that the levels of 5-methylcytosine and 5-hydroxymethylcytosine in DNA play an important role in tumor occurrence and development, and that the methylation characteristics of tumor cells differ significantly from those of normal cells, and that there is a high degree of consistency in DNA methylation characteristics between cfDNA fragments and their tissue-derived genomic DNA, this invention facilitates the convenient extraction of feature information for tumor status prediction from the methylation detection data of the test sample by obtaining the methylation status data of the CpG sites covered by the cfDNA fragments of the test sample.
[0057] At step 204, the computing device 110 determines the target methylation level template associated with the cfDNA fragment among multiple methylation level templates based on CpG site information.
[0058] Each methylation level template is determined based on a series of consecutive CpG sites.
[0059] A method for determining the target methylation level template includes, for example, determining whether the CpG site information contained in each cfDNA fragment data in the methylation detection data of the test sample matches any one of a plurality of methylation level templates (the plurality of methylation level templates are, for example, all methylation level templates that meet predetermined template conditions); and if it is determined that the CpG site information contained in the current cfDNA fragment data in the test result of the test sample matches the current methylation level template among the plurality of methylation level templates (for example, its CpG sites are exactly the same as a certain methylation level template), the current methylation level template is determined as the target methylation level template associated with the current cfDNA fragment data. The above method is used to determine the matching target methylation level template for each cfDNA fragment data. Experiments show that two types of cfDNA fragments ultimately cannot find a matching methylation level template: one type contains fewer than 3 CpG sites; the other type has a fragment length exceeding 200 bp.
[0060] At step 206, computing device 110 generates regional methylation features based on multiple methylation state models associated with the target methylation level template and methylation state data of CpG sites covered by cfDNA fragments, each methylation state model being associated with a tissue or blood.
[0061] In some embodiments, the regional methylation feature includes, for example, T-1 dimensions (T being a natural number greater than or equal to 2). Each dimension corresponds to a regional methylation feature value for a type of tumor tissue. The t-th dimension represents the regional methylation feature value corresponding to the t=1,...,(T-1)-th type of tumor tissue.
[0062] A method for generating regional methylation features includes, for example, the following: a computing device 110 substitutes methylation state data of CpG sites covered by a cfDNA fragment into multiple methylation state models associated with a determined target methylation level template to calculate a posterior probability for each methylation state model; determines a methylation feature value associated with the target methylation level template based on the calculated posterior probability of each methylation state model; and generates regional methylation features based on the determined methylation feature value associated with the target methylation level template.
[0063] It should be understood that a single cfDNA fragment provides limited information and may provide false positive signals. Therefore, this invention summarizes the fragment information within adjacent methylation state models to generate regional methylation features, thereby obtaining richer and more accurate information.
[0064] The following describes the algorithm for calculating the posterior probability of each methylation state model, using Equation (1). Equation (1) shows the functional expression for establishing the methylation state model of the t-th tissue or blood on the j-th template.
[0065] In formula (1) above, t = 1, ..., T represents the source of the t-th tissue (or blood). x represents the methylation status data of the CpG sites covered by the cfDNA fragment. l This represents the methylation status data of the l-th CpG site covered by the cfDNA fragment. t (x) represents the posterior probability. k = 1, ..., K represents the k-th component included in the density function of the methylation state model associated with the j-th methylation level template (the density function of the methylation state model includes, for example, a total of K components). l represents the l-th CpG site. and The model parameters included in the methylation state model are represented, where, The coefficient representing the k-th component included in the methylation state model. Let represent the probability parameter of the Bernoulli distribution at the l-th site in the k-th component.
[0066] The specific methods for generating regional methylation features can include various approaches. In some embodiments, the specific method for generating regional methylation features includes: the computing device 110 converting the posterior probability of each methylation state model associated with the target methylation level template into a probability ratio; and performing a logarithmic summation of the converted probability ratios for all methylation state models associated with the target methylation level template within the region, in order to generate regional methylation features. The algorithm for converting the posterior probability of each model into a probability ratio is explained below with reference to formula (2).
[0067] In the above formula (2), S t (m) (x) represents the transformed probability ratio of the source of the t-th tissue (or blood). f T (x) represents the posterior probability of the T-th model, i.e., the model based on blood from healthy individuals. f t (x) represents the posterior probability of the model before model transformation for the source of the t-th tissue (or blood). It should be understood that by converting the calculated posterior probability of each methylation state model into a probability ratio, this invention can standardize the calculated posterior probability results of different methylation state models.
[0068] The following uses formula (3) to illustrate the algorithm for generating regional methylation features.
[0069] In the above formula (3), S t (x) represents the regional methylation characteristic. t represents the tissue type, t = 1, ..., (T-1). (T-1) represents the source of (T-1) types of tissue (or blood). S t (m) (x) represents the transformed probability ratio of the methylation state model for the t-th tissue (or blood) source. M represents the total number of signal modules included in the region. m represents the m-th signal module in the region. It should be noted that a given region may include one or more different signal modules, each containing multiple consecutive CpG sites that are not entirely consistent. This invention first calculates the methylation feature value of each signal module (this methylation feature value indicates the similarity level of the signal distribution of the test sample and known tissue types or healthy blood on that signal module); then, the probability ratios of the M signal modules included in a region are logarithmically summed to obtain the overall methylation feature value used to indicate the level of that region. It should be understood that for each test sample, a (T-1)-dimensional feature vector indicating the methylation features of a region can be obtained within a region.
[0070] Suppose that within a region, the proportion of signals originating from the t-th tissue (or blood) is α. t If the Tth model is a healthy human blood model, then the likelihood function of the sample data of all cfDNA fragments in this region can be expressed as the following formulas (4) and (5). It should be understood that the likelihood function is a function of the parameters in the statistical model, representing the likelihood of the model parameters. L(x i |α t f t f T )=α t f t (x i )+(1-α t )f T (x i (5)
[0071] In formulas (4) and (5) above, α t ∈[0,1],α t This represents the proportion of signal originating from tissue t (or blood). t = 1, ..., (T-1). It can be estimated using a maximum likelihood estimation algorithm, yielding {α}. t} represents a (T-1) dimensional feature vector within a region. L(x) i |α t f t f T) represents the observational data of all cfDNA fragments within the region. i The likelihood function f. t f represents the posterior probability value of the test sample for tissue type t on that signal module. T f represents the posterior probability of the test sample on a single signal module relative to the blood T of a healthy person on that signal module. t (x i () represents the methylation status of CpG sites on observed cfDNA fragments. i The calculated posterior probability value for model t with respect to organization type. T (x i () represents the methylation status of CpG sites on observed cfDNA fragments. i The calculated posterior probability value for the T-model of blood in healthy individuals.
[0072] At step 208, the computing device 110 generates a prediction result indicating the association between the test sample and one or more tumors based on the generated regional methylation features via a first classification model trained on the sample data.
[0073] A method for generating a prediction result indicating the association between a test sample and one or more tumors includes, for example, a computing device 110 screening for regional methylation features based on screened differentially methylated regions to generate differentially methylated region features; and extracting features of the differentially methylated region features via a first classification model to generate a prediction result indicating the association between a test sample and one or more tumors.
[0074] The method for screening methylation regions that differ for each type of tissue or blood will be described in detail below with reference to Figure 5, and will not be repeated here.
[0075] It should be understood that if the methylation state model f of the t-th tissue (or blood) t The methylation state model f of type T (i.e., blood from healthy individuals) T If the components are very close to each other across most of the modules in the entire region, then the signal proportion α of the t-th tissue (or blood) is... t It can be difficult to estimate accurately, or even impossible to estimate. Therefore, in some embodiments, the present invention will screen for differential methylation regions for each type of tissue or blood, and then accurately predict the association with tumor based on the feature values of each sample in the differential regions.
[0076] The first classification model is, for example, constructed based on a machine learning model. This first classification model is constructed based on typical ElasticNet models, gradient boosting models, random forest models, or support vector machine models. For example, this invention can use an ElasticNet model to construct the first classification model. An ElasticNet model is a linear regression model trained using L1 and L2 norms as prior regularization terms. By using an ElasticNet model to construct the first classification model, this invention can effectively extract and learn features even for sparse sample data with only a small number of non-zero parameters.
[0077] The input data for the first classification model may be, for example, regional methylation features. The output data for the classification model may be, for example, a prediction of the association between the test sample and one or more tumors.
[0078] In some embodiments, the first classification model can be used for early detection prediction of a single cancer type. The input data of the first classification model is, for example, a one-dimensional feature vector indicating the methylation characteristics of differentially expressed regions, and the prediction result of the first classification model indicates the association between the test sample and a type of tumor. Thus, the present invention can accurately and conveniently achieve early detection prediction of a single cancer type. Specifically, the sample data used to train the first classification model is generated by: sorting the distances between each region and the methylation distribution data of a predetermined type of cancer tissue and healthy blood, so as to determine the regions with the largest distances as differentially expressed methylation regions corresponding to the predetermined type of cancer tissue; and generating sample data for training the first classification model based on the feature vectors of the methylation characteristics on the differentially expressed methylation regions that correspond to the dimension of the predetermined type of cancer tissue. Thus, the first classification model trained with the above sample data can be used to predict the association between the test sample and a single tumor of a predetermined type. For example, to predict the association between the test sample and a lung cancer tumor, firstly, the t-th type of tissue is defined as a different lung cancer tumor tissue, and the T-th type is defined as non-cancer blood (i.e., healthy blood). Then, the computing device 110 sorts the distances D(t,T) between the methylation distribution data of the t-th tissue and healthy blood for each region, and identifies the 100 regions with the largest distances D(t,T) as differentially methylated regions. Subsequently, based on the identified 100 differentially methylated regions, the computing device determines the methylation features of the differentially methylated regions associated with lung cancer tumor tissue, for example, based on the t-th dimension (e.g., α) of the methylation features on the differentially methylated regions. t () is used as input data for the first classification model.
[0079] Regarding the training samples, for example, cfDNA fragment samples from lung cancer tumor tissue and healthy blood samples from non-cancer areas, differential methylation features of the training samples are constructed; and Y = 1 / 0 ("1" indicates lung cancer, "0" indicates non-cancer) is used as the response variable (label) for the training samples. The differential methylation features of the training samples are used to train a first classification model constructed based on a machine learning model, and the parameters of the first classification model are adjusted until the difference between the predicted value output by the first classification model and the true value indicated by the response variable meets a predetermined condition.
[0080] To verify the effectiveness of Method 200 in early detection prediction for a single cancer type, this invention uses a set of lung cancer tissue samples and blood samples from healthy individuals to construct a first classification model for lung cancer versus healthy individuals. The first classification model is, for example, based on a support vector machine (SVM) model. For instance, firstly, sequencing data from 45 lung cancer tissue samples and 52 healthy individuals' blood samples are used to construct a model baseline; then, methylation feature values are calculated from blood samples of 82 lung cancer patients and 336 healthy individuals using the model baseline; the 82 lung cancer blood samples and 336 healthy individuals' blood samples are randomly divided into a training set (50 lung cancer blood samples and 194 healthy individuals' blood samples) and a test set (32 lung cancer blood samples and 142 healthy individuals' blood samples); SVM is used to model the training set, and predictions are made for the test set samples. The receiver operating characteristic (ROC) curve for the training set is shown in Figure 3A. The ROC curve for the test set is shown in Figure 3B. As shown in Figures 3A and 3B, Method 200 can accurately achieve early detection and prediction of single cancer types.
[0081] In other embodiments, the first classification model can be used for early detection prediction of multiple cancer types. The input data for the first classification model is multi-dimensional feature vectors representing the methylation features of differentially methylated regions. The prediction results of the first classification model indicate the association between the tested sample and multiple tumors. This invention can perform early detection prediction of multiple cancer types based on the proportion of abnormal DNA from different tissue sources in regions with significant methylation differences. Specifically, the sample data used to train the first classification model is generated by: sorting the distances between each region and the methylation distribution data of each cancer tissue and healthy blood to determine a predetermined number of regions with the largest distances as differentially methylated regions corresponding to each cancer tissue; and merging the feature vectors of the methylation features in the differentially methylated regions with the corresponding dimension of each cancer tissue to generate the sample data used to train the first classification model. Thus, the trained first classification model can be used to predict the association between the tested sample and multiple lung cancer tumors.
[0082] For example, to predict the association between a test sample and various tumors, the computing device 110 first defines tissue type t (t = 1, ..., (T-1)) as a distinct cancerous tissue and tissue type T as non-cancer blood. Then, the computing device 110 sorts the distances D(t,T) between the methylation distribution data of each region for each distinct tissue (t) and non-cancer blood (T) to determine the 100 regions with the largest distances D(t,T) as differentially methylated regions. For example, for tissue type 1, the 100 regions with the largest distances D(1,T) are determined as differentially methylated regions R1; for tissue type 2, the 100 regions with the largest distances D(2,T) are determined as differentially methylated regions R2, and so on, until for tissue type T-1, the 100 regions with the largest distances D(T-1,T) are determined as differentially methylated regions R1. T-1 Then, the computing device 110 uses the feature vector (e.g., α1) of the first dimension of the methylation features on the differentially methylated region R1 and the feature vector (e.g., α1) of the second dimension of the methylation features on the differentially methylated region R2 to the differentially methylated region R... T-1 The second dimension of the eigenvector of the epimethylation feature (e.g., α) T-1 Merge {(R1, α1), ..., (R)}. For example, merge {(R1, α1), ..., (R)}. T-1 α T-1 The data are combined to serve as input data for the classification model.
[0083] Regarding training samples, for example, cfDNA fragment samples from multiple cancer types and non-cancer blood samples are used to construct differentially expressed methylation features of the training samples; and Y = 1 / 0 ("1" indicates lung cancer, "0" indicates non-cancer) is used as the response variable (label) for the training samples. The differentially expressed methylation features of the training samples are used to train a classification model based on a machine learning model, adjusting the parameters of the classification model until the difference between the output of the classification model and the true values indicated by the response variable meets predetermined conditions.
[0084] To verify the effectiveness of Method 200 in early detection prediction for multiple cancer types, this invention used methylation data from 82 lung cancer tissue samples, 78 colon cancer tissue samples, 80 liver cancer tissue samples, 70 ovarian cancer tissue samples, 72 pancreatic cancer tissue samples, 54 esophageal cancer tissue samples, and 336 healthy human samples as input data to construct a classification model. The prediction results of the trained classification model are shown in Figure 4. Figure 4 illustrates a schematic diagram of the distinguishing effect of methylation features constructed according to an embodiment of the present invention and according to a method of the prior art on cancer / health. As shown in Figure 4, the horizontal axis indicates different cancer types. The vertical axis indicates the AUC (Area Under Curve) of methylation features against cancer / health (i.e., the area under the ROC curve enclosed by the coordinate axis). In a set of three graphs corresponding to each cancer type, the left graph indicates the AUC of differential methylation features against cancer / health predicted by the method according to an embodiment of the present invention in different genomic regions, and the middle graph indicates the AUC of methylation features against cancer / health constructed using conventional average methylation levels. The graph on the right indicates the AUC of predictions using conventional methylation region scoring methods. For example, label 410 indicates the AUC of methylation features constructed according to the method of the present invention for lung cancer / health. Label 412 indicates the AUC of methylation features constructed based on average methylation levels for lung cancer / health. Label 414 indicates the AUC of methylation features constructed based on methylation region scoring for lung cancer / health. It should be understood that the closer the AUC is to 1.0, the higher the accuracy of the detection method. As shown in Figure 4, for any one or more types of cancer, including lung cancer, colon cancer, liver cancer, ovarian cancer, pancreatic cancer, and esophageal cancer, the method of the present invention demonstrates significantly higher accuracy in distinguishing between cancer and health based on differentially expressed methylation features in different genomic regions than other methods in the prior art.
[0085] In the above scheme, a target methylation level template associated with the cfDNA fragment is determined by using CpG site information of the cfDNA fragment indicated in the methylation detection data of the sample to be tested; and a regional methylation feature is generated based on multiple methylation state models associated with the target methylation level template and methylation state data of the CpG sites covered by the cfDNA fragment; and based on the generated regional methylation feature, a prediction result indicating the association between the sample to be tested and one or more tumors is generated through a classification model trained on multiple samples. This invention can conveniently and accurately predict the tumor formation risk of the sample to be tested, while causing almost no harm to the human body.
[0086] The following description, in conjunction with FIG5, describes a method for screening methylated regions for each tissue or blood screening difference according to an embodiment of the present invention. FIG5 shows a flowchart of a method 500 for screening methylated regions for each tissue or blood screening difference according to an embodiment of the present invention. It should be understood that method 500 can be performed, for example, at the electronic device 1000 described in FIG10. It can also be performed at the computing device 110 described in FIG1. It should be understood that method 500 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the invention is not limited in this respect.
[0087] At step 502, the computing device 110 determines the distance between each of the plurality of methylation level templates for the methylation distribution of different tissues or blood.
[0088] Methods for calculating the distance between methylation distributions of different tissues or blood include, but are not limited to, methods based on, but not limited to, relative entropy, Kullback-Leibler divergence (or "KL divergence"), Jensen-Shannon divergence (or "JS divergence"), and Wasserstein distance. For example, in some embodiments, using the Wasserstein distance to calculate the distance between methylation distributions of different tissues or blood allows the degree of difference between the two distributions to be reflected even if the support sets of the methylation distributions of two different tissues or blood do not overlap or overlap very little.
[0089] The following example, using formula (6), illustrates an algorithm for calculating the distance between methylation distributions in different tissues or blood based on KL divergence.
[0090] In the above formula (6), and These represent the methylation distribution data for tissue (or blood) at t1 and t2, respectively. D KL (t1, t2) represents the distance between the methylation distributions of t1 and t2 in different tissues or blood. The summation operation requires taking all r values. i 2 L The possible values are given. It should be understood that the greater the difference between the baseline models of methylation distribution for tissue (or blood) t1 and t2, the greater the corresponding distance D between the methylation distributions of t1 and t2 for different tissues or blood. KL (t1,t2) will also be larger.
[0091] At step 504, the computing device 110 calculates the distance between the methylation distribution data of each region for different tissues or blood based on the distance between the methylation distributions of each methylation level template for different tissues or blood.
[0092] For example, computing device 110 can display the methylation level templates within a region and the methylation distribution differences for different tissues or blood (i.e., the distance D between methylation distributions for different tissues or blood). KL The average of (t1, t2) is taken as the distance between the methylation distribution data of that region in two different tissues (or blood). The following example, in conjunction with formula (7), illustrates an algorithm for calculating the distance between the methylation distribution data of each region in different tissues or blood.
[0093] In formula (7) above, D(t1,t2) represents the distance between the methylation distribution data of the region in different tissues (or blood) at t1 and t2. m (t1, t2) represents the distance between the m-th template in the region and t1, t2 in different tissues (or blood).
[0094] At step 506, the computing device 110 sorts multiple regions based on the distances between each region and the methylation distribution data of different tissues or blood. For example, the computing device 110 sorts all regions according to D(t1,t2) calculated at step 504.
[0095] At step 508, the computing device 110, based on the sorting results, identifies methylation regions where there are significant differences in methylation between different tissues or blood. For example, the computing device 110 selects a large number of regions based on the sorting results (e.g., sorting different regions in descending order of distance and selecting the top 100 regions) as regions with significant differences in methylation between tissue (or blood) t1 and t2.
[0096] It should be understood that to determine whether a cfDNA fragment originates from tumor cells, t1 can represent any type of tumor tissue, and t2 can represent blood from a non-tumor patient. Conversely, to determine the origin of a ctDNA fragment (e.g., identifying the primary tumor site), t1 and t2 can represent different tumor tissues.
[0097] By employing the above methods, the present invention can avoid the situation where the methylation distribution data of the healthy sample source model and the methylation distribution data of the tumor tissue (or blood) source model are very similar in most modules of the region, which makes the prediction results difficult to estimate, thereby improving the accuracy of the prediction results.
[0098] The following description, in conjunction with FIG6, describes a method for determining a plurality of associated methylation state models and corresponding model parameters for each methylation level template according to an embodiment of the present invention. FIG6 shows a flowchart of a method 600 for determining a plurality of associated methylation state models and corresponding model parameters for each methylation level template according to an embodiment of the present invention. It should be understood that method 600 can be performed, for example, at the electronic device 1000 described in FIG10. It can also be performed at the computing device 110 described in FIG1. It should be understood that method 600 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the invention is not limited in this respect.
[0099] At step 602, computing device 110 establishes multiple methylation state models associated with each methylation level template based on the Bernoulli mixing model.
[0100] It should be understood that, since the methylation status data at each CpG site on the cfDNA fragment is a binary variable (i.e., methylated / unmethylated), it follows a Bernoulli distribution (BMM) model. Considering the correlation between methylation status data at adjacent CpG sites, this invention employs a Bernoulli distribution model to fit the methylation status data on the cfDNA fragment.
[0101] The following describes the algorithm for establishing multiple models associated with each methylation level template, using formulas (8) to (10).
[0102] In formula (8) above, L represents the number of CpG sites contained in the j-th methylation level template. K represents the K components (polynomial components) of the density function of the model associated with the j-th methylation level template. f(r i |θ,p) represents the density function of the model associated with the j-th methylation level template. i =(r i1 ,…,r iL ) represents the methylation status data at each CpG site on the i-th cfDNA fragment on the j-th template. il =0 / 1 represents the methylation status of the l-th CpG site on the i-th cfDNA fragment as unmethylated / methylated. θ = (θ1,...,θ K ) and p = (p kl ) K×L The parameters represent the model parameters, where θ represents the coefficients of the K components, and p klLet θ represent the probability parameter of the Bernoulli distribution at the l-th site in the k-th component. 0 < θ k ,p kl <1 (9)
[0103] In the above formulas (9) and (10), θ k The coefficient of the k-th component of the density function representing the methylation state model. kl Let represent the probability parameter of the Bernoulli distribution at the l-th site in the k-th component.
[0104] At step 604, computing device 110 uses a maximum likelihood estimation algorithm to estimate model parameters for each methylation state model associated with a methylation level template. For example, model parameters for a methylation state model (e.g., a Bernoulli mixture) are estimated using methylation detection data attributable to the j-th template in the t-th tissue or blood sample.
[0105] The following describes the algorithm for estimating the model parameters of each methylation state model associated with each methylation level template, using formulas (11) to (14). l(θ,p)=∑ i l i (θ,p) (11)
[0106] In the above formulas (11) to (14), l(θ, p) represents the log-likelihood function on the j-th methylation level template. and They represent the logarithmic function l i (θ,p) pairs and The partial derivatives of . The k-th density function representing the model * The coefficients of each component. Indicates the kth * The 1st component * The probability parameters of the Bernoulli distribution at each site. L represents the number of CpG sites contained in the template at the j-th methylation level. K represents the K components (polynomial components) of the density function of the model associated with the j-th methylation level template. * =1,...,(K-1),l * =1, ..., L. f(r) i |θ, p) represents the density function of the methylation state model associated with the j-th methylation level template.
[0107] Research has shown that the partial derivatives of maximum likelihood can be expressed explicitly. Therefore, this invention uses quasi-Newton methods to solve this constrained maximum likelihood estimation problem.
[0108] The number of components K in the model is determined, for example, by the Akaike information criterion (AIC criterion).
[0109] The following uses formula (15) to explain the algorithm for determining the number of components in the methylation state model. AIC = (2K - 2M) / n (15)
[0110] In the above formula (15), AIC represents the Akaike Information Criterion value. K represents the number of parameters in the fitted model. M represents the log-likelihood value. N represents the number of observations. It should be understood that a small K means a simple model. A large M means a precise model. Therefore, determining the number of model components K through the Akaike Information Criterion is beneficial for balancing model complexity and residuals.
[0111] At step 606, the computing device 110 determines a baseline model of cfDNA fragment methylation status associated with each methylation level template based on the model parameters of each estimated methylation state model, in order to determine multiple methylation state models associated with the target methylation level template.
[0112] For example, based on the parameters of the model parameters for each determined methylation state model, the computing device 110 can obtain a baseline model of the cfDNA methylation state for various tissues or blood on each methylation level template. The algorithm for the cfDNA methylation state baseline model is explained below with reference to formula (16).
[0113] In the above formula (16), f(r) i ) represents the methylation status data r at each CpG site on the i-th cfDNA fragment for the j-th methylation level template. i The function of the baseline model of cfDNA methylation status. and The parameters represent the baseline model of the determined methylation state.
[0114] By employing the above methods, this invention can establish methylation state models for different types of tissues or blood on a certain methylation level template, and then use the detection data of a certain tissue or blood sample belonging to the methylation level template to estimate the model parameters, thereby quickly establishing multiple associated methylation state models and corresponding model parameters for each methylation level template.
[0115] The following description, in conjunction with FIG7, describes a method for predicting and determining the origin of tumor tissue in a test sample according to an embodiment of the present invention. FIG7 shows a flowchart of a method 700 for predicting and determining the origin of tumor tissue in a test sample according to an embodiment of the present invention. It should be understood that method 700 may be performed, for example, at the electronic device 1000 described in FIG10. It may also be performed at the computing device 110 described in FIG1. It should be understood that method 700 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the invention is not limited in this respect.
[0116] At step 702, computing device 110 acquires methylation detection data about the sample to be tested, the methylation detection data including at least information on CpG sites covered by cell-free DNA (cfDNA) fragments and methylation status data of each CpG site.
[0117] At step 704, the computing device 110 determines a target methylation level template associated with the cfDNA fragment among multiple methylation level templates based on CpG site information, so as to generate regional methylation features based at least on the methylation status data of the CpG sites covered by the target methylation level template and the cfDNA fragment.
[0118] At step 706, the computing device 110 filters the regional methylation features based on the generated regional methylation features and the differentially methylated regions in order to generate differentially methylated regional features.
[0119] For example, to predict and determine the tumor tissue origin of a sample to be tested, the computing device 110 first defines tissue type t (t = 1, ..., (T-1)) as different cancer tissues and type T as non-cancer blood. Then, the computing device 110 sorts the distance D(t1,t2) between the methylation distribution data of each pair of different cancer tissues and filters for differentially methylated regions. For example, for tissues 1 and 2, the 100 regions with the largest D(1,2) are identified as the differentially methylated regions R between tissues 1 and 2. 12 For tissues t1 and t2, the 100 regions with the largest D(t1,t2) are identified as the methylation regions representing the differences between tissues t1 and t2. And so on. Then, computing device 110 uses the differential methylation regions R of tissues 1 and 2. 12 The first and second dimensions of methylation characteristics, such as α1 and α2; the methylation regions differing between the two tissues, t1 and t2. The t1 and t2 dimensions of the methylation features on the surface, for example And so on; The combined data are used as input to the second classification model based on the training samples, for example, using cfDNA fragment samples from multiple cancer types of tumor tissues, to construct differentially expressed methylation features about the training samples.
[0120] At step 708, the computing device 110 generates a prediction result indicating the tumor tissue origin of the test sample based on the differential methylation features via a second classification model trained on the sample data.
[0121] For example, computing device 110 uses cancer type as a response variable (label) to the training samples. It trains a second classification model based on a machine learning model by using the methylation features of the differential regions of the training samples, adjusting the parameters of the second classification model until the difference between the true values indicated by the response variable output by the second classification model satisfies a predetermined condition (e.g., minimizing the loss function).
[0122] In the above scheme, the present invention can accurately and conveniently predict the origin of tumor tissue in the sample to be tested.
[0123] The following description, in conjunction with FIG8, describes a method for constructing a classification model for tumor prediction according to an embodiment of the present invention. FIG8 shows a flowchart of a method 800 for constructing a classification model for tumor prediction. It should be understood that method 800 can be performed, for example, at the electronic device 1000 described in FIG10. It can also be performed at the computing device 110 described in FIG1 according to an embodiment of the present invention. It should be understood that method 800 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the invention is not limited in this respect.
[0124] At step 802, computing device 110 acquires methylation detection data for various tissue and blood samples, the methylation detection data including at least information on CpG sites covered by cell-free DNA (cfDNA) fragments in the blood and methylation status data for each CpG site.
[0125] Regarding tissue samples, they undergo pathological evaluation to ensure that the proportion of tumor cells in the tissue sample sections exceeds 50%.
[0126] In some embodiments, negative and positive standards can be used to simulate tissue samples with different levels of tumor components (ctDNA) in cell-free blood DNA (cfDNA). For example, using NA24385 as a negative control and H2228 as a positive standard, H2228 is incorporated into the negative control at different ratio gradients to simulate cfDNA samples with different ctDNA levels. In some embodiments, the included gradients are 5%, 0.5%, 0.05%, 0.01%, and 0.001%. Figure 9 schematically illustrates the content gradient results of simulated tissue sample data according to an embodiment of the present invention.
[0127] At step 804, the computing device 110 determines multiple methylation level templates based on multiple consecutive CpG sites within the probe region, so as to determine multiple associated methylation state models and corresponding model parameters for each methylation level template, each methylation state model being associated with a tissue or blood.
[0128] In some embodiments, the method for determining multiple methylation level templates includes, for example, determining within a probe region, whether a current group of consecutive CpG sites satisfies predetermined template conditions, the predetermined template conditions including: the number of consecutive CpG sites is greater than or equal to a predetermined number threshold; the distance between the first CpG site and the last CpG site is less than or equal to a predetermined distance threshold; and the CpG sites are different from those included in a determined methylation level template; in response to determining that the current group of consecutive CpG sites satisfies the predetermined template conditions, determining the current group of consecutive CpG sites as methylation level templates; and in response to determining that the current group of consecutive CpG sites does not satisfy the predetermined template conditions, determining whether the next group of consecutive CpG sites satisfies the predetermined template conditions.
[0129] Regarding the predetermined threshold, it is, for example, but not limited to, 3. It should be understood that if there are fewer than 3 consecutive CpG sites, such reads can be considered to lack methylation information and therefore lack analytical value.
[0130] Regarding the predetermined distance threshold, it is, for example, but not limited to, 200 bp. The reason for setting the predetermined distance threshold to 200 bp is that the length of cfDNA fragments is usually between 150-200 bp, and the possibility of cfDNA fragments exceeding 200 bp in length is relatively small.
[0131] At step 806, the computing device 110 screens for differentially methylated regions for each type of tissue or blood, the differentially methylated regions being associated with significant methylation differences.
[0132] At step 808, computing device 110 generates sample data based on methylation status data of CpG sites covered by the cfDNA fragments of the sample, multiple methylation status models associated with methylation level templates matching the cfDNA fragments of the sample, and differential methylation regions, so as to train a classification model based on the sample data.
[0133] Regarding the classification model, it can be a first classification model used to predict the association between the test sample and one or more tumors, or a second classification model used to predict and determine the origin of the tumor tissue in the test sample.
[0134] A method for training a first classification model for predicting associations among various tumors using sample data includes, for example: sorting the distances between each region and methylation distribution data of a predetermined type of cancer tissue and healthy blood, so as to determine a predetermined number of regions with the largest distances as differentially methylated regions corresponding to the predetermined type of cancer tissue; and generating sample data for training the first classification model based on feature vectors of methylation features on the differentially methylated regions that correspond to the dimension of the predetermined type of cancer tissue.
[0135] A method for using sample data to train a second classification model includes, for example, the following: a computing device 110 sorts the distances between methylation distribution data of two different cancer tissues to select a predetermined number of regions with the largest distances as methylation regions corresponding to the differences between the two different cancer tissues; and merges the feature vectors of the corresponding two dimensions of the methylation features on the methylation regions corresponding to the differences between the two different cancer tissues to generate sample data for training the second classification model. Thereby, the trained second classification model can be used to predict and determine the tumor tissue origin of a test sample.
[0136] By employing the above methods, the present invention can obtain accurate prediction of tumor tissue origin and classification models for early detection of one or more tumors.
[0137] [Correction 17.12.2024 according to Rule 91] FIG10 schematically illustrates a block diagram suitable for implementing an electronic device 1000 of embodiments of the present invention. The electronic device 1000 may be used to implement methods 200, 500 to 800 shown in FIG2, 5 to 8. As shown in FIG10, the electronic device 1000 includes a central processing unit (i.e., CPU 1001), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (i.e., ROM 1002) or loaded from a storage unit 1008 into a random access memory (i.e., RAM 1003). Various programs and data required for the operation of the electronic device 1000 may also be stored in the RAM 1003. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output interface (i.e., I / O interface 1005) is also connected to the bus 1004.
[0138] Multiple components in electronic device 1000 are connected to I / O interface 1005, including: input unit 1006, output unit 1007, and storage unit 1008. CPU 1001 executes the various methods and processes described above, such as methods 200, 500 to 800. For example, in some embodiments, methods 200, 500 to 800 may be implemented as computer software programs stored in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by CPU 1001, one or more operations of methods 200, 500 to 800 described above may be performed. Alternatively, in other embodiments, CPU 1001 may be configured to execute one or more actions of methods 200, 500 to 800 by any other suitable means (e.g., by means of firmware).
[0139] It should be further noted that the present invention can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the present invention.
[0140] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0141] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0142] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.
[0143] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0144] These computer-readable program instructions can be provided to a processor in a voice interaction device, a general-purpose computer, a special-purpose computer, or a processing unit of another programmable data processing device, thereby producing a machine such that, when executed by the processing unit of the computer or other programmable data processing device, these instructions create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing device, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0145] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0147] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0148] The above are merely optional embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting the association between a test sample and a tumor, characterized in that, include: Acquire methylation detection data for the sample to be tested, wherein the methylation detection data indicates at least the CpG site information covered by the cell-free DNA (cfDNA) fragment and the methylation status data of each CpG site; Based on CpG site information, target methylation level templates associated with cfDNA fragments were identified among multiple methylation level templates; Based on multiple methylation state models associated with the target methylation level template and methylation state data of CpG sites covered by cfDNA fragments, regional methylation features are generated, with each methylation state model associated with a tissue or blood. as well as Based on the generated regional methylation features, a first classification model trained on sample data is used to generate predictions indicating the association between the test sample and one or more tumors.
2. The method according to claim 1, characterized in that, Generating predictive results that indicate the association between a test sample and one or more tumors includes: Based on the selected differentially methylated regions, regional methylation features are further refined to generate differentially methylated region features; and Features of methylation characteristics in differentially classified regions are extracted using a first classification model to generate predictions that indicate the association between the test sample and one or more tumors.
3. The method according to claim 2, characterized in that, The methylation features of the differential regions are one-dimensional feature vectors, and the prediction results of the first classification model indicate the association between the sample to be tested and a type of tumor.
4. The method according to claim 2, characterized in that, The methylation features of the differential regions are multidimensional feature vectors, and the prediction results of the first classification model indicate the association between the test sample and various tumors.
5. The method according to claim 3, characterized in that, The sample data used to train the first classification model is generated from the following: The distance between each region and the methylation distribution data of a predetermined type of cancer tissue and healthy blood is sorted in order to determine the regions with the largest distance as the methylation regions corresponding to the differences in cancer tissue of the predetermined type. as well as Based on the feature vectors of the methylation features on the differentially methylated regions that correspond to the dimensions of the predetermined type of cancer tissue, sample data for training the first classification model is generated.
6. The method according to claim 4, characterized in that, The sample data used to train the first classification model is generated from the following: The distance between each region and the methylation distribution data of each cancer tissue and healthy blood is sorted in order to determine the regions with the largest distances as the methylation regions corresponding to the differences in each cancer tissue. as well as The feature vectors of the methylation features on the differentially methylated regions, corresponding to the dimension of each cancer tissue, are merged to generate sample data for training the first classification model.
7. A method for predicting the origin of tumor tissue in a sample to be tested, characterized in that, include: Acquire methylation detection data for the sample to be tested, wherein the methylation detection data indicates at least the CpG site information covered by the cell-free DNA (cfDNA) fragment and the methylation status data of each CpG site; Based on CpG site information, target methylation level templates associated with cfDNA fragments are identified among multiple methylation level templates, so as to generate regional methylation features based at least on the methylation status data of the CpG sites covered by the target methylation level templates and the cfDNA fragments. Based on the generated regional methylation features and differentially methylated regions, the regional methylation features are screened in order to generate differentially methylated regional features. as well as Based on differential methylation features, a second classification model trained on sample data is used to generate a prediction of the tumor tissue origin of the sample to be tested.
8. The method according to claim 1 or 7, characterized in that, Generative region methylation features include: The methylation status data of CpG sites covered by the cfDNA fragment are substituted into multiple methylation status models associated with the template of the determined target methylation level in order to calculate the posterior probability of each methylation status model. Based on the calculated posterior probability of each methylation state model, methylation feature values associated with the target methylation level template are determined; and Regional methylation features are generated based on the methylation feature values associated with the target methylation level template.
9. The method according to claim 8, characterized in that, The methylation features generated based on the determined methylation feature values associated with the target methylation level template include: The posterior probability of each methylation state model associated with the target methylation level template is converted into a probability ratio; and Logarithmic summation is performed on the transformed probability ratios of all methylation state models associated with the target methylation level template within the region to generate regional methylation features.
10. The method according to claim 1 or 7, characterized in that, Target methylation level templates associated with cfDNA fragments were identified across multiple methylation level templates, including: To determine whether the CpG site information contained in each cfDNA fragment in the methylation detection data of the sample matches any one of the multiple methylation level templates; and In response to determining that the CpG site information contained in the current cfDNA fragment data in the test result of the sample matches the current methylation level template among multiple methylation level templates, the current methylation level template is identified as the target methylation level template associated with the current cfDNA fragment data.
11. The method according to claim 1 or 7, characterized in that, Also includes any of the following: Within the probe region, multiple methylation levels are templated based on consecutive CpG sites; and For each type of tissue or blood, differentially methylated regions were screened, and these differentially methylated regions were associated with significant differences in methylation.
12. The method according to claim 7, characterized in that, Sample data is generated from the following: The distances between the methylation distribution data of each two different cancer tissues are sorted in order to select a predetermined number of regions with the largest distances as the methylation regions corresponding to the differences between each two different cancer tissues; as well as The feature vectors of the corresponding two dimensions of methylation features on the methylation regions corresponding to the differences between the two types of cancer tissues are merged to generate sample data for training the second classification model.
13. A method for constructing a classification model for tumor prediction, characterized in that, include: Acquire methylation detection data for various tissue and blood samples, wherein the methylation detection data indicates at least the CpG site information covered by cell-free DNA (cfDNA) fragments in the blood and the methylation status data of each CpG site; Within the probe region, multiple methylation level templates are determined based on a series of consecutive CpG sites, so as to determine multiple associated methylation state models and corresponding model parameters for each methylation level template, with each methylation state model associated with a tissue or blood. For each type of tissue or blood, differentially methylated regions were screened, and these differentially methylated regions were associated with significant differences in methylation. Sample data is generated based on the methylation status data of CpG sites covered by the cfDNA fragments of the sample, multiple methylation status models associated with methylation level templates that match the cfDNA fragments of the sample, and differentially methylated regions, so as to train a classification model based on the sample data.
14. The method according to claim 13, characterized in that, The classification model is any one of the following: A first classification model used to predict the association between a test sample and one or more tumors; or A secondary classification model used to predict and determine the origin of tumor tissue in a test sample.
15. The method according to claim 11 or 13, characterized in that, Multiple methylation level templates were determined based on consecutive CpG sites, including: Within the probe region, determine whether multiple consecutive CpG sites in the current group meet the predetermined template conditions, which include: The number of consecutive CpG sites is greater than or equal to a predetermined number threshold; and The distance between the first CpG site and the last CpG site is less than or equal to a predetermined distance threshold; and Unlike the continuous CpG sites included in the established methylation level template; In response to determining that multiple consecutive CpG sites in the current group meet predetermined template conditions, the multiple consecutive CpG sites in the current group are determined as methylation level templates; and In response to determining that the current consecutive multiple CpG sites do not meet the predetermined template conditions, determine whether the next group of consecutive multiple CpG sites meets the predetermined template conditions.
16. The method according to claim 15, characterized in that, The predetermined bit threshold is 3, and the predetermined distance threshold is 200bp.
17. The method according to claim 11 or 13, characterized in that, For each type of tissue or blood, the differentially methylated regions screened include: Determine the distance between the methylation distributions of each methylation level template in multiple methylation level templates for different tissues or blood. Based on the distance between the methylation distributions of each methylation level template for different tissues or blood, calculate the distance between the methylation distribution data of each region for different tissues or blood. Based on the calculated distances between methylation distribution data of each region for different tissues or blood, multiple regions are sorted; and Based on the sorting results, methylation regions with significant differences in methylation across different tissues or blood were identified.
18. The method according to claim 15, characterized in that, For each methylation level template, multiple associated methylation state models and corresponding model parameters are determined, including: For each methylation level template, based on the Bernoulli mixing model, multiple methylation state models associated with each methylation level template are established; Maximum likelihood estimation was used to estimate the model parameters for each methylation state model associated with the methylation level template; and Based on the model parameters of each estimated methylation state model, a baseline model of cfDNA fragment methylation state associated with each methylation level template is determined to identify multiple methylation state models associated with the target methylation level template.
19. A computing device, characterized in that, include: At least one processing unit; At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the device to perform the steps of the method according to any one of claims 1 to 18.
20. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program that, when executed by a machine, implements the method according to any one of claims 1 to 18.
21. A computer program product, characterized in that, The computer program product includes instructions that, when executed by a machine, implement the method according to any one of claims 1 to 18.
Citation Information
Patent Citations
Method and system for evaluating tumor formation risk and tumor tissue source
CN115132273A
Method, device, medium and program for predicting incidence relation between sample and tumor
CN118522357A
Tumor fraction estimation using methylation variants
US20230272486A1
Detecting cancer, cancer tissue of origin, and / or a cancer cell type
WO2020163410A1