Marker screening method based on methylation sequencing, cancer detection method and device

CN116312739BActive Publication Date: 2026-09-25BIOCHAIN BEIJING SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111563505.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-20
Publication Date
2026-09-25
Estimated Expiration
2041-12-20

AI Technical Summary

Technical Problem

[0009](1)极度依赖marker的选择,而marker的选择又是一个高度依赖经验的过程,而且由于marker的选择与最终的分类过程并非强偶联,上游选择出的marker在最终的独立测试集中常常没有表现出期望的分类性能;

Benefits of technology

[0103]本发明所述的方法采用健康样本血浆、正常组织和肿瘤组织等三种组织来源的混合模型假设,而非传统的健康样本血浆和肿瘤组织的二组分混合假设,理论上这种假设不仅能对待检测样本给出健康或者患病的判别,而且对于判别为患病的样本,还能给出是否为癌症的准确区分;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_3
    Figure QLYQS_3
  • Figure QLYQS_15
    Figure QLYQS_15
  • Figure QLYQS_16
    Figure QLYQS_16
Patent Text Reader

Abstract

The application discloses a marker screening method and device based on methylation sequencing and a cancer detection method, and the method comprises the following steps: selecting L candidate markers; calculating the methylation level distribution mode characteristic value of each candidate marker in three groups of samples, i.e., healthy plasma samples, normal tissue samples and tumor tissue samples; performing methylation sequencing on a to-be-detected sample to obtain a methylation sequencing result of the to-be-detected sample; based on the methylation sequencing result, using the methylation level distribution mode characteristic value to calculate the tissue tracing characteristic vector of each marker of the L candidate markers of the to-be-detected sample; based on the tissue tracing characteristic vector and the cancer or health condition of the to-be-detected sample, constructing a linear classification model, and screening Q markers as markers for subsequent cancer detection based on the average prediction accuracy of the linear classification model, so that accurate distinction can be given as to whether it is cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of in vitro cancer detection technology, and in particular to a biomarker screening method based on methylation sequencing, a cancer detection method, and an apparatus thereof. Background Technology

[0002] Diagnosis and prognosis play a crucial role in the prevention and treatment of diseases. Traditionally, various types of surgical biopsies, such as bone marrow or needle biopsies, are performed in clinical settings. However, surgical biopsies are generally not the first choice due to their invasiveness and the potential sampling bias in tumor biopsies. As an alternative to surgical biopsies, researchers and clinicians have been searching for molecular biomarkers for disease diagnosis that can help in disease state detection, subtype classification, progression prediction, and treatment response characterization. Thanks to advances in high-throughput genomics technologies such as microarrays and next-generation sequencing, the scope of molecular biomarker discovery has expanded dramatically over the past two decades.

[0003] Abnormal DNA methylation patterns are highly promising molecular biomarkers for disease diagnosis. Cancer cells typically exhibit aberrant DNA methylation patterns, such as hypermethylation in tumor suppressor gene promoter regions and widespread hypomethylation in intergenic regions. Therefore, DNA methylation is an ideal target for cancer diagnosis in clinical practice. Hypermethylated / hypomethylated tumor DNA fragments can be released into the bloodstream through apoptosis or necrosis, becoming part of circulating cell-free DNA (cfDNA) in plasma. The non-invasive nature of cfDNA methylation analysis makes it a promising strategy for general cancer screening.

[0004] Cell-free DNA (cfDNA) enters circulation during cell death as part of normal cell renewal or pathological processes. In healthy individuals, most cell-free DNA molecules in plasma originate from blood cells. When an organ or tissue experiences increased cell death, more cfDNA from the affected tissue or organ will be present in circulation. Due to significant tissue-specific differences in DNA methylation, an increase in cfDNA from a specific tissue or organ will alter its normal methylation pattern. Furthermore, changes in physiological state leading to alterations in the methylation pattern of a specific tissue or organ, when entering the bloodstream, will also affect the plasma cfDNA methylation pattern. Therefore, changes in the proportion of cfDNA from a specific tissue, or changes in the methylation pattern of cfDNA from a particular tissue, can lead to abnormalities in the final plasma cell-free DNA methylation pattern. Thus, identifying abnormal methylation patterns in a patient's plasma cell-free DNA can effectively diagnose their disease state and pinpoint the affected tissue or organ.

[0005] For the reasons mentioned above, traditional disease discrimination methods based on methylation generally follow two approaches:

[0006] (1) Identify sites or regions where there are differences in methylation levels between the control group and the disease group, and make predictions directly based on the differences in methylation levels of these sites or regions;

[0007] (2) Based on the differences in DNA methylation patterns from different tissue sources, estimate the proportion of cfDNA from each tissue source, and predict disease status based on the proportion of abnormal tissue sources.

[0008] However, the above-mentioned traditional methods have the following problems:

[0009] (1) It is highly dependent on the selection of markers, which is a process that is highly dependent on experience. Moreover, since the selection of markers is not strongly coupled with the final classification process, the markers selected upstream often do not show the expected classification performance in the final independent test set.

[0010] (2) In the traditional tissue tracing approach, the tissue origin of the patient's plasma cfDNA is often simply assumed to be a mixture of two components: healthy human plasma and tumor tissue. However, since in most cases, tumor tissue exists in the form of a mixture of normal tissue cells and cancer cells, the purity of the tumor is generally between 20% and 80%. This means that the tumor tissue contains tissue-specific information of the corresponding tissue, as well as information on the tumor carcinogenesis of that tissue. Therefore, according to this two-tissue mixture assumption, both carcinogenesis and ordinary lesions of the corresponding tissue may lead to a high proportion of corresponding tissue tracing, resulting in difficulties in distinguishing between carcinogenesis and ordinary lesions. Summary of the Invention

[0011] In view of the above problems, the present invention provides a biomarker screening and cancer detection method and device based on plasma cell-free DNA methylation sequencing. The method adopts a mixed model assumption of three tissue sources: healthy plasma, normal tissue and tumor tissue, rather than the traditional two-component mixed assumption of healthy plasma and tumor tissue. Theoretically, this assumption can not only give a judgment of whether the sample to be tested is healthy or diseased, but also give an accurate distinction of whether the sample is cancerous for the sample judged to be diseased.

[0012] The specific technical solution of this invention is as follows:

[0013] 1. A biomarker screening method based on methylation sequencing, comprising:

[0014] Select L candidate markers;

[0015] For each candidate biomarker, the methylation level distribution pattern characteristic value of its respective methylation level in three groups of samples (healthy plasma sample, normal tissue sample, and tumor tissue sample) was calculated, that is, 3L characteristic values ​​were obtained.

[0016] The methylation sequencing results of the sample to be tested are obtained by performing methylation sequencing on the sample to be tested;

[0017] Based on the methylation sequencing results of the sample to be tested, the tissue origination feature vector of each of the L candidate biomarkers in the sample to be tested is calculated using the methylation level distribution pattern feature value.

[0018] A linear classification model is constructed based on the tissue source feature vector and the cancer or health status of the sample to be tested. Q biomarkers are selected as biomarkers for subsequent cancer detection based on the average prediction accuracy of the linear classification model, where L is greater than or equal to Q and both L and Q are integers greater than or equal to 1.

[0019] 2. The biomarker screening method according to item 1, wherein the candidate biomarkers are selected from known biomarkers, preferably by the following method: the detection sites of the 450K methylation chip are merged into the same biomarker according to their proximity in the genome, and finally only biomarkers with no less than 3 detection sites are retained as candidate biomarkers.

[0020] 3. The biomarker screening method according to item 1 or 2, wherein the healthy plasma sample is derived from a known healthy plasma sample, preferably from the NCBI SRA database GSE164600 dataset;

[0021] Preferably, the normal tissue sample is a normal lung tissue sample, preferably derived from a known normal lung tissue sample, and more preferably derived from adjacent normal tissue samples from the LUSC and LUAD projects in the TCGA database;

[0022] Preferably, the tumor tissue sample is a lung cancer tumor tissue sample, preferably derived from a known lung cancer tumor tissue sample, and more preferably derived from tumor samples from the LUSC and LUAD projects in the TCGA database, and is a sample with a tumor purity higher than 90%.

[0023] 4. The marker screening method according to any one of items 1-3, wherein the methylation level distribution pattern feature value is represented by a beta distribution, denoted as Beta(α,β), where α and β are the shape parameters of the distribution, and the calculation formula is as follows:

[0024]

[0025]

[0026] Where μ is the mean methylation level of a candidate biomarker in multiple samples within a group of three sample groups: healthy plasma sample, normal tissue sample, and tumor tissue sample, and σ is the variance of the methylation level of the biomarker in multiple samples within the same group.

[0027] 5. The biomarker screening method according to any one of items 1-4, wherein calculating the tissue origination feature vector of each of the L candidate biomarkers of the test sample includes:

[0028] Calculate the likelihood p of a sequencing fragment in the test sample originating from one of three source components: healthy plasma, normal tissue, or tumor tissue. r,t To obtain the proportion of a certain sequencing fragment in the sample to be tested that originates from each of the aforementioned source components;

[0029] The expectation-maximization algorithm is used to determine the tissue tracing ratio of the test sample based on the detection results within the range of each candidate biomarker.

[0030] 6. The marker screening method according to item 5, wherein the likelihood value p r,t The formula is:

[0031]

[0032] in, The methylation level distribution of a candidate marker m in a certain source component t;

[0033] and Let be the shape parameter of the beta distribution;

[0034] B(·) is the beta function;

[0035] t is 1, 2, or 3, representing the three source components: healthy plasma, normal tissue, and tumor tissue, respectively.

[0036] r = r1r2…r j …r n It is the methylation encoding form of a sequencing fragment in the sample to be tested, r j ∈{0,1}, where 0 represents the unmethylated state and 1 represents the methylated state.

[0037] 7. The marker screening method according to item 5 or 6, wherein the expectation-maximization algorithm employs the following method:

[0038] Initialization: Randomly initialize the proportions of healthy plasma (t=1), normal tissue (t=2), and tumor tissue (t=3) as θ1, θ2 (0≤θ1, θ2≤1 and θ1+θ2≤1) and θ3=1-θ1-θ2, respectively;

[0039] E-Step:

[0040] M-step:

[0041] Repeat the E-step and M-step iterations until the values ​​of θ1 and θ2 converge;

[0042] Where, r (i) This represents the i-th sequencing fragment of a sample in the test sample, which has r (i) =r1r2…r j …r n The methylation coding sequence is in the form of N, and the total number of sequencing fragments is z. (i) This indicates the tissue origin of the sequencing fragment.

[0043] 8. The biomarker screening method according to any one of items 1-7, wherein when the average prediction accuracy is greater than 0.7, Q biomarkers are screened as biomarkers for subsequent cancer detection.

[0044] 9. The marker selection method according to item 8, wherein a linear classification model with an average prediction accuracy greater than 0.7 is selected based on 5-fold cross-validation.

[0045] 10. A method for screening for cancer, comprising:

[0046] The Q markers selected by the marker screening method described in any one of items 1-9 are used to vote on the unknown samples, thereby interpreting the unknown samples.

[0047] 11. The method for screening cancer according to item 10, wherein the voting method adopts the following formula:

[0048]

[0049] Where, X = {x m} represents the organizational origin feature vector of each marker of the unknown sample;

[0050] x m =(x m1 ,x m2 ,x m3 ) represents the tissue origination feature vector of the unknown sample on marker m;

[0051] h m (·)∈{0,1} is a linear classification model trained on marker m;

[0052] The percentage of votes for a "1" prediction given by a linear classification model with an average prediction accuracy greater than 0.7;

[0053] The set threshold for the percentage of votes cast.

[0054] 12. The cancer detection method according to item 11, wherein h m (·)∈{0,1} gives a prediction result of "0" or "1", when the voting proportion does not exceed the threshold. If the sample is classified as "0", it is considered a normal sample; otherwise, it is classified as "1", which is considered a disease sample.

[0055] 13. A biomarker screening device based on methylation sequencing, comprising a candidate biomarker selection unit, a methylation level distribution pattern feature value calculation unit, a methylation sequencing unit, a tissue origination feature vector calculation unit, and a biomarker screening unit;

[0056] The candidate marker selection unit is used to select L candidate markers;

[0057] The methylation level distribution pattern feature value calculation unit is used to calculate the methylation level distribution pattern feature value of each candidate biomarker in three groups of samples: healthy plasma sample, normal tissue sample and tumor tissue sample, i.e., 3L feature values ​​are obtained by calculation.

[0058] The methylation sequencing unit is used to perform methylation sequencing on the sample to be tested to obtain the methylation sequencing results of the sample to be tested;

[0059] The tissue origination feature vector calculation unit is used to calculate the tissue origination feature vector of each of the L candidate biomarkers in the test sample based on the methylation sequencing results of the test sample and using the methylation level distribution pattern feature value.

[0060] The biomarker screening unit is used to construct a linear classification model based on the tissue traceability feature vector and the cancer or health status of the sample to be tested, and to screen out Q biomarkers as biomarkers for subsequent cancer detection based on the average prediction accuracy of the linear classification model, where L is greater than or equal to Q, and both L and Q are integers greater than or equal to 1.

[0061] 14. The biomarker screening device according to item 13, wherein the candidate biomarkers are obtained from known biomarkers, preferably selected by the following method: the detection sites of the 450K methylation chip are combined into the same biomarker according to their proximity in the genome, and finally only biomarkers with no less than 3 detection sites are retained as candidate biomarkers.

[0062] 15. The biomarker screening device according to item 13 or 14, wherein the unit for calculating the methylation level distribution pattern feature value includes a subunit for acquiring healthy plasma samples, a subunit for acquiring normal tissue, a subunit for acquiring tumor tissue, and a unit for calculating the methylation level distribution pattern feature value of each candidate biomarker.

[0063] The subunit for obtaining healthy plasma samples is used to obtain them from known healthy plasma samples, preferably from the NCBI SRA database GSE164600 dataset;

[0064] The subunit for obtaining normal tissue samples is used to obtain normal tissue samples from known normal tissue samples, preferably from adjacent normal tissue samples of the LUSC and LUAD projects in the TCGA database, and preferably, the normal tissue samples are normal lung tissue samples.

[0065] The tumor tissue sample acquisition subunit is used to obtain tumor samples from known tumor tissue samples, preferably from the LUSC and LUAD projects in the TCGA database, and preferably, the tumor samples are lung cancer tumor samples.

[0066] 16. The marker screening apparatus according to any one of items 13-15, wherein the methylation level distribution pattern characteristic value is represented by a beta distribution, denoted as Beta(α,β), where α and β are shape parameters of the distribution, and the calculation formula is as follows:

[0067]

[0068]

[0069] Where μ is the mean methylation level of a candidate biomarker in multiple samples within a group of three sample groups: healthy plasma sample, normal tissue sample, and tumor tissue sample, and σ is the variance of the methylation level of the biomarker in multiple samples within the same group.

[0070] 17. The biomarker screening apparatus according to any one of items 13-16, wherein the unit for calculating the tissue traceability feature vector includes a subunit for calculating the likelihood value of a sequencing fragment in the test sample originating from each of the three source components: healthy plasma, normal tissue, and tumor tissue, and a subunit for determining the tissue traceability ratio of the test sample based on the detection results within each candidate biomarker range;

[0071] The subunit that calculates the likelihood value of a sequencing fragment in the test sample originating from one of the three source components—healthy plasma, normal tissue, and tumor tissue—is used to calculate the likelihood value p of the sequencing fragment in the test sample originating from that specific source component. r,t To determine the probability that a particular sequencing fragment in the sample to be tested originates from each of the aforementioned source components;

[0072] The sub-unit that determines the proportion of tissue traceability of the test sample based on the detection results within each marker range is used to determine the proportion of the test sample originating from a certain source component based on the detection results within each candidate marker range using the expectation-maximization algorithm.

[0073] 18. The marker screening device according to item 17, wherein the likelihood value p r,t The formula is as follows:

[0074]

[0075] in, The methylation level distribution of a candidate marker m in a certain source component t;

[0076] and Let be the shape parameter of the beta distribution;

[0077] B(·) is the beta function;

[0078] t is 1, 2, or 3, representing the three source components: healthy plasma, normal tissue, and tumor tissue, respectively.

[0079] r = r1r2…r j …r n It is the methylation encoding form of a sequencing fragment in the sample to be tested, r j ∈{0,1}, where 0 represents the unmethylated state and 1 represents the methylated state.

[0080] 19. The marker screening device according to item 17 or 18, wherein the expectation-maximization algorithm is as follows:

[0081] Initialization: Randomly initialize the proportions of healthy plasma (t=1), normal lung tissue (t=2), and tumor tissue (t=3) as θ1, θ2 (0≤θ1, θ2≤1 and θ1+θ2≤1) and θ3=1-θ1-θ2, respectively;

[0082] E-Step:

[0083] M-step:

[0084] Repeat the E-step and M-step iterations until the values ​​of θ1 and θ2 converge, thereby obtaining the tissue source tracing feature vector;

[0085] Where, r (i) This represents the i-th sequencing fragment of a sample in the test sample, which has r. (i) =r1r2…r j …r n The methylation coding sequence is in the form of N, and the total number of sequencing fragments is z. (i) This indicates the tissue origin of the sequencing fragment.

[0086] 20. The marker screening apparatus according to any one of items 13-19, wherein the marker screening unit includes a linear classification model construction subunit and a linear classification model selection subunit, the linear classification model selection subunit being used to screen out Q markers based on the average prediction accuracy of the linear classification model, preferably, the average prediction accuracy of the selected linear classification model is >0.7.

[0087] 21. The marker screening device according to item 20, wherein the selection is performed using 5-fold cross-validation.

[0088] 22. A cancer detection device based on methylation sequencing, comprising a biomarker screening device, a voting unit, and an interpretation unit as described in any one of items 13-21;

[0089] The voting unit is used to vote on unknown samples based on the Q selected markers;

[0090] The interpretation unit is used to interpret unknown samples to determine whether the unknown sample is a cancer sample or a normal sample.

[0091] 23. The cancer detection device according to item 22, wherein voting is performed using the following formula:

[0092]

[0093] Where, X = {x m} represents the organizational origin tracing results of various markers in the unknown sample;

[0094] xm =(x m1 ,x m2 ,x m3 ) represents the tissue origination feature vector of the unknown sample on marker m;

[0095] h m (·)∈{0,1} is a linear classification model trained on marker m;

[0096] The percentage of votes for a "1" prediction given by a linear classification model with an average prediction accuracy greater than 0.7; The set threshold for the percentage of votes cast.

[0097] 23. The cancer detection device according to item 18, wherein h m (·)∈{0,1} gives a prediction result of "0" or "1", when the voting proportion does not exceed the threshold. If the sample is classified as "0", it is considered a normal sample; otherwise, it is classified as "1", which is considered a disease sample.

[0098] 24. A biomarker screening system based on methylation sequencing, comprising a first processor and a first memory, the first processor and the first memory communicating via a first bus, the first memory storing first program instructions, the first processor executing the first program instructions, the first program instructions executing any one of items 1-9 of the biomarker screening method.

[0099] 25. A cancer detection system based on methylation sequencing, comprising a second processor and a second memory, the second processor and the second memory communicating via a second bus, the second memory storing second program instructions, the second processor executing the second program instructions, the second program instructions executing any one of items 10-12 of the cancer detection method.

[0100] 26. A non-transitory computer-readable storage medium for biomarker screening based on methylation sequencing, the non-transitory computer-readable storage medium storing computer instructions A, the computer instructions A causing a computer to execute the biomarker screening method described in any one of items 1-9.

[0101] 27. A non-transitory computer-readable storage medium for cancer detection based on methylation sequencing, the non-transitory computer-readable storage medium storing computer instructions B, the computer instructions B causing a computer to perform the cancer detection method described in any one of items 10-12.

[0102] The effects of the invention

[0103] The method described in this invention adopts a mixed model assumption of three tissue sources: healthy plasma, normal tissue, and tumor tissue, rather than the traditional two-component mixed assumption of healthy plasma and tumor tissue. Theoretically, this assumption can not only give a judgment on whether the sample to be tested is healthy or diseased, but also give an accurate distinction on whether the sample is cancerous.

[0104] The method described in this invention uses the methylation linkage information of CpG sites in each sequencing fragment as the basic detection unit of the model. Compared with the traditional method that uses biomarkers as the basic unit, it can provide more accurate tissue tracing. Even when the methylation signal on the corresponding biomarker is weak, a strong methylation signal can still be identified at the sequencing fragment level. Therefore, it can still provide relatively accurate interpretation in the early stage of cancer when the content of tumor-specific source DNA is still very low. Thus, it has high application value in early cancer screening.

[0105] The method described in this invention fully utilizes the potential discriminative power of a large number of candidate markers and adopts an ensemble learning strategy to provide more accurate prediction results while also ensuring higher generalization performance. Attached Figure Description

[0106] Figure 1 This is a schematic diagram of the model classification effect in Embodiment 3 of the present invention. Detailed Implementation

[0107] The present invention will now be described in detail with reference to the accompanying drawings, wherein the same numerals in all the drawings denote the same features. While specific embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.

[0108] It should be noted that certain terms are used in the specification and claims to refer to specific components. Those skilled in the art will understand that different terms may be used to refer to the same component. This specification and claims do not distinguish components based on differences in terminology, but rather on differences in function. The terms "comprising" or "including" used throughout the specification and claims are open-ended and should be interpreted as "comprising but not limited to." The following descriptions are preferred embodiments for carrying out the invention; however, these descriptions are for the purpose of understanding the general principles of the specification and are not intended to limit the scope of the invention. The scope of protection of this invention is determined by the appended claims.

[0109] This invention provides a biomarker screening method based on methylation sequencing, comprising:

[0110] Select L candidate markers;

[0111] For each candidate biomarker, the methylation level distribution pattern characteristic value of its respective methylation level in three groups of samples (healthy plasma sample, normal tissue sample, and tumor tissue sample) was calculated, that is, 3L characteristic values ​​were obtained.

[0112] The methylation sequencing results of the sample to be tested are obtained by performing methylation sequencing on the sample to be tested;

[0113] Based on the methylation sequencing results of the sample to be tested, the tissue origination feature vector of each of the L candidate biomarkers in the sample to be tested is calculated using the methylation level distribution pattern feature value.

[0114] A linear classification model is constructed based on the tissue source feature vector and the cancer or health status of the sample to be tested. Q biomarkers are selected as biomarkers for subsequent cancer detection based on the average prediction accuracy of the linear classification model, where L is greater than or equal to Q and both L and Q are integers greater than or equal to 1.

[0115] In one implementation, the candidate biomarkers are obtained from known biomarkers, preferably selected as follows: the detection sites of the 450K methylation chip are merged into the same biomarker according to their proximity in the genome, and only biomarkers with no less than 3 detection sites are retained as candidate biomarkers. For example, the candidate biomarkers are from the article published by Kang et al. in 2017 (Kang S, Li Q, Chen Q, et al. Cancer Locator: non-invasive cancer diagnosis and tissue-of-origin prediction using methylation profiles of cell-free DNA. Genome Biol. 2017; 18(1):53.), which can yield 42,374 biomarkers. Based on the standard that at least half of the samples in each tissue group can obtain valid data, L = 33,961 candidate biomarkers are finally screened.

[0116] In one embodiment, the healthy plasma sample is a known source of healthy plasma, preferably from the NCBI SRA database GSE164600 dataset, and more preferably from 12 healthy individuals' WGBS data samples. In one embodiment, the tumor tissue sample is preferably a lung cancer tumor tissue sample, preferably from known lung cancer tumor samples, and more preferably from tumor samples from the LUSC and LUAD projects in the TCGA database, selecting only samples with tumor purity higher than 90%. More preferably, the lung cancer tumor tissue sample comprises 35 samples, including 6 LUAD (lung adenocarcinoma) samples and 29 LUSC (lung squamous cell carcinoma) samples. In one embodiment, the normal tissue sample is a normal lung tissue sample, preferably from known normal lung tissue samples, and more preferably from adjacent normal tissue samples from the LUSC and LUAD projects in the TCGA database. More preferably, the normal lung tissue sample comprises 71 samples.

[0117] In one implementation, the method for assessing tumor purity is derived from an article published by Aran D et al. in 2015 (Aran D, Sirota M, Butte AJ. Systematic pan-cancer analysis of tumor purity. Nat Commun. 2015; 6:8971.).

[0118] In one implementation, for each biomarker methylation level, it refers to the arithmetic mean of the methylation levels of the biomarker-covered detection sites in the sample.

[0119] In one embodiment, the methylation level distribution pattern feature values ​​are represented by a beta distribution, denoted as Beta(α,β), where α and β are the shape parameters of the distribution, and the calculation formula is as follows:

[0120]

[0121]

[0122] Where μ is the mean methylation level of a candidate biomarker in multiple samples within a group of three sample groups: healthy plasma sample, normal tissue sample, and tumor tissue sample, and σ is the variance of the methylation level of the biomarker in multiple samples within the same group.

[0123] For example, for a healthy plasma sample, a normal tissue sample, or a tumor tissue sample, μ refers to the mean methylation level of a candidate biomarker among multiple samples in the healthy plasma sample, normal tissue sample, or tumor tissue sample, and σ refers to the variance of the methylation level of the biomarker among multiple samples in the healthy plasma sample, normal tissue sample, or tumor tissue sample.

[0124] In one implementation, calculating the tissue origination feature vector of each of the L candidate biomarkers for the test sample includes:

[0125] Calculate the likelihood p of a sequencing fragment in the test sample originating from one of three source components: healthy plasma, normal tissue, or tumor tissue. r,t To obtain the proportion of a certain sequencing fragment in the sample to be tested that originates from each of the aforementioned source components;

[0126] The expectation-maximization algorithm is used to determine the tissue tracing ratio of the test sample based on the detection results within the range of each candidate biomarker.

[0127] The tissue origin feature vector refers to the proportion of cfDNA from the three source components in the sample to be tested.

[0128] In one implementation, the likelihood value p r,t The formula is:

[0129]

[0130] The methylation level distribution of a candidate marker m in a certain source component t;

[0131] and Let be the shape parameter of the beta distribution;

[0132] B(·) is the beta function;

[0133] t is 1, 2, or 3, representing the three source components: healthy plasma, normal tissue, and tumor tissue, respectively.

[0134] r = r1r2…r j …r n It is the methylation encoding form of a sequencing fragment in the sample to be tested, r j ∈{0,1}, where 0 represents the unmethylated state and 1 represents the methylated state;

[0135] n refers to the n CpG sites on a sequencing fragment.

[0136] Based on the calculation results obtained from the above steps, the likelihood value of a certain sequencing fragment of the test sample originating from each source component is calculated. Using the expectation-maximization algorithm, the tissue tracing ratio of the test sample based on the detection results within the range of each candidate biomarker can be obtained.

[0137] Preferably, the expectation maximization algorithm employs the following method:

[0138] Initialization: Randomly initialize the proportions of healthy plasma (t=1), normal tissue such as normal lung sample (t=2), and tumor tissue such as lung cancer tissue (t=3) as θ1, θ2 (0≤θ1, θ2≤1 and θ1+θ2≤1) and θ3=1-θ1-θ2, respectively;

[0139] E-Step:

[0140] M-step:

[0141] Where, r (i) This represents the i-th sequencing fragment of a sample in the test sample, which has r (i) =r1r2…r j …r n The methylation coding sequence is in the form of N, and the total number of sequencing fragments is z. (i) This indicates the tissue origin of the sequencing fragment, and P refers to the probability.

[0142] Repeat the E-step and M-step iterations until the values ​​of θ1 and θ2 converge.

[0143] For organizational tracing using the aforementioned Expectation-Maximization (EM) algorithm, compared to the brute-force method represented by GridSearch, it is an adaptive heuristic learning method. Especially when the amount of data is large and the theoretical search space is relatively large, it can greatly shorten the time consumption of the optimization solution process. However, the EM algorithm cannot guarantee that it will obtain the global optimum, but rather a local optimum related to the initial value. However, the risk of local optima can be reduced by multiple random initializations.

[0144] In one implementation, when the average prediction accuracy is greater than 0.7, Q biomarkers are selected as biomarkers for subsequent cancer detection. Preferably, a linear classification model with an average prediction accuracy greater than 0.7 is selected based on 5-fold cross-validation.

[0145] The linear classification model is a linear classification model trained based on the organizational origin feature vector of each candidate marker in the test sample. The classification model with excellent classification performance (average prediction accuracy above 0.7) is selected to screen out Q markers.

[0146] Preferably, the base learner is the tissue origination result of each marker in the sample to be tested, that is, the proportion of cfDNA from the three source components is used as the input feature value for discrimination.

[0147] Preferably, for the base learner, a logistic regression algorithm is used for binary classification. If the output result p≤0.5, it is classified as class "0", which is a normal sample; otherwise, it is classified as class "1", which is a disease sample.

[0148] This invention provides a method for cancer screening, comprising:

[0149] Based on the marker screening method described above, the Q markers selected are used to vote on the unknown samples, thereby interpreting the unknown samples.

[0150] In one implementation scheme, the voting method uses the following formula:

[0151]

[0152] Where, X = {x m} represents the organizational origin feature vector of each marker of the unknown sample;

[0153] x m =(x m1 ,x m2 ,x m3 ) represents the tissue origination feature vector of the unknown sample on marker m;

[0154] h m (·)∈{0,1} is a linear classification model trained on marker m;

[0155] The percentage of votes given for a "1" prediction for a selected linear classification model with an average prediction accuracy greater than 0.7 (referred to as the voting percentage);

[0156] The set threshold for the percentage of votes cast.

[0157] In one implementation, h m (·)∈{0,1} gives a prediction result of "0" or "1", that is, when the voting proportion does not exceed the threshold. If the sample is classified as "0", it is considered a normal sample; otherwise, it is classified as "1", which is considered a disease sample.

[0158] In one implementation, voting on unknown samples includes the following steps:

[0159] The methylation sequencing results of the unknown sample were obtained by performing methylation sequencing on the unknown sample.

[0160] Based on the methylation sequencing results of the unknown sample, the tissue origin feature vector of each of the Q markers of the unknown sample is calculated using the methylation level distribution pattern feature value;

[0161] The unknown samples are voted on using a voting method. The voting formula is as follows:

[0162]

[0163] Where, X = {x m} represents the organizational origin feature vector of each marker of the unknown sample;

[0164] x m =(x m1 ,x m2 ,x m3 ) represents the tissue origination feature vector of the unknown sample on marker m;

[0165] h m (·)∈{0,1} is a linear classification model trained on marker m;

[0166] The percentage of votes given for a "1" prediction for a selected linear classification model with an average prediction accuracy greater than 0.7 (referred to as the voting percentage);

[0167] The set threshold for the percentage of votes;

[0168] The unknown sample is determined to be either a normal sample or a cancerous sample based on the voting results.

[0169] The present invention provides a biomarker screening device based on methylation sequencing, which includes a candidate biomarker selection unit, a methylation level distribution pattern feature value calculation unit, a methylation sequencing unit, a tissue origination feature vector calculation unit, and a biomarker screening unit;

[0170] The candidate marker selection unit is used to select L candidate markers;

[0171] The methylation level distribution pattern feature value calculation unit is used to calculate the methylation level distribution pattern feature value of each candidate biomarker in three groups of samples: healthy plasma sample, normal tissue sample and tumor tissue sample, i.e., to calculate and obtain 3L feature values.

[0172] The methylation sequencing unit is used to perform methylation sequencing on the sample to be tested to obtain the methylation sequencing results of the sample to be tested;

[0173] The tissue origination feature vector calculation unit is used to calculate the tissue origination feature vector of each of the L candidate biomarkers in the test sample based on the methylation sequencing results of the test sample and using the methylation level distribution pattern feature value.

[0174] The biomarker screening unit is used to construct a linear classification model based on the tissue traceability feature vector and the cancer or health status of the sample to be tested, and to screen out Q biomarkers as biomarkers for subsequent cancer detection based on the average prediction accuracy of the linear classification model, where L is greater than or equal to Q, and both L and Q are integers greater than or equal to 1.

[0175] In one implementation, the candidate biomarkers are obtained from known biomarkers, preferably selected as follows: the detection sites of the 450K methylation chip are merged into a single biomarker based on their proximity in the genome, with adjacent detection sites within a distance of no more than 200 bases. Finally, only biomarkers with at least 3 detection sites are retained as candidate biomarkers. For example, the candidate biomarkers are obtained from an article published by Kang et al. in 2017, resulting in 42,374 biomarkers. Based on the criterion that at least half of the samples in each tissue group can obtain valid data, 33,961 usable candidate biomarkers are finally selected, i.e., L = 33,961.

[0176] In one implementation, the unit for calculating the methylation level distribution pattern feature value includes a sub-unit for acquiring healthy plasma samples, a sub-unit for acquiring normal tissue, a sub-unit for acquiring tumor tissue, and a unit for calculating the methylation level distribution pattern feature value of each candidate biomarker.

[0177] The subunit for obtaining healthy plasma samples is used to obtain healthy plasma samples from known healthy plasma samples, preferably from the NCBI SRA database GSE164600 dataset, and preferably from 12 healthy individuals' WGBS data samples.

[0178] The subunit for obtaining normal tissue samples is used to obtain normal tissue samples from known normal tissue samples, preferably from adjacent normal tissue samples of the LUSC and LUAD projects in the TCGA database, preferably, the normal tissue samples are normal lung tissue samples, preferably, the normal lung tissue samples are 71 cases;

[0179] The tumor tissue sample acquisition subunit is used to obtain tumor samples from known tumor tissue samples, preferably from the LUSC and LUAD projects in the TCGA database. Preferably, the tumor samples are lung cancer tumor samples, and more preferably 6 LUAD (lung adenocarcinoma) samples and 29 LUSC (lung squamous cell carcinoma) samples.

[0180] Preferably, the method for assessing tumor purity is derived from an article published by Aran D et al. in 2015.

[0181] In one embodiment, the methylation level distribution pattern feature value is represented by a beta distribution, denoted as Beta(α,β), where α and β are the shape parameters of the distribution, and the calculation formula is as follows:

[0182]

[0183]

[0184] Where μ is the mean methylation level of a candidate biomarker among multiple samples in one of the three groups of samples: healthy plasma sample, normal tissue sample, and tumor tissue sample, and σ is the variance of the methylation level of the biomarker among multiple samples in the same group.

[0185] For example, for a healthy plasma sample, a normal tissue sample, or a tumor tissue sample, μ refers to the mean methylation level of a candidate biomarker among multiple samples in the healthy plasma sample, normal tissue sample, or tumor tissue sample, and σ refers to the variance of the methylation level of the biomarker among multiple samples in the healthy plasma sample, normal tissue sample, or tumor tissue sample.

[0186] In one implementation, the unit for calculating tissue origin feature vectors includes a subunit for calculating the likelihood value of a sequencing fragment in the test sample originating from each of the three source components: healthy plasma, normal tissue, and tumor tissue, and a subunit for determining the tissue origin ratio of the test sample based on the detection results within each candidate biomarker range.

[0187] The subunit that calculates the likelihood value of a sequencing fragment in the test sample originating from one of the three source components—healthy plasma, normal tissue, and tumor tissue—is used to calculate the likelihood value p of the sequencing fragment in the test sample originating from that specific source component. r,t To determine the probability that a sequencing fragment in the sample to be tested originates from each source component;

[0188] The sub-unit for determining the proportion of tissue traceability of the test sample based on the detection results within each marker range is used to determine the proportion of the test sample originating from a certain source component based on the detection results within each candidate marker range using the expectation-maximization algorithm.

[0189] The preferred likelihood value p r,t The formula is:

[0190]

[0191] in, This refers to the methylation level distribution of a candidate biomarker m in a given source component t. and Let B(·) be the shape parameter of the beta distribution, and B(·) be the beta function.

[0192] t is 1, 2, or 3, representing the three source components: healthy plasma, normal tissue, and tumor tissue, respectively.

[0193] r = r1r2…r j …r n It is the methylation coding form of a sequencing fragment in the sample to be tested, r j ∈{0,1}, where 0 represents the unmethylated state and 1 represents the methylated state.

[0194] The preferred expectation-maximization algorithm is as follows:

[0195] Initialization: Randomly initialize the proportions of healthy plasma (t=1), normal tissue (t=2), and tumor tissue (t=3) as θ1, θ2 (0≤θ1, θ2≤1 and θ1+θ2≤1) and θ3=1-θ1-θ2, respectively;

[0196] E-Step:

[0197] M-step:

[0198] Where, r (i) This represents the i-th sequencing fragment of a sample in the test sample, which has r. (i) =r1r2…r j …r n The methylation coding sequence is in the form of N, and the total number of sequencing fragments is z. (i) This indicates the tissue origin of the sequencing fragment.

[0199] Repeat the E-step and M-step iterations until the values ​​of θ1 and θ2 converge, thereby obtaining the organizational source tracing feature vector.

[0200] In one implementation, the marker selection unit includes a linear classification model construction subunit and a linear classification model selection subunit. The linear classification model selection subunit is used to select Q markers based on the average prediction rate of the linear classification model. Preferably, the average prediction accuracy of the selected linear classification model is >0.7.

[0201] It is a linear classification model trained based on the organizational origin feature vector of each marker, and the linear classification model with excellent classification performance is selected based on the average prediction accuracy being greater than 0.7, that is, Q markers are selected.

[0202] The chosen method involves performing 5-fold cross-validation on the linear classification model to provide the classification performance of the base learners trained on each marker. The linear classification model with an average prediction accuracy greater than 0.7 is selected, i.e., Q markers are chosen.

[0203] The present invention provides a cancer detection device based on methylation sequencing, which includes the biomarker screening device, voting unit and interpretation unit described above. The voting unit is used to vote on unknown samples based on the Q biomarkers selected.

[0204] The interpretation unit is used to interpret unknown samples to determine whether the unknown sample is a cancer sample or a normal sample.

[0205] In one implementation scheme, voting is conducted using the following formula:

[0206]

[0207] Where, X = {x m} represents the organizational origin feature vector of each marker of the unknown sample;

[0208] x m =(x m1 ,x m2 ,x m3 ) represents the tissue origination feature vector of the unknown sample on marker m;

[0209] h m (·)∈{0,1} is a linear classification model trained on marker m;

[0210] The percentage of votes given for a "1" prediction for a selected linear classification model with an average prediction accuracy greater than 0.7 (referred to as the voting percentage);

[0211] The set threshold for the percentage of votes cast.

[0212] In one implementation, the tissue origination feature vector of each marker of the unknown sample is obtained by the following method:

[0213] Methylation sequencing is performed on an unknown sample to obtain the methylation sequencing results of the unknown sample;

[0214] Based on the methylation sequencing results of the unknown sample, the tissue origination feature vector of each of the Q biomarkers of the unknown sample is calculated using the methylation level distribution pattern feature value. The specific operation method is the same as the biomarker screening method.

[0215] When the voting ratio does not exceed the threshold If the sample is classified as "0", it is considered a normal sample; otherwise, it is classified as "1", which is considered a disease sample.

[0216] This invention provides a biomarker screening system based on methylation sequencing, which includes a first processor and a first memory. The first processor and the first memory communicate through a first bus. The first memory stores first program instructions. The first processor executes the first program instructions, and the first program instructions perform the biomarker screening method described above.

[0217] The present invention provides a cancer detection system based on methylation sequencing, which includes a second processor and a second memory. The second processor and the second memory communicate through a second bus. The second memory stores second program instructions. The second processor executes the second program instructions, and the second program instructions execute the cancer detection method described above.

[0218] The present invention provides a non-transitory computer-readable storage medium for biomarker screening based on methylation sequencing, wherein the non-transitory computer-readable storage medium stores computer instructions A, which cause a computer to execute the biomarker screening method described above.

[0219] The present invention provides a non-transitory computer-readable storage medium for cancer detection based on methylation sequencing, wherein the non-transitory computer-readable storage medium stores computer instructions B, which cause a computer to execute the cancer detection method described above.

[0220] Example

[0221] This invention provides a general and / or specific description of the materials and methods used in the experiments. In the following examples, unless otherwise specified, % represents wt%, i.e., weight percentage. Reagents or instruments used, unless otherwise specified, are all commercially available conventional reagent products.

[0222] Example 1: Biomarker Screening Method

[0223] (1) Select L candidate biomarkers: The candidate biomarkers are from an article published by Kang et al. in 2017. They combined the detection sites of the 450K methylation chip according to their proximity in the genome, and merged the detection sites with an adjacent distance of no more than 200 bases into the same biomarker. Finally, only biomarkers with no less than 3 detection sites were retained, resulting in a total of 42,374 biomarkers. Based on the standard that at least half of the samples in each tissue group can obtain valid data, 33,961 usable biomarkers were finally selected as candidate biomarkers, i.e., L = 33,961.

[0224] (2) For each candidate biomarker, the methylation level distribution pattern characteristic value was calculated in the three groups of samples: healthy plasma sample, normal tissue sample, and tumor tissue sample. That is, 3L characteristic values ​​were obtained:

[0225] The healthy plasma samples were WGBS data samples from 12 healthy individuals from the NCBI SRA database GSE164600 dataset.

[0226] The lung cancer tumor tissue samples were tumor samples from the LUSC and LUAD projects in the TCGA database, and only samples with a tumor purity higher than 90% were selected, totaling 35 cases, including 6 LUAD (lung adenocarcinoma) samples and 29 LUSC (lung squamous cell carcinoma) samples. The tumor purity assessment method CPE (consensus measurement of purity estimates) was derived from an article published by Aran D et al. in 2015.

[0227] The normal lung tissue samples were 71 cases of adjacent normal tissue samples from the LUSC and LUAD projects in the TCGA database.

[0228] The biomarker methylation level refers to the mean methylation level of the biomarker covering the detection sites in the sample;

[0229] The methylation distribution pattern characteristic values ​​are represented by a beta distribution, denoted as Beta(α,β), where α and β are the shape parameters (i.e., key parameters) of the distribution. The parameter estimation method is moment estimation, that is, calculating the mean μ and variance σ. Based on the relationship between the key parameters and the mean and standard deviation, the estimated values ​​of the key parameters are calculated, and the calculation formula is as follows:

[0230]

[0231]

[0232] μ is the mean methylation level of a candidate biomarker among multiple samples in one of the three groups of samples: healthy plasma sample, normal tissue sample, and tumor tissue sample; σ is the variance of the methylation level of the biomarker among multiple samples in the same group.

[0233] For example, the methylation level of a certain biomarker detected in healthy population sample data (the raw sequencing results were obtained using the conventional WGBS high-throughput sequencing method, and then the methylation status of each sequencing fragment was analyzed using Bismark software) is as follows. This refers to the methylation level of a certain biomarker detected in multiple samples of healthy plasma:

[0234] [0.389121338912134,0.370950888192268,0.277044854881267,0.267344497607655,0.302179962894249,0.334645669291339,0.212361331220285,0.430724637681159,0.16 [4156090617958,0.19335587825741,0.0460829493087558,0.344767441860465], the corresponding mean and variance can be calculated as μ = 0.277728 and σ = 0.01189968, respectively. Based on the relationship between the key parameters and the mean and variance, α = 4.403983 and β = 11.453 can be calculated. For this biomarker, there are also methylation data from normal lung tissue and tumor tissue. Their various corresponding beta distribution parameters can be calculated using the same method, which will not be listed here.

[0235] (3) Perform methylation sequencing on the sample to be tested to obtain the methylation sequencing results of the sample to be tested;

[0236] (4) Based on the methylation sequencing results of the sample to be tested, the tissue origination feature vector of each of the L candidate biomarkers of the sample to be tested is calculated using the methylation level distribution pattern feature value;

[0237] Calculating the tissue origination feature vector of each of the L candidate biomarkers for the test sample includes:

[0238] Calculate the likelihood p of a sequencing fragment in the test sample originating from a certain source component. r,t To obtain the probability that a certain sequencing fragment in the sample to be tested originates from each of the said source components, where the likelihood value p r,t The formula is:

[0239]

[0240] in, The methylation level distribution of a candidate biomarker m in a certain source component t;

[0241] and Let be the shape parameter of the beta distribution;

[0242] B(·) is the beta function;

[0243] t is 1, 2, or 3, representing the three source components: healthy plasma, normal tissue, and tumor tissue, respectively.

[0244] r = r1r2…r j …r n It is the methylation encoding form of a sequencing fragment in the sample to be tested, r j ∈{0,1}, where 0 represents the unmethylated state and 1 represents the methylated state.

[0245] Next, the expectation-maximization algorithm is used to determine the tissue tracing ratio of the test sample based on the detection results within each candidate biomarker range. The expectation-maximization algorithm employs the following method:

[0246] Initialization: Randomly initialize the proportions of healthy plasma (t=1), normal tissue (t=2), and tumor tissue (t=3) as θ1, θ2 (0≤θ1, θ2≤1 and θ1+θ2≤1) and θ3=1-θ1-θ2, respectively;

[0247] E-Step:

[0248] M-step:

[0249] Repeat the E-step and M-step iterations until the values ​​of θ1 and θ2 converge. For example, for one of the test samples, the marker numbered 12 eventually converges to (0.812661, 0.059183, 0.128156).

[0250] Where, r (i) This represents the i-th sequencing fragment of a sample to be tested, which has r (i) =r1r2…r j …r n The methylation coding sequence is in the form of N, and the total number of sequencing fragments is z. (i) This indicates the tissue origin of the sequencing fragment.

[0251] (5) Based on the tissue traceability feature vector and the cancer or health status of the sample to be tested, a linear classification model is constructed, and Q biomarkers are selected as biomarkers for subsequent cancer detection based on the average prediction accuracy of the linear classification model, where L is greater than or equal to Q, and L and Q are both integers greater than 1: the average prediction accuracy is greater than 0.7, and 5-fold cross-validation is used to select the linear classification model with an average prediction accuracy greater than 0.7, thereby selecting 298 biomarkers, i.e. Q = 298.

[0252] Example 2: Method for screening lung cancer

[0253] The unknown samples were identified by voting on the 298 markers selected in Example 1.

[0254] The voting method uses the following formula:

[0255]

[0256] Where, X = {x m} represents the organizational origin feature vector of each marker of the unknown sample;

[0257] x m =(x m1 ,x m2 ,x m3 ) represents the tissue origination feature vector of the unknown sample on marker m;

[0258] h m (·)∈{0,1} is a linear classification model trained on marker m, which ultimately gives a prediction result of "0" or "1";

[0259] The voting percentage (or simply voting percentage) is the percentage of "1" predictions given for the selected linear classification model with an average prediction accuracy greater than 0.7. The value is between 0 and 1.

[0260] In this embodiment, the voting ratio of one unknown sample is 0.173; The set voting ratio threshold is set to 0.231 in this embodiment. When the voting ratio does not exceed this threshold, the sample is classified as "0", which is a normal sample. Otherwise, it is classified as "1", which is a disease sample. In this embodiment, the sample is classified as a normal sample because the voting ratio is less than the voting ratio threshold. The actual health status of the sample to be tested is also labeled as healthy, so an accurate classification is obtained.

[0261] The method for calculating the organizational origin feature vectors of each marker of the unknown sample is the same as in Example 1.

[0262] Example 3

[0263] The dataset from Bocheng (Beijing) Technology Co., Ltd. was used to screen biomarkers according to the method described in Example 1, and voting interpretation was performed according to the method described in Example 2. The dataset contained 60 cases, including 38 healthy individuals, 19 LUAD patients, and 3 LUSC patients. 80% of the dataset was used as the training set for five-fold cross-validation to screen cancer detection biomarkers, and 20% of the dataset was used as the test set to evaluate the predictive performance of the voting interpretation method. Regardless of the splitting of the training and test sets, or the splitting of samples for five-fold cross-validation within the training set, stratified sampling was used to ensure that the proportion of samples in each group remained unchanged before and after the split. After each random splitting of the training and test sets, a model evaluation result was obtained. To avoid potential bias in a single model evaluation result, 20 such evaluations were performed, and the results are as follows. Figure 1 As shown, the vertical axis in the figure has three meanings: the box plot represents the voting ratio, the AUC curve represents the AUC value ranging from 0 to 1, and the accuracy curve represents the accuracy value ranging from 0 to 1. The AUC value is the area under the ROC curve (AUC). Generally, a higher AUC value indicates better classification performance. The ROC curve is the receiver operating characteristic curve, plotted with the true positive rate (TPR) (sensitivity) on the vertical axis and the false positive rate (FPR) (1-specificity) on the horizontal axis, based on a series of set thresholds. It reflects the changes in TPR and FPR under different thresholds. The closer the curve is to the upper left corner, the better the model's classification performance. Accuracy describes the classification model's ability to judge the overall data; Accuracy = (TP + TN) / (TP + NP + TN + FN).

[0264] from Figure 1 As can be seen, in 20 model evaluations, the voting proportions (box plot part) obtained by the samples with the actual class "0" (normal samples) in the test set were significantly lower than those of the samples with the actual class "1" (cancer samples) in the ensemble discrimination. The average AUC and average accuracy were 0.8919643 and 0.9, respectively. Moreover, the optimal voting proportion threshold (star points in the figure) in each cross-validation was relatively stable, with an average value of 0.25, indicating that the model has good classification performance.

[0265] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A biomarker screening method based on methylation sequencing, comprising: Select L candidate markers; For each candidate biomarker, the methylation level distribution pattern characteristic value of its respective methylation level in three groups of samples (healthy plasma sample, normal tissue sample, and tumor tissue sample) was calculated, that is, 3L characteristic values ​​were obtained. The methylation sequencing results of the sample to be tested are obtained by performing methylation sequencing on the sample to be tested; Based on the methylation sequencing results of the sample to be tested, the tissue origination feature vector of each of the L candidate biomarkers in the sample to be tested is calculated using the methylation level distribution pattern feature value. A linear classification model is constructed based on the tissue source feature vector and the cancer or health status of the sample to be tested. Based on the average prediction accuracy of the linear classification model, Q biomarkers are selected as biomarkers for subsequent cancer detection, where L is greater than or equal to Q, and both L and Q are integers greater than or equal to 1. Calculating the tissue origination feature vector of each of the L candidate biomarkers for the test sample includes: Calculate the likelihood that a sequencing fragment in the sample to be tested originates from one of three source components: healthy plasma, normal tissue, or tumor tissue. To obtain the proportion of a certain sequencing fragment in the sample to be tested that originates from each of the aforementioned source components; Likelihood value The formula is: in, For a certain source component A candidate marker methylation level distribution; and Let be the shape parameter of the beta distribution; For beta function; t is 1, 2, or 3, representing the three source components: healthy plasma, normal tissue, and tumor tissue, respectively. It is the methylation encoding form of a sequencing fragment in the sample to be tested. 0 indicates the unmethylated state, and 1 indicates the methylated state; The expectation-maximization algorithm is used to determine the tissue tracing ratio of the test sample based on the detection results within the range of each candidate biomarker.

2. The marker screening method according to claim 1, wherein, The candidate markers are selected from known markers.

3. The biomarker screening method according to claim 2, wherein the candidate biomarkers are selected according to the following method: the detection sites of the 450K methylation chip are merged into the same biomarker according to their proximity in the genome, and finally only biomarkers with no less than 3 detection sites are retained as candidate biomarkers.

4. The marker screening method according to claim 1, wherein, The healthy plasma sample was derived from a known healthy plasma sample.

5. The biomarker screening method according to claim 4, wherein the healthy plasma sample is derived from the NCBI SRA database GSE164600 dataset.

6. The marker screening method according to claim 1, wherein, The normal tissue sample is a normal lung tissue sample.

7. The biomarker screening method according to claim 6, wherein the normal tissue sample is derived from a known normal lung tissue sample.

8. The biomarker screening method according to claim 6, wherein the normal tissue sample is a neighboring tissue sample from the LUSC and LUAD projects in the TCGA database.

9. The marker screening method according to claim 1, wherein, The tumor tissue sample is a lung cancer tumor tissue sample.

10. The biomarker screening method according to claim 9, wherein the tumor tissue sample is derived from a known lung cancer tumor tissue sample.

11. The biomarker screening method according to claim 9, wherein the tumor tissue sample is a tumor sample from the LUSC and LUAD projects in the TCGA database, and is a sample with a tumor purity of more than 90%.

12. The marker screening method according to claim 1, wherein, The methylation level distribution pattern characteristic values ​​are represented by a beta distribution. It means that, among them and The shape parameter of this distribution is calculated using the following formula: in, It is the mean methylation level of a candidate biomarker in multiple samples from a specific group of three sample groups: healthy plasma samples, normal tissue samples, and tumor tissue samples. It is the variance of the methylation level of a certain marker among multiple samples in a certain group of samples.

13. The marker screening method according to claim 1, wherein, The expectation maximization algorithm employs the following method: Initialization: Randomly initialize healthy plasma ( ), normal tissue ( ) and tumor tissue ( The proportions are respectively , and ,in and ; E-Step: ; M-step: ; Repeat the E-step and M-step iterations until... and The values ​​converge. in, This indicates the first sample in the test sample. Sequencing fragments, which have The methylation coding sequence is in the form of N, and the total number of sequencing fragments is N. This indicates the tissue origin of the sequencing fragment.

14. The marker screening method according to any one of claims 1-13, wherein, When the average prediction accuracy is greater than 0.7, Q biomarkers are selected as biomarkers for subsequent cancer detection.

15. The marker screening method according to claim 14, wherein, Linear classification models with an average prediction accuracy greater than 0.7 were selected based on 5-fold cross-validation.

16. A method for screening for cancer, comprising: The Q markers selected by the marker screening method according to any one of claims 1-15 are used to vote on the unknown samples, thereby interpreting the unknown samples.

17. The method for screening cancer according to claim 16, wherein, The voting method uses the following formula: in, For the organizational origination feature vectors of various markers of unknown samples; For unknown samples in markers Organizational source traceability feature vectors; For the marker The linear classification model trained on it; The voting percentage for the "1" prediction given by the selected linear classification model with an average prediction accuracy greater than 0.7; The set threshold for the percentage of votes cast.

18. The method for screening cancer according to claim 17, wherein, Provide a prediction result of "0" or "1" when the voting proportion does not exceed the threshold. If the result is positive, the sample is classified as "0" and is considered a normal sample; otherwise, it is classified as "1" and is considered a disease sample.

19. A biomarker screening device based on methylation sequencing, comprising a candidate biomarker selection unit, a methylation level distribution pattern feature value calculation unit, a methylation sequencing unit, a tissue origination feature vector calculation unit, and a biomarker screening unit; The candidate marker selection unit is used to select L candidate markers; The methylation level distribution pattern feature value calculation unit is used to calculate the methylation level distribution pattern feature value of each candidate biomarker in three groups of samples: healthy plasma sample, normal tissue sample and tumor tissue sample, i.e., 3L feature values ​​are obtained by calculation. The methylation sequencing unit is used to perform methylation sequencing on the sample to be tested to obtain the methylation sequencing results of the sample to be tested; The tissue origination feature vector calculation unit is used to calculate the tissue origination feature vector of each of the L candidate biomarkers in the test sample based on the methylation sequencing results of the test sample and using the methylation level distribution pattern feature value. The biomarker screening unit is used to construct a linear classification model based on the tissue traceability feature vector and the cancer or health status of the sample to be tested, and to screen out Q biomarkers as biomarkers for subsequent cancer detection based on the average prediction accuracy of the linear classification model, where L is greater than or equal to Q, and L and Q are both integers greater than or equal to 1. The unit for calculating the tissue origin feature vector includes a subunit for calculating the likelihood value of a sequencing fragment in the test sample originating from each of the three source components: healthy plasma, normal tissue, and tumor tissue, and a subunit for determining the tissue origin ratio of the test sample based on the detection results within the range of each candidate biomarker. The subunit that calculates the likelihood of a sequencing fragment in the test sample originating from one of the three source components—healthy plasma, normal tissue, and tumor tissue—is used to calculate the likelihood of the sequencing fragment in the test sample originating from that specific source component. To determine the probability that a particular sequencing fragment in the sample to be tested originates from each of the aforementioned source components; Likelihood value The formula is as follows: in, For a certain source component A candidate marker methylation level distribution; and Let be the shape parameter of the beta distribution; For beta function; t is 1, 2, or 3, representing the three source components: healthy plasma, normal tissue, and tumor tissue, respectively. It is the methylation encoding form of a sequencing fragment in the sample to be tested. 0 indicates the unmethylated state, and 1 indicates the methylated state. The sub-unit for determining the tissue tracing ratio of the test sample based on the detection results within each marker range is used to determine the proportion of the test sample originating from a certain source component based on the detection results within each candidate marker range using the expectation-maximization algorithm.

20. The marker screening device according to claim 19, wherein, The candidate markers are obtained from known markers.

21. The biomarker screening device according to claim 20, wherein the candidate biomarkers are selected by the following method: the detection sites of the 450K methylation chip are merged into the same biomarker according to their proximity in the genome, and finally only biomarkers with no less than 3 detection sites are retained as candidate biomarkers.

22. The marker screening device according to claim 19, wherein, The methylation level distribution pattern feature value unit includes a sub-unit for acquiring healthy plasma samples, a sub-unit for acquiring normal tissues, a sub-unit for acquiring tumor tissues, and a methylation level distribution pattern feature value unit for calculating the methylation level data of each candidate biomarker. The healthy plasma sample acquisition subunit is used to acquire plasma from known healthy plasma samples; The subunit for obtaining normal tissue samples is used to obtain them from known normal tissue samples; The tumor tissue sample acquisition subunit is used to obtain tumor tissue samples from known tumor tissue samples.

23. The marker screening device according to claim 22, wherein, The healthy plasma sample acquisition subunit is used to acquire healthy plasma samples from the NCBI SRA database GSE164600 dataset.

24. The marker screening device according to claim 22, wherein, The subunit for obtaining normal tissue samples is used to obtain normal tissue samples from adjacent normal tissue samples of the LUSC and LUAD projects in the TCGA database.

25. The marker screening device according to claim 22, wherein, The normal tissue sample is a normal lung tissue sample.

26. The marker screening device according to claim 22, wherein, The tumor tissue sample acquisition subunit is used to obtain tumor samples from the LUSC and LUAD projects in the TCGA database.

27. The marker screening device according to claim 22, wherein, The tumor sample is a lung cancer tumor sample.

28. The marker screening device according to claim 19, wherein, The methylation level distribution pattern feature values ​​are represented by a beta distribution, denoted as: ,in and The shape parameter of this distribution is calculated using the following formula: in, It is the mean methylation level of a candidate biomarker in multiple samples from a specific group of three sample groups: healthy plasma samples, normal tissue samples, and tumor tissue samples. It is the variance of the methylation level of a certain marker among multiple samples in a certain group of samples.

29. The marker screening device according to claim 19, wherein, The expectation-maximization algorithm is as follows: Initialization: Randomly initialize healthy plasma ( ), normal lung tissue ( ) and tumor tissue ( The proportions are respectively , ( and )and ; E-Step: ; M-step: ; Repeat the E-step and M-step iterations until... and The values ​​converge, thus yielding the organizational origin feature vector; in, This represents the i-th sequencing fragment of a sample in the test sample, which has... The methylation coding sequence is in the form of N, and the total number of sequencing fragments is N. This indicates the tissue origin of the sequencing fragment.

30. The marker screening device according to any one of claims 19-29, wherein, The marker selection unit includes a linear classification model construction subunit and a linear classification model selection subunit. The linear classification model selection subunit is used to select Q markers based on the average prediction accuracy of the linear classification model.

31. The marker screening device according to claim 30, wherein, The average prediction accuracy of the selected linear classification model is >0.

7.

32. The marker screening device according to claim 30, wherein, We chose to use 5-fold cross-validation for the selection.

33. A cancer detection device based on methylation sequencing, comprising a biomarker screening device, a voting unit, and an interpretation unit as described in any one of claims 19-32; The voting unit is used to vote on unknown samples based on the Q selected markers; The interpretation unit is used to interpret unknown samples to determine whether the unknown sample is a cancer sample or a normal sample.

34. The cancer detection device according to claim 33, wherein, The voting will be conducted using the following formula: in, The organizational origin tracing results for various markers in unknown samples; For unknown samples in markers Organizational source traceability feature vectors; For the marker The linear classification model trained on it; The voting percentage for the "1" prediction given by the selected linear classification model with an average prediction accuracy greater than 0.7; The set threshold for the percentage of votes cast.

35. The cancer detection device according to claim 34, wherein, Provide a prediction result of "0" or "1" when the voting proportion does not exceed the threshold. If the result is positive, the sample is classified as "0" and is considered a normal sample; otherwise, it is classified as "1" and is considered a disease sample.

36. A biomarker screening system based on methylation sequencing, comprising a first processor and a first memory, the first processor and the first memory communicating via a first bus, the first memory storing first program instructions, the first processor executing the first program instructions, the first program instructions performing the biomarker screening method according to any one of claims 1-15.

37. A cancer detection system based on methylation sequencing, comprising a second processor and a second memory, the second processor and the second memory communicating via a second bus, the second memory storing second program instructions, the second processor executing the second program instructions, the second program instructions performing the cancer detection method according to any one of claims 16-18.

38. A non-transitory computer-readable storage medium for biomarker screening based on methylation sequencing, the non-transitory computer-readable storage medium storing computer instructions A, the computer instructions A causing a computer to perform the biomarker screening method according to any one of claims 1-15.

39. A non-transitory computer-readable storage medium for cancer detection based on methylation sequencing, the non-transitory computer-readable storage medium storing computer instructions B, the computer instructions B causing a computer to perform the cancer detection method according to any one of claims 16-18.

Citation Information

Patent Citations

  • Lung cancer diagnostic assay

    CN103245784A

  • Lung cancer DNA methylation molecular markers and application thereof in preparation of kit for early diagnosis of lung cancer

    CN112941180A