Affected tissue or organ tracing method and apparatus, storage medium, and program product
By acquiring and analyzing methylation data of healthy tissues or organs, an average methylation rate matrix is established. Unsupervised learning methods are used to screen for specific methylation sites or regions. Combined with the methylation data of the sample to be traced, the source can be traced, which solves the limitations and high costs of PET-CT and biopsy, and realizes low-cost, non-invasive diagnosis of disease-affected tissues or organs.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BOE TECHNOLOGY GROUP CO LTD
- Filing Date
- 2024-10-29
- Publication Date
- 2026-05-07
AI Technical Summary
When diagnosing diseased tissues or organs, current technologies such as PET-CT are costly and invasive, and the number of biopsy sites is limited, making it difficult to accurately determine the extent of organ or tissue involvement.
By acquiring methylation data from multiple healthy tissues or organs, an average methylation rate matrix is established to identify specific methylation sites or regions. Unsupervised learning methods are then used for cluster analysis to screen out specific methylation sites or regions. These sites or regions are then combined with the methylation data of the samples to be traced to achieve non-invasive diagnosis.
This provides a low-cost, non-invasive method to identify disease-affected tissues or organs at an early stage, reducing patient risk and improving diagnostic accuracy and efficiency.
Smart Images

Figure CN2024128247_07052026_PF_FP_ABST
Abstract
Description
Methods, devices, storage media, and program products for tracing the source of disease-affected tissues or organs. Technical Field
[0001] This disclosure relates to, but is not limited to, the field of biotechnology, and in particular to a method and apparatus, storage medium and program product for tracing the source of disease-affected tissues or organs. Background Technology
[0002] Some tumors are aggressive during their development, affecting certain organs or tissues. For example, diffuse large B-cell lymphoma (DLBCL) is an aggressive non-Hodgkin's lymphoma that can affect multiple tissues or organs, including lymph nodes, bone marrow, lungs, spleen, liver, and gastrointestinal tract. Currently, diagnosis primarily relies on positron emission tomography (PET-CT), often combined with biopsy for further investigation to determine if organs or tissues are affected by inflammation, metastasis, or other abnormalities. However, biopsies have limitations in sampling sites and are invasive, posing certain risks to patients; PET-CT also has drawbacks, such as high cost.
[0003] Summary of the Invention
[0004] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.
[0005] This disclosure provides a method for tracing the source of a disease-affected tissue or organ, including:
[0006] Acquire methylation data of multiple tissues or organs, wherein the methylation data of the tissues or organs are methylation data of healthy tissues or organs;
[0007] An average methylation rate matrix is constructed based on methylation data from multiple tissues or organs. The average methylation rate matrix includes the average methylation rate of each tissue or organ at multiple methylation sites or methylation regions.
[0008] Specific methylation sites or specific methylation regions are determined based on the average methylation rate matrix;
[0009] Acquire methylation data of the sample to be traced, and trace the source based on the average methylation rate of multiple tissues or organs at the specific methylation site or specific methylation region and the methylation data of the sample to be traced.
[0010] This disclosure also provides a device for tracing diseased tissues or organs, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of the method for tracing diseased tissues or organs according to any embodiment of this disclosure based on the instructions stored in the memory.
[0011] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for tracing the source of disease-affected tissues or organs as described in any embodiment of this disclosure.
[0012] This disclosure also provides a program product including instructions that, when executed by a computer, perform a method for tracing disease-affected tissues or organs as described in any embodiment of this disclosure.
[0013] This disclosure also provides a device for tracing the source of disease-affected tissues or organs, including a data acquisition module, a calculation module, a determination module, and a tracing module, wherein:
[0014] The data acquisition module is configured to acquire methylation data of multiple tissues or organs, wherein the methylation data of the tissues or organs are methylation data of healthy tissues or organs;
[0015] The calculation module is configured to establish an average methylation rate matrix based on methylation data from multiple tissues or organs, wherein the average methylation rate matrix includes the average methylation rate of each tissue or organ at multiple methylation sites or methylation regions.
[0016] The determining module is configured to determine specific methylation sites or specific methylation regions based on the average methylation rate matrix.
[0017] The tracing module is configured to acquire methylation data of the sample to be traced, and to trace the source based on the average methylation rate of multiple tissues or organs at the specific methylation site or specific methylation region and the methylation data of the sample to be traced.
[0018] Other features and advantages of this disclosure will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the disclosure. Other advantages of this disclosure may be realized and obtained by means of the methods described in the description and the accompanying drawings. Attached Figure Description
[0019] The accompanying drawings are used to provide an understanding of the technical solutions of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.
[0020] Figure 1 is a flowchart illustrating a method for tracing the source of a disease-affected tissue or organ according to an exemplary embodiment of this disclosure;
[0021] Figure 2 is a schematic diagram of the changes in elbow clustering analysis according to an exemplary embodiment of the present disclosure;
[0022] Figure 3 is a schematic diagram of the contribution ratio of organs affected by eight diseases in an exemplary embodiment of this disclosure;
[0023] Figure 4 is a schematic diagram of the structure of another disease-affected tissue or organ tracing device according to an exemplary embodiment of the present disclosure;
[0024] Figure 5 is a schematic diagram of another disease-affected tissue or organ tracing device according to an exemplary embodiment of this disclosure. Detailed Implementation
[0025] This disclosure describes several embodiments, but these descriptions are exemplary and not limiting, and it will be apparent to those skilled in the art that many more embodiments and implementations are possible within the scope of the embodiments described herein. Although many possible combinations of features are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with, or may replace, any feature or element of any other embodiment.
[0026] This disclosure includes and contemplates combinations of features and elements known to those skilled in the art. The embodiments, features, and elements disclosed in this disclosure may also be combined with any conventional features or elements to form a unique inventive scheme as defined by the claims. Any feature or element of any embodiment may also be combined with features or elements from other inventive schemes to form another unique inventive scheme as defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in this disclosure may be implemented individually or in any suitable combination. Therefore, the embodiments are not limited except by the limitations imposed by the appended claims and their equivalents. Furthermore, various modifications and changes may be made within the scope of the appended claims.
[0027] Furthermore, in describing representative embodiments, the specification may have presented methods and / or processes as a specific sequence of steps. However, the method or process should not be limited to the specific order of steps described herein, to the extent that the method or process does not depend on the specific order of steps described herein. As will be understood by those skilled in the art, other sequences of steps are also possible. Therefore, the specific order of steps set forth in the specification should not be construed as a limitation of the claims. Moreover, the claims relating to the method and / or process should not be limited to the steps performed in the order written, and those skilled in the art will readily understand that these orders can be varied and still remain within the spirit and scope of the embodiments disclosed herein.
[0028] As shown in Figure 1, this embodiment of the present disclosure provides a method for tracing the source of a disease-affected tissue or organ, including:
[0029] Step 101: Obtain methylation data of multiple tissues or organs, wherein the obtained methylation data of multiple tissues or organs are methylation data of healthy tissues or organs;
[0030] Step 102: Establish an average methylation rate matrix based on methylation data from multiple tissues or organs, wherein the average methylation rate matrix includes the average methylation rate of each tissue or organ at multiple methylation sites or methylation regions.
[0031] Step 103: Determine specific methylation sites or specific methylation regions based on the average methylation rate matrix;
[0032] Step 104: Obtain the methylation data of the sample to be traced, and trace the source based on the average methylation rate of multiple tissues or organs at the specific methylation site or specific methylation region and the methylation data of the sample to be traced.
[0033] In this disclosure, "affected" refers to the influence of a disease or condition on a specific organ, tissue, or system. The method for tracing the affected tissue or organ in this disclosure involves acquiring methylation data from multiple tissues or organs, establishing an average methylation rate matrix based on this data, determining specific methylation sites or regions based on the matrix, and tracing the source based on the average methylation rate of multiple tissues or organs at these specific methylation sites or regions, as well as the methylation data of the sample to be traced. This method can identify affected tissues or organs from the methylation sequencing data of circulating free DNA (cfDNA) in the patient's peripheral blood. It is convenient to sample, highly patient-friendly, and cost-effective.
[0034] In some exemplary embodiments, in step 101, multiple tissues or organs may include: bone, brain, colon, kidney, liver, lung, pancreas, stomach; however, this disclosure is not limiting in this regard.
[0035] In some exemplary embodiments, in step 101, the methylation data of multiple tissues or organs can be methylation sequencing data or microarray data.
[0036] In this embodiment of the disclosure, the methylation data of multiple tissues or organs obtained can be methylation sequencing data or microarray data obtained from public databases. The sequencing library type of the methylation sequencing data can be various methylation sequencing libraries, including simplified genome methylation sequencing (RRBS), whole genome DNA methylation sequencing (WGBS), TET-assisted sulfite sequencing (TAB-Seq), etc.; the microarray data can be 450K or 850K DNA methylation chip data, however, this disclosure does not limit it.
[0037] In this embodiment, the acquired methylation sequencing data undergoes quality control and cleaning before being aligned to a human reference genome to obtain alignment results and extract the methylation rate of methylation sites. For the acquired microarray data (methylation chip data), the methylation rate of methylation sites can be directly obtained, and genomic coordinate information can be obtained through chip database conversion. This embodiment, by utilizing methylation sequencing data and microarray data from tissues or organs, can minimize noise introduced by the involvement of other tissues or organs, ensuring the accuracy of subsequent identification of tissue or organ methylation-specific regions.
[0038] For example, taking the methylation microarray data of tissues or organs obtained from the public database MethBank (a comprehensive DNA methylation database) as an example, the microarray type can be 450K or 850K. The following uses 450K microarray data as an example. Assuming that multiple tissue or organ types are bone, brain, colon, kidney, liver, lung, pancreas, and stomach, methylation data similar to that shown in Table 1 will be obtained for each tissue or organ.
[0039] Table 1
[0040] In Table 1, the first column is the probe number, followed by the sample number and the methylation rate measured for each sample under the probe number. The probe number can be converted into genomic coordinates using a 450k table, as illustrated in Table 2.
[0041] Table 2
[0042] In Table 2, the first column is the chromosome number, the second column is the genomic coordinates, and the third column is the probe number.
[0043] In some exemplary embodiments, step 102, establishing an average methylation rate matrix based on methylation data from multiple tissues or organs, includes:
[0044] In the methylation data of each tissue or organ, delete methylation sites or methylation regions where the methylation data is empty;
[0045] Find the intersection of methylation sites or methylation regions in methylation data from multiple tissues or organs;
[0046] For each tissue or organ, calculate the average methylation rate of multiple samples at each methylation site or methylation region obtained by taking the intersection;
[0047] An average methylation rate matrix is generated based on the calculated average methylation rate of multiple tissues or organs at multiple methylation sites or methylation regions.
[0048] In this embodiment of the disclosure, since the methylation data of some probes for each tissue or organ may be missing, as shown in Table 3, the methylation data of probes cg01671070 and cg01737421 are both missing (NA).
[0049] Table 3
[0050] Therefore, probe numbers containing NA sites in the methylation data of all tissues or organs were excluded. The intersection of the remaining probe numbers in all tissues or organs was calculated. Then, the methylation site data of all tissues or organs under the intersection probe numbers were extracted and concatenated column by column to form a methylation rate matrix, as shown in Table 4.
[0051] Table 4
[0052] In Table 4, each column represents the methylation data of a sample at different methylation sites. The type of tissue to which each sample belongs is indicated by “XXX_” before the sample number.
[0053] In this embodiment of the disclosure, when the acquired methylation data of multiple tissues or organs is methylation sequencing data, the methylation site data of multiple tissues or organs is divided into n methylation regions according to a pre-set window length W, where n is a natural number greater than 1. The methylation rate of each sample in each methylation region of the multiple tissue or organ methylation data is calculated, resulting in a window of length W as shown in Table 5. m i(where i is the tissue or organ number, m) i The methylation rate matrix M1, with a width of n (n columns representing n methylation regions), is shown in Table 5. (This represents the number of samples from a specific tissue or organ.) t The organization or organ is numbered, where t is a positive integer greater than or equal to 1, and T tm This is the sample number in the methylation data of a certain tissue or organ, where m is also a positive integer greater than or equal to 1 (note that in Table 5, the length is the number of samples and the width is the number of methylation sites or methylation regions; in Table 4, the length is the number of methylation sites or methylation regions and the width is the number of samples, that is, the methylation rate matrix in Table 5 and the methylation rate matrix in Table 4 are transposes of each other).
[0054] Table 5
[0055] For each tissue or organ, calculate the average methylation rate of each of the n methylation regions across all samples. That is, for each tissue or organ, there exists an average methylation rate value across n methylation regions. The specific calculation method is as follows:
[0056] The formula for calculating the methylation rate (MR) of a single methylated region in a single sample is as follows:
[0057] Where, N C The number of methylated C bases in the region, N T This represents the number of unmethylated T bases within the region.
[0058] The formula for calculating the average methylation rate (TMR) of a single methylated region in each tissue or organ is as follows:
[0059] Where m is the number of samples in each tissue or organ, MR j represents the methylation rate of a single methylated region in the j-th sample from each tissue or organ.
[0060] After the above steps, the average methylation rate matrix M2 of the methylation region with length t and width n is obtained as shown in Table 6.
[0061] Table 6
[0062] When the methylation data of multiple tissues or organs is obtained as microarray data, the number of sites S is directly counted as n. Other steps are similar to those for methylation sequencing data, and a matrix of length as shown in Table 5 can be formed. m iThe matrix M1 has a methylation rate of n width (n columns represent n methylation sites) and the matrix M2 has an average methylation rate of methylation sites of length t and width n, as shown in Table 6.
[0063] For microarray data, the methylation rate of a single methylation site in a single sample is assumed to be MR. The average methylation rate (TMR) of a single methylation site for each tissue or organ is calculated using the following formula:
[0064] Where m is the number of samples in each tissue or organ, MR j The methylation rate of a single methylation site in the j-th sample of each tissue or organ.
[0065] In related technologies, obtaining specific methylation sites or specific methylation regions of tissues or organs requires setting specific thresholds. For example, by judging whether the methylation level of a candidate methylation site or candidate methylation region of a certain tissue is higher or lower than the average methylation level of all remaining tissues by the standard deviation k, k needs to be set manually. A fixed k value may not be suitable for different datasets or tissue types.
[0066] This disclosure embodiment uses an unsupervised learning method to cluster methylation data. Without manually setting thresholds, it automatically identifies specific methylation sites or methylation regions. Clustering can group sites or regions with high similarity between tissues together, thereby enabling more systematic analysis of methylation data. Furthermore, since it does not judge whether each site or region is a specific methylation site or region individually, but rather compares whether a cluster of sites or regions is a specific methylation site or region, it can discover epigenomic information patterns hidden in the methylation data.
[0067] In some exemplary embodiments, in step 103, determining specific methylation sites or specific methylation regions based on the average methylation rate matrix includes:
[0068] Cluster analysis is performed on methylation sites or methylation regions in the average methylation rate matrix to obtain the number of clusters;
[0069] For each tissue or organ in each cluster, determine the significant difference between the average methylation rate of the current tissue or organ and the mean average methylation rate of all remaining tissues or organs, and calculate the number of tissues or organs with significant differences in each cluster.
[0070] All methylation sites or methylation regions within a cluster where the number of tissues or organs with significant differences is greater than or equal to a preset threshold are identified as specific methylation sites or specific methylation regions.
[0071] For example, the clustering analysis method can be K-means; however, this disclosure is not limiting in this regard.
[0072] In this embodiment of the disclosure, for the average methylation rate matrix M2, specific methylation sites or specific methylation regions can be identified using unsupervised learning. Taking cluster analysis as an example, the number of clusters generated by the cluster analysis is extracted. For each tissue or organ in each cluster, the significant difference between the average methylation rate of the current tissue or organ and the mean of the average methylation rates of all remaining tissues or organs is calculated. The number of tissues or organs in each cluster that meet the calculated significant difference threshold is obtained. If the number of tissues or organs in a cluster that meet the calculated significant difference threshold is greater than a preset threshold, the methylation site or methylation region of that cluster is considered a specific methylation site or specific methylation region. This process is repeated to obtain the specific methylation sites or specific methylation regions in all clusters.
[0073] This embodiment utilizes an unsupervised learning method to cluster methylation sites or methylation regions. From these clusters, a significance test is used to determine the differences between the target tissue or organ and other tissues or organs, thereby identifying specific methylation sites or regions within the tissue or organ. This embodiment uses an unsupervised learning method to screen for specific methylation sites or regions in tissues or organs without setting a fixed threshold, making it adaptable to different datasets and tissue or organ types, thus exhibiting good universality.
[0074] Taking the 450K microarray data of the aforementioned eight tissues or organs as an example, the average methylation rate of the eight tissues or organs at each methylation site S is obtained according to Formula 3, forming an average methylation rate matrix M2. Using cluster analysis, the average methylation rate matrix M2 is clustered column-wise to obtain multiple clusters. Each cluster includes at least one column of data from the average methylation rate matrix M2. For example, one cluster may include columns S1, S2, and S3 of the average methylation rate matrix M2, and another cluster may include columns S6 and S7 of the average methylation rate matrix M2.
[0075] In some exemplary embodiments, the method further includes: using the elbow method to filter the number of clusters to obtain the optimal number of clusters.
[0076] This disclosure embodiment can utilize the Elbow Method to select the optimal number of clusters. An exemplary Elbow Method clustering analysis variation diagram is shown in Figure 2, where the vertical axis represents the Total Within-cluster Sum of Squares (WSS), which is the sum of the squared distances from all data points within a cluster to their respective cluster centers. The horizontal axis represents the number of clusters, set to a range of 1 to 100. The WSS value in the figure decreases as the number of clusters increases. When the number of clusters increases, the WSS typically decreases because more clusters can better fit the data. However, the rate of decrease in WSS gradually diminishes. When an "elbow" appears in the figure, indicating a very slow decrease in the rate of WSS decrease, this point is often considered an indicator for selecting the optimal number of clusters. In practical applications, the rate of change of WSS can be calculated first, and then the minimum value of the rate of change can be calculated to determine the location of the elbow point. Calculations show that the optimal number of clusters in this example is 91.
[0077] In some exemplary implementations, the determined significance difference can be a significance p-value or a q-value (corrected p-value).
[0078] In some exemplary embodiments, determining a significant difference between the average methylation rate of each tissue or organ in a cluster and the mean average methylation rate of all remaining tissues or organs includes:
[0079] For each tissue or organ, perform the following procedures:
[0080] Calculate the mean average methylation rate of all remaining tissues or organs other than the currently targeted tissue or organ, and calculate the significance p-value between the mean average methylation rate of the currently targeted tissue or organ and the mean average methylation rate of all remaining tissues or organs. Compare the significance p-value with a preset significance difference threshold. If the significance p-value is less than the preset significance difference threshold, it is determined that there is a significant difference between the mean average methylation rate of the currently targeted tissue or organ and the mean average methylation rate of all other tissues or organs.
[0081] In some exemplary embodiments, determining a significant difference between the average methylation rate of each tissue or organ in a cluster and the mean average methylation rate of all remaining tissues or organs includes:
[0082] For each tissue or organ, perform the following procedures:
[0083] Calculate the mean average methylation rate of all tissues or organs other than the currently targeted tissue or organ, and calculate the significance p-value between the mean average methylation rate of the currently targeted tissue or organ and the mean average methylation rate of all other tissues or organs. Correct the calculated significance p-value to obtain the significance q-value. Compare the significance q-value with the preset significance difference threshold. If the significance q-value is less than the preset significance difference threshold, it is determined that there is a significant difference between the mean average methylation rate of the currently targeted tissue or organ and the mean average methylation rate of all other tissues or organs.
[0084] In some exemplary embodiments, the preset quantity threshold can be the number of multiple tissues or organs. However, this disclosure does not limit this, and the preset quantity threshold can be t-1 or t-2, etc., where t is the number of multiple tissues or organs.
[0085] Taking a predetermined significance level as the q-value and a preset threshold equal to the number of tissues or organs as an example, cluster analysis is performed using the optimal number of clusters determined above. A T-test is used to calculate the significance p-value between the mean methylation rate of each tissue or organ's methylation sites (or methylation regions) and the mean methylation rate of the remaining tissues or organs. Subsequently, the Bonferroni method is used to correct the significance p-value to obtain the significance q-value, retaining the mean methylation rate of each tissue or organ relative to the mean methylation rate of other tissues or organs. Clusters with significantly different values are identified. For example, in this embodiment, the number of multiple tissues or organs is 8. If the number of tissues or organs that meet the significance q value less than the preset significance difference threshold is 8, then all methylation sites (or methylation regions) under this cluster are considered to be specific methylation sites (or specific methylation regions) of that tissue or organ. Otherwise (i.e., the number of tissues or organs that meet the significance q value less than the preset significance difference threshold is less than 8), then all methylation sites (or methylation regions) under this cluster are considered not to be specific methylation sites (or specific methylation regions) of that tissue or organ. This process is repeated to obtain the specific methylation sites (or specific methylation regions) under all clusters. The union of the specific methylation sites (or specific methylation regions) under all clusters is then used to obtain the final specific methylation sites (or specific methylation regions) of the tissue or organ.
[0086] In some exemplary embodiments, the preset significance difference threshold may be 0.01 or 0.05. For example, the preset significance difference threshold may be 0.01; however, this disclosure is not limited thereto.
[0087] In some exemplary embodiments, in step 104, the sample to be traced can be a tumor patient sample. However, this disclosure is not limiting in this regard, and the sample to be traced can also be a patient sample from any other type of disease.
[0088] In some exemplary embodiments, the methylation data of the sample to be traced is cfDNA methylation sequencing data. This disclosure identifies affected tissues or organs by analyzing cfDNA methylation sequencing data from a patient's peripheral blood, which is convenient to sample, highly patient-friendly, and low-cost.
[0089] In some exemplary embodiments, the method further includes:
[0090] Perform data quality control and data filtering on the methylation data of the samples to be traced.
[0091] The methylation data of the samples to be traced after data quality control and data filtering were compared with the human reference genome to obtain the methylation rate of the samples to be traced at multiple methylation sites or methylation regions.
[0092] One or more specific methylation sites or specific methylation regions are deleted from multiple specific methylation sites or specific methylation regions, wherein the methylation rate of the sample to be traced in the deleted specific methylation sites or specific methylation regions is empty.
[0093] For example, for the cfDNA methylation sequencing data of the sample to be traced generated after sequencing, data quality control and data filtering are performed after obtaining the sequencing data. The filtered sequencing data is then aligned to the human reference genome hg19, and the methylation data of each methylation site is extracted based on the alignment results. After obtaining the methylation data of each methylation site in the sample to be traced, the intersection of the finally obtained specific methylation site and all methylation sites in the sample to be traced is taken. That is, the methylation site after the intersection exists in both the specific methylation site and the sample to be traced, and there is no missing methylation rate.
[0094] An example of the obtained dataset is shown in Table 7.
[0095] Table 7
[0096] In Table 7, the first two columns are the genomic coordinates of the intersection sites, columns 3 to 10 are the methylation rates of eight tissues or organs (bone, brain, colon, kidney, liver, lung, pancreas, and stomach), and the last column, freqc, is the methylation rate of the sample to be traced at specific methylation sites.
[0097] In some exemplary embodiments, step 104 involves tracing the source based on the average methylation rate of multiple tissues or organs at specific methylation sites or specific methylation regions and the methylation data of the sample to be traced, including:
[0098] Establish an objective function and constraints. The objective function includes the sum of the squares of the differences between the product of the probability of involvement of multiple tissues or organs and the average methylation rate of their respective tissues or organs at all specific methylation sites or regions and the methylation rate of the sample to be traced. The constraints include that the probability of involvement of each tissue or organ is greater than or equal to 0 and the sum of the probabilities of involvement of multiple tissues or organs is 1.
[0099] Based on the set constraints, calculate the probability of each tissue or organ being affected when the objective function reaches its minimum value.
[0100] In some exemplary implementations, the objective function can be: Where s is the number of specific methylation sites or specific methylation regions, t is the number of multiple tissues or organs, and T is the number of tissues or organs. i,j p represents the average methylation rate of the i-th specific methylation site or specific methylation region in the j-th tissue or organ. j Let Me be the probability of involvement of the j-th tissue or organ. i The methylation rate of the sample to be traced is the methylation rate at the i-th specific methylation site or specific methylation region.
[0101] Compared to related technologies that use PET-CT to detect affected tissues, the embodiments of this disclosure use a quadratic programming method to consider the tissue or organ origin at the molecular level. This method has a relatively lower detection limit and can identify possible tissue or organ involvement earlier, facilitating early intervention in clinical practice and thus avoiding cancer metastasis and complications such as inflammation in cancer patients to some extent.
[0102] In this embodiment of the disclosure, the objective function of the quadratic programming is set as follows: for each specific methylation site (or specific methylation region), the difference between the sum of the products of each tissue or organ's contribution ratio (i.e., the probability of involvement) and the corresponding average methylation rate of the tissue or organ, and the average methylation rate of the sample to be traced at the specific methylation site (or specific methylation region), is used as the objective function. Since the difference may be positive or negative, as long as the difference between the sum of the products of each specific methylation site (or specific methylation region) and the methylation rate of the sample to be traced is minimized, the difference is squared and the results of each specific methylation site (or specific methylation region) are accumulated. The minimum value is the final optimized result of the objective function. In addition, to ensure that extreme cases do not occur, such as the contribution ratio of a certain tissue or organ being too high or too low, constraints are set to ensure that the sum of the contribution ratios is 1 and the contribution ratio of each tissue or organ is greater than or equal to 0.
[0103] Assuming the number of tissues or organs is t, and setting the same interval range as the specific methylation sites (or specific methylation regions) R mentioned above, the methylation level Me of the sample to be traced is obtained using the methylation sequencing data of the sample to be traced within these intervals. Assuming the contribution ratio of cfDNA from t tissues or organs is set to p1 to p2... t Using quadratic programming, the objective function is defined as:
[0104] Where s is the total number of specific methylation sites (or specific methylation regions) R determined in the above steps, and T i,j The average methylation rate is defined for each specific methylation site (or specific methylation region) in each tissue or organ obtained in the above steps. Furthermore, the following constraints are defined:
[0105] Let the objective function be f(p1,p2,…,p t When p1 reaches the minimum value, p... t The value of is taken as the contribution ratio of t tissues or organs in each sample to cfDNA.
[0106] In some exemplary embodiments, the method further includes: sorting the severity of involvement of the affected tissues or organs according to the calculated probability of involvement of each tissue or organ, and / or generating an involvement result display diagram, which includes the probability of involvement of each tissue or organ.
[0107] This embodiment of the disclosure assists in diagnosis by sorting the contribution ratio of each sample to be traced in t tissues or organs in descending order, and identifying the tissues or organs with higher probability values of involvement as possible affected sites.
[0108] Taking the dataset in Table 7 above as an example, the proportions of cfDNA sources from bone, brain, colon, kidney, liver, lung, pancreas, and stomach are set as p1 to p8, respectively. Using the objective function of Formula 4 and the constraints of Formula 5, the values of p1 to p8 when f(p1,p2,…,p8) reaches its minimum are obtained. The values of p1 to p8 are sorted from largest to smallest to represent the severity of involvement of the eight affected tissues or organs in the patient to be traced. The results can be illustrated by the contribution ratio diagram shown in Figure 3. As can be seen from Figure 3, the organs involved in the sample to be traced are bone and stomach, with bone showing the most significant involvement.
[0109] The method for tracing the source of disease-affected tissues or organs in this disclosure involves acquiring methylation data from multiple tissues or organs, analyzing and obtaining tissue- or organ-specific methylation sites or regions, and using the methylation data of the specific methylation sites or regions and the methylation data of the sample to be traced to obtain information on the affected tissues or organs of the sample. This method can be used as an auxiliary diagnostic tool for clinical diagnosis of affected sites.
[0110] As shown in Figure 4, this embodiment of the present disclosure also provides a device for tracing the source of disease-affected tissues or organs, including a data acquisition module 410, a calculation module 420, a determination module 430, and a tracing module 440, wherein:
[0111] The data acquisition module 410 is configured to acquire methylation data of multiple tissues or organs, wherein the acquired methylation data of multiple tissues or organs are methylation data of healthy tissues or organs.
[0112] The calculation module 420 is configured to build an average methylation rate matrix based on methylation data from multiple tissues or organs, wherein the average methylation rate matrix includes the average methylation rate of each tissue or organ at multiple methylation sites or methylation regions.
[0113] The determination module 430 is configured to determine specific methylation sites or specific methylation regions based on the average methylation rate matrix.
[0114] The tracing module 440 is configured to acquire methylation data of the sample to be traced and to trace the source based on the average methylation rate of multiple tissues or organs at the specific methylation site or specific methylation region and the methylation data of the sample to be traced.
[0115] In some exemplary embodiments, the tracing module 440 performs tracing based on the average methylation rate of multiple tissues or organs at the specific methylation site or specific methylation region and the methylation data of the sample to be traced, including:
[0116] Establish objective function and constraints. The objective function includes the sum of squares of the difference between the product of the probability of involvement of multiple tissues or organs and the average methylation rate of their respective tissues or organs at all specific methylation sites or regions and the methylation rate of the sample to be traced. The constraints include that the probability of involvement of each tissue or organ is greater than or equal to 0 and the sum of the probabilities of involvement of multiple tissues or organs is 1.
[0117] Based on the set constraints, calculate the probability of each tissue or organ being affected when the objective function reaches its minimum value.
[0118] In some exemplary implementations, the objective function is: Where s is the number of specific methylation sites or specific methylation regions, t is the number of multiple tissues or organs, and T i,j p represents the average methylation rate of the i-th specific methylation site or specific methylation region in the j-th tissue or organ. j Let Me be the probability of involvement of the j-th tissue or organ. i The methylation rate of the sample to be traced is the methylation rate at the i-th specific methylation site or specific methylation region.
[0119] In some exemplary embodiments, the methylation data of the sample to be traced is cfDNA methylation sequencing data.
[0120] In some exemplary embodiments, the traceability module 440 is further configured to:
[0121] The methylation data of the samples to be traced are subjected to data quality control and data filtering.
[0122] The methylation data of the sample to be traced, after data quality control and data filtering, is compared to the human reference genome to obtain the methylation rate of the sample to be traced at multiple methylation sites or methylation regions.
[0123] One or more of the specific methylation sites or specific methylation regions are deleted from the plurality of the specific methylation sites or specific methylation regions, wherein the methylation rate of the sample to be traced is empty at the deleted specific methylation sites or specific methylation regions.
[0124] In some exemplary embodiments, the determining module 430 determines specific methylation sites or specific methylation regions based on the average methylation rate matrix, including:
[0125] Cluster analysis is performed on the methylation sites or methylation regions of the average methylation rate matrix to obtain the number of clusters;
[0126] For each tissue or organ in each cluster, determine the significant difference between the average methylation rate of the current tissue or organ and the mean average methylation rate of all remaining tissues or organs, and calculate the number of tissues or organs with significant differences in each cluster.
[0127] All methylation sites or methylation regions within a cluster where the number of tissues or organs with significant differences is greater than or equal to a preset threshold are identified as specific methylation sites or specific methylation regions.
[0128] In some exemplary embodiments, the determining module 430 is also configured to filter the number of clusters using the elbow method to obtain the optimal number of clusters.
[0129] In some exemplary embodiments, the preset quantity threshold is equal to the number of multiple tissues or organs.
[0130] In some exemplary embodiments, the determining module 430 determines a significant difference between the average methylation rate of each of the tissues or organs and the mean average methylation rate of all remaining tissues or organs, including:
[0131] For each of the aforementioned tissues or organs, the following operations shall be performed:
[0132] Calculate the average methylation rate of all remaining tissues or organs other than the currently targeted tissue or organ, and calculate the significance p-value between the average methylation rate of the currently targeted tissue or organ and the average methylation rate of all remaining tissues or organs. Compare the significance p-value with a preset significance difference threshold. If the significance p-value is less than the preset significance difference threshold, it is determined that there is a significant difference between the average methylation rate of the currently targeted tissue or organ and the average methylation rate of all other tissues or organs.
[0133] In some exemplary embodiments, the determining module 430 determines a significant difference between the average methylation rate of each of the tissues or organs and the mean average methylation rate of all remaining tissues or organs, including:
[0134] For each of the aforementioned tissues or organs, the following operations shall be performed:
[0135] Calculate the mean average methylation rate of all remaining tissues or organs other than the currently targeted tissue or organ, and calculate the significance p-value between the mean average methylation rate of the currently targeted tissue or organ and the mean average methylation rate of all remaining tissues or organs. Correct the calculated significance p-value to obtain a significance q-value. Compare the significance q-value with a preset significance difference threshold. If the significance q-value is less than the preset significance difference threshold, it is determined that there is a significant difference between the mean average methylation rate of the currently targeted tissue or organ and the mean average methylation rate of all other tissues or organs.
[0136] In some exemplary embodiments, multiple tissues or organs may include: bone, brain, colon, kidney, liver, lung, pancreas, and stomach.
[0137] In some exemplary embodiments, the methylation data of multiple tissues or organs are methylation sequencing data or microarray data.
[0138] In some exemplary embodiments, the calculation module 420 establishes an average methylation rate matrix based on methylation data from multiple tissues or organs, including:
[0139] In the methylation data of each tissue or organ, delete methylation sites or methylation regions where the methylation data is empty;
[0140] Find the intersection of methylation sites or methylation regions in methylation data from multiple tissues or organs;
[0141] For each tissue or organ, calculate the average methylation rate of multiple samples at each methylation site or methylation region obtained by taking the intersection;
[0142] An average methylation rate matrix is generated based on the calculated average methylation rate of multiple tissues or organs at multiple methylation sites or methylation regions.
[0143] This disclosure also provides a device for tracing diseased tissues or organs, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the method for tracing diseased tissues or organs as described in any embodiment of this disclosure based on the instructions stored in the memory.
[0144] As shown in Figure 5, in one example, the source tracing device for diseased tissues or organs may include: a processor 510, a memory 520, a bus system 530, and a transceiver 540. The processor 510, the memory 520, and the transceiver 540 are connected via the bus system 530. The memory 520 is used to store instructions, and the processor 510 is used to execute the instructions stored in the memory 520 to control the transceiver 540 to send and receive signals. Specifically, transceiver 540, under the control of processor 510, can acquire methylation data of multiple tissues or organs and methylation data of the sample to be traced, wherein the methylation data of the tissues or organs are methylation data of healthy tissues or organs; processor 510 establishes an average methylation rate matrix based on the methylation data of multiple tissues or organs, wherein the average methylation rate matrix includes the average methylation rate of each tissue or organ at multiple methylation sites or methylation regions; determines specific methylation sites or specific methylation regions based on the average methylation rate matrix; and performs traceability based on the average methylation rate of multiple tissues or organs at the specific methylation sites or specific methylation regions and the methylation data of the sample to be traced.
[0145] It should be understood that processor 510 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0146] Memory 520 may include read-only memory and random access memory, and provides instructions and data to processor 510. A portion of memory 520 may also include non-volatile random access memory. For example, memory 520 may also store device type information.
[0147] In addition to the data bus, the bus system 530 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 530 in Figure 5.
[0148] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the processor 510 or through software instructions. That is, the method steps of this embodiment can be executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in memory 520, and the processor 510 reads information from memory 520 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, further details are omitted here.
[0149] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method for tracing disease-affected tissues or organs as described in any embodiment of this disclosure. The method for tracing disease-affected tissues or organs driven by executing executable instructions is substantially the same as the method for tracing disease-affected tissues or organs provided in the above embodiments of this disclosure, and will not be described in detail here.
[0150] In some possible implementations, various aspects of the method for tracing disease-affected tissues or organs provided in this disclosure can also be implemented in the form of a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps in the method for tracing disease-affected tissues or organs according to various exemplary embodiments of this disclosure as described above. For example, the computer device can perform the method for tracing disease-affected tissues or organs described in the embodiments of this disclosure.
[0151] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0152] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0153] It should be noted that the above embodiments or implementation methods are merely exemplary and not restrictive. Therefore, this disclosure is not limited to the content specifically shown and described herein. Various modifications, substitutions, or omissions can be made to the form and details of the implementations without departing from the scope of this disclosure.
Claims
1. A method for tracing the origin of a disease-affected tissue or organ, comprising: Acquire methylation data of multiple tissues or organs, wherein the methylation data of the tissues or organs are methylation data of healthy tissues or organs; An average methylation rate matrix is established based on methylation data from multiple tissues or organs, wherein the average methylation rate matrix includes the average methylation rate of each tissue or organ at multiple methylation sites or methylation regions; Specific methylation sites or specific methylation regions are determined based on the average methylation rate matrix. Acquire methylation data of the sample to be traced, and trace the source based on the average methylation rate of multiple tissues or organs at the specific methylation site or specific methylation region and the methylation data of the sample to be traced.
2. The method according to claim 1, wherein, The tracing based on the average methylation rate of multiple tissues or organs at the specific methylation site or specific methylation region and the methylation data of the sample to be traced includes: Establish an objective function and constraints. The objective function includes the sum of the squares of the differences between the product of the probability of involvement of multiple tissues or organs and the average methylation rate of their respective tissues or organs at all specific methylation sites or regions and the methylation rate of the sample to be traced. The constraints include that the probability of involvement of each tissue or organ is greater than or equal to 0 and the sum of the probabilities of involvement of multiple tissues or organs is 1. Based on the set constraints, calculate the probability of each tissue or organ being affected when the objective function reaches its minimum value.
3. The method according to claim 2, wherein, The objective function is: Where s is the number of specific methylation sites or specific methylation regions, t is the number of multiple tissues or organs, and T i,j p represents the average methylation rate of the i-th specific methylation site or specific methylation region in the j-th tissue or organ. j Let Me be the probability of involvement of the j-th tissue or organ. i The methylation rate of the sample to be traced is the methylation rate at the i-th specific methylation site or specific methylation region.
4. The method according to claim 2, wherein, The method further includes: sorting the severity of the affected tissues or organs according to the calculated probability of each tissue or organ being affected, and / or generating an affected result display diagram, wherein the affected result display diagram includes the probability of each tissue or organ being affected.
5. The method according to claim 1, wherein, The methylation data of the sample to be traced is plasma free DNA methylation sequencing data.
6. The method according to claim 1, wherein, The method further includes: The methylation data of the samples to be traced are subjected to data quality control and data filtering. The methylation data of the sample to be traced, after data quality control and data filtering, is compared to the human reference genome to obtain the methylation rate of the sample to be traced at multiple methylation sites or methylation regions. One or more of the specific methylation sites or specific methylation regions are deleted from the plurality of the specific methylation sites or specific methylation regions, wherein the methylation rate of the sample to be traced is empty at the deleted specific methylation sites or specific methylation regions.
7. The method according to claim 1, wherein, The step of determining specific methylation sites or specific methylation regions based on the average methylation rate matrix includes: Cluster analysis is performed on the methylation sites or methylation regions of the average methylation rate matrix to obtain the number of clusters; For each tissue or organ in each cluster, determine the significant difference between the average methylation rate of the current tissue or organ and the mean average methylation rate of all remaining tissues or organs, and calculate the number of tissues or organs with significant differences in each cluster. All methylation sites or methylation regions within a cluster where the number of tissues or organs with significant differences is greater than or equal to a preset threshold are identified as specific methylation sites or specific methylation regions.
8. The method according to claim 7, wherein, The method further includes: using the elbow method to filter the number of clusters to obtain the optimal number of clusters.
9. The method according to claim 7, wherein, The preset quantity threshold is equal to the number of multiple tissues or organs.
10. The method according to claim 7, wherein, The determination of the significant difference between the average methylation rate of each of the tissues or organs and the mean average methylation rate of all remaining tissues or organs includes: For each of the aforementioned tissues or organs, the following operations shall be performed: Calculate the average methylation rate of all remaining tissues or organs other than the currently targeted tissue or organ. The significance p-value between the average methylation rate of the currently targeted tissue or organ and the mean average methylation rate of all other tissues or organs is calculated. The significance p-value is compared with a preset significance difference threshold. If the significance p-value is less than the preset significance difference threshold, it is determined that there is a significant difference between the average methylation rate of the currently targeted tissue or organ and the mean average methylation rate of all other tissues or organs.
11. The method according to claim 7, wherein, The determination of the significant difference between the average methylation rate of each of the tissues or organs and the mean average methylation rate of all remaining tissues or organs includes: For each of the aforementioned tissues or organs, the following operations shall be performed: Calculate the mean average methylation rate of all remaining tissues or organs other than the currently targeted tissue or organ, and calculate the significance p-value between the mean average methylation rate of the currently targeted tissue or organ and the mean average methylation rate of all remaining tissues or organs. Correct the calculated significance p-value to obtain a significance q-value. Compare the significance q-value with a preset significance difference threshold. If the significance q-value is less than the preset significance difference threshold, it is determined that there is a significant difference between the mean average methylation rate of the currently targeted tissue or organ and the mean average methylation rate of all other tissues or organs.
12. The method according to claim 10 or 11, wherein, The preset significance threshold can be 0.01 or 0.
05.
13. The method according to claim 1, wherein, The aforementioned tissues or organs include: bone, brain, colon, kidney, liver, lung, pancreas, and stomach.
14. The method according to claim 1, wherein, The methylation data of the multiple tissues or organs are methylation sequencing data or microarray data.
15. The method according to claim 1, wherein, The step of establishing an average methylation rate matrix based on methylation data from multiple tissues or organs includes: In the methylation data of each tissue or organ, delete methylation sites or methylation regions where the methylation data is empty; Find the intersection of methylation sites or methylation regions in methylation data from multiple tissues or organs; For each tissue or organ, calculate the average methylation rate of multiple samples at each methylation site or methylation region obtained by taking the intersection; Based on the calculated average methylation rate of multiple tissues or organs at multiple methylation sites or methylation regions Generate the average methylation rate matrix.
16. A source tracing device for diseased tissues or organs, comprising a memory; and a processor connected to the memory, the memory for storing instructions, the processor being configured to perform the steps of the source tracing method for diseased tissues or organs as claimed in any one of claims 1 to 15 based on the instructions stored in the memory.
17. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for tracing the source of disease-affected tissues or organs as described in any one of claims 1 to 15.
18. A computer program product comprising instructions that, when executed by a computer, perform a method for tracing the source of a disease-affected tissue or organ as described in any one of claims 1 to 15.
19. A device for tracing the source of a disease-affected tissue or organ, comprising a data acquisition module, a calculation module, a determination module, and a tracing module, wherein: The data acquisition module is configured to acquire methylation data of multiple tissues or organs, wherein the methylation data of the tissues or organs are methylation data of healthy tissues or organs; The calculation module is configured to establish an average methylation rate matrix based on methylation data from multiple tissues or organs, wherein the average methylation rate matrix includes the average methylation rate of each tissue or organ at multiple methylation sites or methylation regions. The determining module is configured to determine specific methylation sites or specific methylation regions based on the average methylation rate matrix. The tracing module is configured to acquire methylation data of the sample to be traced, and to trace the source based on the average methylation rate of multiple tissues or organs at the specific methylation site or specific methylation region and the methylation data of the sample to be traced.