Tissue tracing method for cell-free DNA
By analyzing the TSS region coverage and depth of tissue-specific expressed gene sets and calculating the tissue contribution index, the complexity and high cost of cfDNA tracing in existing technologies are solved, achieving efficient, low-cost, and highly accurate tissue tracing.
Patent Information
- Application Number
- PCT/CN2024/110139
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2026-02-12
AI Technical Summary
Existing methylation-based cfDNA tracing methods are complex and costly, while chromatin-based tracing methods have low accuracy and are only applicable to paired-end sequencing, making it difficult to achieve efficient, low-cost, and high-accuracy tissue tracing.
By acquiring tissue-specific expressed gene sets, analyzing the coverage and sequencing depth of cell-free DNA in the transcription start site (TSS) region, and calculating the tissue contribution index, tissue tracing can be achieved.
It provides a tissue tracing method that is easy to operate, low-cost, highly accurate, widely applicable, and has an efficient analytical workflow. It is suitable for both single-end and paired-end sequencing data, reducing the need for plasma volume and experimental complexity.
Smart Images

Figure PCTCN2024110139-FTAPPB-I100001 
Figure PCTCN2024110139-FTAPPB-I100002 
Figure PCTCN2024110139-FTAPPB-I100003
Abstract
Description
Method for tissue tracing of cell-free DNA TECHNICAL FIELD
[0001] The present disclosure relates to the field of biotechnology, in particular to a method for tissue tracing of cell-free DNA, a method for indicating risk of tissue lesion, a method for indicating risk of cancer based on cell-free DNA, a method for indicating risk of gestational disease based on cell-free DNA, a method for indicating risk of recipient tolerance based on cell-free DNA, a method for evaluating or prognosticating treatment effect of disease, and devices, electronic devices, computer-readable storage media, computer program products, and computer programs thereof. BACKGROUND
[0002] Cell-free DNA (cfDNA) is fragmented DNA released into the blood after cell death of various tissues in the human body. The development and treatment of many diseases, such as cancer, autoimmune diseases and sepsis, will affect the cell death rate, thereby affecting the cfDNA fraction of various tissues in the blood. Therefore, the change of abnormal tissue-derived cfDNA components in human plasma can reveal the change of tissue homeostasis due to disease and the collateral tissue damage due to treatment. The cfDNA tissue tracing provides a new risk indication method for non-invasive assessment of various tissues in the whole body and overall health status.
[0003] The cfDNA tissue tracing has great clinical potential in helping disease diagnosis, prognosis and treatment monitoring. The cfDNA tissue tracing related technologies include tissue tracing methods based on tissue-specific methylation sites, tissue tracing methods based on cell-specific methylation and tissue tracing methods based on chromatin open state, etc. However, the methylation-based tracing methods have problems such as too much plasma required, complex operation, high cost and destructive to cfDNA in specific practice; and the tracing based on chromatin open state is only suitable for paired-end sequencing and data with higher sequencing depth, and has lower tracing accuracy.
[0004] Therefore, there is an urgent need for a method for tissue tracing of cell-free DNA which is simple in operation, low in cost, high in accuracy, widely applicable, efficient in analysis process and high in clinical feasibility.
[0005] SUMMARY
[0006] The present disclosure aims to at least partially solve one of the technical problems in the related art.
[0007] To this end, the first aspect of the present disclosure provides a method for tissue provenance of cell-free DNA, comprising: obtaining a tissue-specific expressed gene set of a tissue; obtaining first distribution information of cell-free DNA originated from a sample on a specific region of each gene in the tissue-specific expressed gene set; and obtaining a tissue contribution index of the tissue in the cell-free DNA based on the first distribution information, to realize tissue provenance of the cell-free DNA, wherein the specific region comprises a transcription start site (TSS) region.
[0008] In some embodiments, the TSS region is selected from a region of 5Kb upstream and downstream of the TSS, preferably from a region of 3Kb upstream and downstream of the TSS, more preferably from a region of 1.5Kb upstream and downstream of the TSS, and most preferably from a region of 1Kb upstream and downstream of the TSS.
[0009] In some embodiments, the first distribution information comprises coverage and / or sequencing depth of the cell-free DNA on the TSS region of each gene in the tissue-specific expressed gene set.
[0010] In some embodiments, the obtaining of the first distribution information of the cell-free DNA originated from a sample on a specific region of each gene in the tissue-specific expressed gene set comprises: obtaining second distribution information, wherein the second distribution information comprises coverage and / or sequencing depth of the cell-free DNA on the TSS region of each gene in the whole genome of the sample; and obtaining the first distribution information based on the second distribution information and the tissue-specific expressed gene set.
[0011] In some embodiments, the obtaining of the first distribution information and / or the obtaining of the second distribution information further comprises: screening the cell-free DNA mapped to the TSS region of each gene in the tissue-specific expressed gene set according to a length criterion; and / or screening the cell-free DNA mapped to the specific region in the whole genome corresponding to the sample according to a length criterion, wherein the length criterion for screening the cell-free DNA is 600nt or less, preferably, the length criterion for screening the cell-free DNA is 150-210nt.
[0012] In some embodiments, after the obtaining of the first distribution information and / or the obtaining of the second distribution information, the method comprises: normalizing the first distribution information to obtain normalized first distribution information; and / or normalizing the second distribution information to obtain normalized second distribution information.
[0013] In some embodiments, the obtaining the tissue-specific expressed gene set of the tissue comprises: performing tissue-specific gene differential expression analysis based on the gene expression data of the tissue to obtain the tissue-specific expressed gene set, optionally, the tissue-specific gene differential expression analysis comprises TAU value threshold method and / or clustering analysis method.
[0014] In some embodiments, the tissue-specific gene differential expression analysis comprises: determining a tissue-specific expressed gene candidate set from the gene expression data of the tissue; ranking each gene in the tissue-specific expressed gene set based on the expression level of each gene in the tissue-specific expressed gene candidate set; and selecting the tissue-specific expressed genes ranked N1-N2 in the ranking to obtain the tissue-specific expressed gene set.
[0015] In some embodiments, the N1-N2 is 1-1000, 1-500, 1-250, more preferably, the N1-N2 is 1-500 or 1-250.
[0016] In some embodiments, wherein the statistical indicator of the expression level of the gene is TPM, FPKM, RPKM, CPM or reads counts, preferably, the statistical indicator of the expression level of the gene is TPM, and the ranking is in descending order.
[0017] In some embodiments, the standardization comprises: in the first distribution information or the second distribution information, based on the number of coverage bases of each of the cell-free DNA aligned to each gene in the tissue-specific expressed gene set or aligned to the TSS region of the autosome corresponding to the sample, and the total number of coverage bases of the TSS region, performing the standardization to obtain the standardized first distribution information,
[0018] Wherein the standardization method comprises constructing a mathematical model and / or a machine learning model.
[0019] In some embodiments, wherein the standardized first distribution information is TSS region coverage score, and the mathematical model is:
[0020] Wherein h is the total number of TSS regions in the whole genome, Base_count(i) is the coverage base number of the i-th TSS region of the cell-free DNA, Cov_score(i) is the standardized TSS coverage score of the i-th TSS region, and Base_count(j) is the total coverage base number of the j-th TSS region of the cell-free DNA in the h TSS regions.
[0021] In some embodiments, the formula of the tissue contribution index is:
[0022] where k is the number of genes in the tissue-specific expression set, Cov_score(i, t) is the TSS coverage score of the ith tissue-specific expression gene of the tth tissue, TCI(t) is the normalized tissue contribution index, and x is an arbitrary constant greater than 0.
[0023] In some embodiments, the sample is derived from at least one of plasma, seminal plasma, saliva, urine, amniotic fluid, serum, pleural effusion, cerebrospinal fluid, synovial fluid.
[0024] In some embodiments, the sample is derived from plasma.
[0025] In some embodiments, the one or more tissues are derived from one or more of liver, kidney, heart, pancreas, brain, lung, stomach, spleen, gallbladder, minor salivary glands, colon, small intestine, large intestine, prostate, ovary, uterus, pituitary, whole blood, vagina, breast, thyroid, adrenal gland, placenta.
[0026] A second aspect embodiment of the present disclosure provides a method of suggesting a tissue lesion risk, comprising: obtaining sequencing data of cell-free DNA of a test sample and a reference sample; calculating tissue contribution indices of one or more tissues of the test sample and the reference sample according to the method of tracing tissue origin of cell-free DNA of any one of the first aspect embodiments of the present disclosure; and suggesting that the tissue of a sampled individual of the test sample is suspected of or has a lesion based on a significant difference between the tissue contribution index of the test sample and the tissue contribution index of the reference sample for the same tissue.
[0027] In some embodiments, the one or more tissues are derived from one or more of liver, kidney, heart, pancreas, brain, lung, stomach, spleen, gallbladder, minor salivary glands, colon, small intestine, large intestine, prostate, ovary, uterus, pituitary, whole blood, vagina, breast, thyroid, adrenal gland, placenta.
[0028] In some embodiments, the method comprises suggesting that the tissue lesion occurs in at least one of a pregnancy disorder, a neurodegenerative disease, a cancer, and an autoimmune disease based on a significant difference between the tissue contribution index of the test sample and the tissue contribution index of the reference sample.
[0029] A third aspect embodiment of the present disclosure provides a method for indicating cancer risk based on cell-free DNA, comprising: obtaining cell-free DNA sequencing data of a test sample from a subject; calculating tissue contribution indices of one or more tissues of the test sample and a reference sample according to the method for tracing tissue of cell-free DNA of any one of the first aspect embodiments of the present disclosure; and indicating that the subject is suspected of or has cancer lesions in the same tissue based on a significant difference between the tissue contribution index of the test sample and the tissue contribution index of the reference sample.
[0030] A fourth aspect embodiment of the present disclosure provides a method for indicating pregnancy disease risk based on cell-free DNA, comprising: obtaining cell-free DNA sequencing data of a test sample from a subject; calculating tissue contribution indices of placental tissue of the test sample and a reference sample according to the method for tracing tissue of cell-free DNA of any one of the first aspect embodiments of the present disclosure; and indicating that the subject is suspected of or has a pregnancy disease based on a significant difference between the tissue contribution index of the test sample and the tissue contribution index of the reference sample.
[0031] A fifth aspect embodiment of the present disclosure provides a method for indicating recipient tolerance risk based on cell-free DNA, comprising: obtaining cell-free DNA sequencing data of a test sample from a subject; calculating tissue contribution indices of donor tissue of the test sample and a reference sample according to the method for tracing tissue of cell-free DNA of any one of the first aspect embodiments of the present disclosure; and indicating that the subject is suspected of or has recipient intolerance based on a significant difference between the tissue contribution index of the test sample and the tissue contribution index of the reference sample.
[0032] A sixth aspect embodiment of the present disclosure provides a method for evaluating or prognosing disease treatment effect, comprising: obtaining cell-free DNA sequencing data of test samples from a subject at different time periods; calculating tissue contribution indices of donor tissue of the test samples according to the method for tracing tissue of cell-free DNA of any one of the first aspect embodiments of the present disclosure; and indicating the disease treatment effect or prognosis of the subject based on the tissue contribution indices of the test samples at different time periods.
[0033] A seventh aspect embodiment of the present disclosure provides a tissue tracing device, comprising: a first obtaining module configured to obtain a tissue-specific expression gene set of a tissue; a second obtaining module configured to obtain first distribution information of cell-free DNA from a sample on a specific region of each gene in the tissue-specific expression gene set; and a third obtaining module configured to obtain a tissue contribution index of the tissue in the cell-free DNA to realize tissue tracing of the cell-free DNA.
[0034] An eighth aspect of the present disclosure provides an organization tracing device, comprising: a processor; a memory for storing executable instructions; wherein the processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the method for tracing organization of free DNA according to any one of the first aspect of the present disclosure, or the method for prompting risk of organization lesion according to any one of the second aspect of the present disclosure, or the method for indicating risk of cancer based on free DNA according to the third aspect of the present disclosure, or the method for indicating risk of disease during pregnancy based on free DNA according to the fourth aspect of the present disclosure, or the method for prompting risk of recipient tolerance based on free DNA according to the fifth aspect of the present disclosure, or the method for evaluating treatment effect or prognosis of disease according to the sixth aspect of the present disclosure.
[0035] An ninth aspect of the present disclosure provides a computer readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the method for tracing organization of free DNA according to any one of the first aspect of the present disclosure, or the method for prompting risk of organization lesion according to any one of the second aspect of the present disclosure, or the method for indicating risk of cancer based on free DNA according to the third aspect of the present disclosure, or the method for indicating risk of disease during pregnancy based on free DNA according to the fourth aspect of the present disclosure, or the method for prompting risk of recipient tolerance based on free DNA according to the fifth aspect of the present disclosure, or the method for evaluating treatment effect or prognosis of disease according to the sixth aspect of the present disclosure.
[0036] An tenth aspect of the present disclosure provides a computer program product, wherein the computer program product comprises a computer program, and when the computer program is executed by a processor, the processor implements the method for tracing organization of free DNA according to any one of the first aspect of the present disclosure, or the method for prompting risk of organization lesion according to any one of the second aspect of the present disclosure, or the method for indicating risk of cancer based on free DNA according to the third aspect of the present disclosure, or the method for indicating risk of disease during pregnancy based on free DNA according to the fourth aspect of the present disclosure, or the method for prompting risk of recipient tolerance based on free DNA according to the fifth aspect of the present disclosure, or the method for evaluating treatment effect or prognosis of disease according to the sixth aspect of the present disclosure.
[0037] The eleventh aspect of the present disclosure provides a computer program, wherein the computer program comprises computer program codes, when the computer program codes are run on a computer, to make the computer execute the method for tissue tracing of free DNA as any one of the first aspect of the present disclosure, or the method for prompting tissue lesion risk as any one of the second aspect of the present disclosure, or the method for cancer risk indication based on free DNA as any one of the third aspect of the present disclosure, or the method for pregnancy disease risk indication based on free DNA as any one of the fourth aspect of the present disclosure, or the method for receptor tolerance risk prompting based on free DNA as any one of the fifth aspect of the present disclosure, or the disease treatment effect evaluation or prognosis method as the sixth aspect of the present disclosure.
[0038] The advantages and technical effects of the present disclosure are as follows:
[0039] (1) Simple operation and low cost: In the method provided by the embodiments of the present disclosure, when sequencing cfDNA, high initial plasma volume and complex experimental processing are not required, and only conventional cfDNA whole genome sequencing is required, and the sequencing depth requirement is low, only 1-10X, so the cost is relatively low.
[0040] (2) High accuracy: The method provided by the embodiments of the present disclosure is based on the analysis of specific expression genes of tissue organs, and has high accuracy in tissue tracing.
[0041] (3) Wide applicability: The method provided by the embodiments of the present disclosure is not only suitable for paired-end sequencing data, but also suitable for single-end sequencing data, and is not limited by sequencing read length. Therefore, it can adapt to data of various sequencing platforms and sequencing modes, and has wide application prospects.
[0042] (4) Efficient analysis process: The method provided by the embodiments of the present disclosure only needs to analyze the coverage of the TSS region, the analysis region is small, and the computing resource occupation is small, thereby ensuring the efficiency and rapidity of the analysis, and greatly reducing the required time.
[0043] (5) High clinical feasibility: Compared with the methylation tissue tracing method, the method provided by the embodiments of the present disclosure requires lower initial plasma volume, simple experimental operation, and no additional methylation conversion processing, so it has higher clinical feasibility.
[0044] In summary, the method provided by the embodiments of the present disclosure has the advantages of simple operation, low cost, high accuracy, wide applicability, efficient analysis process and high clinical feasibility, and provides a new powerful tool for tissue tracing and other feature analysis. BRIEF DESCRIPTION OF DRAWINGS
[0045] FIG. 1 is a schematic diagram of a methylation sequencing-based tissue tracing method in the related art.
[0046] FIG. 2 is a schematic diagram of a methylation sequencing-based cell tracing method in the related art.
[0047] FIG. 3 is a schematic diagram of a tissue-specific chromatin openness state-based tissue tracing method in the related art.
[0048] FIG. 4 is a schematic diagram of the results of a tissue-specific expression gene atlas provided in an embodiment of the present disclosure.
[0049] FIG. 5 is a schematic diagram of the results of tissue tracing of maternal plasma cell-free DNA in an embodiment of the present disclosure.
[0050] FIG. 6 is a schematic diagram of the results of tissue tracing of plasma cell-free DNA of a bone marrow transplant recipient in an embodiment of the present disclosure.
[0051] FIG. 7 is a schematic diagram of the results of tissue tracing of plasma cell-free DNA of a liver transplant recipient in an embodiment of the present disclosure.
[0052] FIG. 8 is a schematic diagram of the results of tissue tracing of plasma cell-free DNA of a liver cancer patient in an embodiment of the present disclosure.
[0053] FIG. 9 is a schematic diagram of the results of tissue tracing of maternal plasma cell-free DNA at an ultra-low sequencing depth in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0054] Embodiments of the present disclosure are described in detail below with reference to the accompanying drawings. The embodiments described below by reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, and cannot be understood as limiting the present disclosure.
[0055] Cell-free DNA (cfDNA) is fragmented DNA released into the blood after cell death of various tissues in the human body. The development of many diseases such as cancer, autoimmune diseases and sepsis will all cause abnormal cell death, thereby affecting the cfDNA portion from each tissue in the blood. Therefore, through tissue tracing of cfDNA, a new risk warning method can be provided for non-invasive evaluation of each tissue in the body and overall health status.
[0056] The existing technologies for tissue tracing of cfDNA mainly include a tissue tracing method based on methylation sequencing and a tissue tracing method based on tissue-specific chromatin openness. The tissue tracing method based on methylation sequencing in the related art is shown in FIG. 1. The method performs methylation sequencing on sample plasma cfDNA, analyzes the methylation rate of tissue-specific methylation sites in cfDNA, and analyzes the source proportion of each tissue organ by a quadratic programming method to complete tissue tracing. On this basis, as shown in the cell tracing method based on methylation sequencing in FIG. 2, different cell types in cfDNA can be analyzed, and tissue tracing can be completed based on the cell types. However, the method based on methylation sequencing has high requirements for samples and needs more plasma. On the other hand, the experimental operation of methylation sequencing itself is complex and costly, and the bisulfite treatment in the experimental process will also cause fragmentation of cfDNA, making it difficult to use the data for other dimensional analysis.
[0057] The tissue tracing method based on tissue-specific chromatin openness in the related art is shown in FIG. 3. The chromatin state of the genome is tissue-specific, and the nucleosome is regularly arranged in the chromatin conserved region. DNA is not easily degraded by nucleases in plasma due to the protection of nucleosomes, while the DNA in the chromatin open region is fragmented by degradation. By obtaining the distribution and frequency of cfDNA end positions, the degradation degree of tissue-specific open chromatin regions can be quantified. By analyzing the degradation degree of these tissue-specific open chromatin regions, the contribution degree of tissue organs in plasma free DNA can be calculated to complete tissue tracing. However, the tissue tracing method based on tissue-specific chromatin openness needs high sequencing depth and has insufficient accuracy, and is only suitable for data of paired-end sequencing.
[0058] Therefore, based on the correlation between the openness of the transcription start site (TSS) on the genome and gene expression, the embodiment of the present disclosure provides a method for tracing the tissue of free DNA, which comprises: obtaining a tissue-specific expression gene set of a tissue; obtaining first distribution information of free DNA derived from a sample on a specific region of each gene in the tissue-specific expression gene set; and obtaining a tissue contribution index of the tissue in the free DNA based on the first distribution information, to realize tissue tracing of the free DNA, wherein the specific region comprises a transcription start site (TSS) region.
[0059] It should be noted that the method for tissue tracing of free DNA in the embodiments of the present disclosure is achieved by correlating the openness of chromatin at the transcription start site (TSS) on the genome with gene expression, wherein the TSS region of highly expressed genes usually presents an open chromatin state, thus the cfDNA originating from this TSS region is difficult to capture in sequencing due to easier degradation, thereby showing low coverage; on the contrary, the TSS region of lowly expressed genes presents a closed state, which is difficult to be degraded due to the protection of nucleosomes, thus is easy to be captured in sequencing, thereby showing high coverage, thereby achieving tissue tracing with high accuracy and small data volume based on tissue-specific genes.
[0060] In some embodiments, the TSS region is selected from a region of 5Kb upstream and downstream of the TSS, preferably from a region of 3Kb upstream and downstream of the TSS, more preferably from a region of 1.5Kb upstream and downstream of the TSS, most preferably from a region of 1Kb upstream and downstream of the TSS. It can be understood that the transcription start site (TSS) is the starting point of transcription activity, located at the 3' end of the transcription unit. The TSS region is the region upstream and downstream of the TSS, in which there are various promoters, transcription factors and regulatory sequences that work together to regulate the transcription expression of genes. Based on the method of the embodiments of the present disclosure, those skilled in the art can expand or reduce the range of the TSS region in different samples to obtain the optimal effect. The present disclosure does not intend to limit this.
[0061] In some embodiments, the first distribution information includes the coverage and / or sequencing depth of the free DNA at the transcription start site (TSS) region of each gene in the set of tissue-specific expression genes.
[0062] In some embodiments, the obtaining the first distribution information of the free DNA originating from the sample at the specific region of each gene in the set of tissue-specific expression genes comprises: obtaining second distribution information, wherein the second distribution information includes the coverage and / or sequencing depth of the free DNA at the TSS region of each gene in the whole genome of the sample; and obtaining the first distribution information based on the second distribution information and the set of tissue-specific expression genes.
[0063] It should be noted that by comparing the tissue contribution index of the same tissue at different times in the same sample, or the tissue contribution index of the same tissue in different samples, the change in the contribution of the tissue can be obtained, thereby achieving disease risk indication and / or prognosis indication.
[0064] It can be understood that as long as an index capable of being used to analyze the distribution (such as coverage / depth, etc.) of cfDNA in the TSS region of each tissue-specific expression gene to reflect the open or closed state of the chromatin of the TSS region of each tissue-specific expression gene can be referred to as first distribution information. Therefore, the first distribution information can also be the sequencing depth of cfDNA in the TSS region of a tissue-specific expression gene, the total sequencing depth, the depth standardized in different ways, the diversity of fragment length, etc., to quantify the open or closed state of the chromatin of the TSS region.
[0065] In some embodiments, obtaining the first distribution information and / or obtaining the second distribution information further comprises: screening the free DNA aligned to the TSS region of each gene in the set of tissue-specific expression genes according to a length criterion; and / or screening the free DNA aligned to the specific region in the whole genome corresponding to the sample according to a length criterion, wherein the length criterion for screening the free DNA is 600 nt or less, preferably, the length criterion for screening the free DNA is 150-210 nt.
[0066] In the embodiments of the present application, by setting the length threshold (criterion) of cfDNA aligned to the TSS region and screening the cfDNA matched to each TSS region based on the criterion, cfDNA originating from the open chromatin region of each tissue can be effectively obtained, which cooperates with each highly expressed gene in the set of tissue-specific genes of each tissue, that is, cfDNA originating from the TSS region of the specific highly expressed gene of each tissue can be traced back to the corresponding tissue, thereby simply and quickly achieving high-accuracy tissue tracing. It can be understood that based on the method of the embodiments of the present disclosure, different lengths of TSS length fragments can be selected based on different first distribution information (such as the expression level of each gene in the set of tissue-specific genes, the coverage / depth of cfDNA) to obtain the optimal effect. The present disclosure does not intend to make such a limitation.
[0067] In some embodiments, after obtaining the first distribution information and / or obtaining the second distribution information, the method comprises: standardizing the first distribution information to obtain standardized first distribution information; and / or standardizing the second distribution information to obtain standardized second distribution information. It can be understood that by standardizing the first distribution information, the reliability and performance of the data can be improved, and the standardized data can be applied to subsequent algorithm analysis, thereby achieving more accurate tissue tracing.
[0068] In some embodiments, the normalized first distribution information is in the form of a first distribution score. It can be understood that the normalized first distribution information corresponds to the quantification standard of the first distribution information and the manner of normalization, for example, when the first distribution information is the number of covered bases of cfDNA at the TSS region of each gene in the tissue-specific expression gene set, the manner of normalization can be normalization based on the total number of covered bases of cfDNA at the TSS region of each gene in the tissue-specific expression gene set, and the obtained coverage score of each TSS region is the first distribution score.
[0069] In addition, in the embodiments of the present application, as long as it can be used to analyze the distribution (such as coverage / depth, etc.) of cfDNA at the TSS region of each gene in the whole genome (autosome) corresponding to the sample of the tissue to reflect the distribution state of cfDNA at the TSS region of the whole genome (autosome), it can be called second distribution information. Therefore, the second distribution information can also be the sequencing depth of cfDNA at the TSS region of each gene in the whole genome (autosome), the total sequencing depth, the depth standardized in different ways, the diversity of fragment length, etc., to quantify the chromatin open or closed state of the TSS region in the genomic dimension.
[0070] It can be understood that by normalizing the first distribution information according to the second distribution information, the reliability and performance of the data can be improved, and the normalized data can be applied to subsequent algorithm analysis, thereby realizing more accurate tissue tracing.
[0071] In some embodiments, the tissue-specific expression gene set of the tissue comprises: performing tissue-specific gene differential expression analysis based on gene expression data of the tissue to obtain the tissue-specific expression gene set. In some embodiments, the tissue-specific gene differential expression analysis comprises TAU value threshold method and / or clustering analysis method and / or other self-developed methods, such as specific gene analysis / screening command line, script or software developed based on different algorithms. It can be understood that those skilled in the art can use various methods to perform tissue-specific gene differential expression analysis to obtain the tissue-specific expression gene set of the tissue, and can also obtain the tissue-specific expression gene set of the tissue from known public data. The present disclosure does not intend to limit the way of obtaining the tissue-specific expression gene set.
[0072] In some embodiments, the tissue-specific gene differential expression analysis comprises: determining a candidate set of tissue-specific expression genes according to gene expression data of the tissue; ranking each gene in the set of tissue-specific expression genes based on expression level of each gene in the candidate set of tissue-specific expression genes; and selecting the tissue-specific expression genes ranked N1-N2 in the ranking to obtain the set of tissue-specific expression genes.
[0073] In some embodiments, the N1-N2 is 1-1000, 1-500, 1-250, more preferably, the N1-N2 is 1-500 or 1-250.
[0074] In some embodiments, the statistical indicator of expression level of the gene is TPM, FPKM, RPKM, CPM or reads counts, preferably, the statistical indicator of expression level of the gene is TPM, and the ranking is in descending order. That is, the genes with higher expression level in the candidate set of tissue-specific genes are included in the set of tissue-specific expression genes.
[0075] It can be understood that in some embodiments, a specific gene in the set of tissue-specific genes has a lower expression amount, but has strong tissue specificity and plays an important role. Based on the method of the embodiments of the present disclosure, those skilled in the art can adjust the range of expression amount of the selected genes, the number of selected tissue-specific genes, the order and strategy of screening to obtain the optimal effect. The present disclosure does not intend to limit this.
[0076] In some embodiments, the standardization comprises: in the first distribution information or the second distribution information, based on the number of covered bases of each of the cell-free DNAs aligned to each gene in the set of tissue-specific expression genes or aligned to the TSS region of the autosome corresponding to the sample, and the total number of covered bases of the TSS region, the standardization is performed to obtain the standardized first distribution information,
[0077] The standardization manner comprises constructing a mathematical model and / or a machine learning model.
[0078] In some embodiments, the standardized first distribution information is TSS region coverage fraction, and the mathematical model is:
[0079] wherein h is the total number of TSS regions in the whole genome, Base_count(i) is the number of free DNA coverage bases of the ith TSS region, Cov_score(i) is the normalized TSS coverage score of the ith TSS region, and Base_count(j) is the total number of free DNA coverage bases of the jth TSS region among the h TSS regions.
[0080] In some embodiments, the formula for calculating the tissue contribution index is:
[0081] wherein k is the number of genes in the tissue-specific expression of the tissue, Cov_score(i, t) is the TSS coverage score of the ith tissue-specific expression gene of tissue t, TCI(t) is the normalized tissue contribution index, and x is an arbitrary constant greater than 0. Since Cov_score is mostly distributed between 20 and 30, in order to make the distribution of TCI comparable to Cov_score, x can be set to 50 in the embodiments of the present disclosure.
[0082] In some embodiments, the sample is derived from at least one of plasma, seminal plasma, saliva, urine, amniotic fluid, serum, pleural effusion, cerebrospinal fluid, synovial fluid.
[0083] In some embodiments, the sample is derived from plasma.
[0084] In some embodiments, the tissue is one or more, and the tissue is derived from one or more of liver, kidney, heart, pancreas, brain, lung, stomach, spleen, gallbladder, small salivary glands, colon, small intestine, large intestine, prostate, ovary, uterus, pituitary, whole blood, vagina, breast, thyroid, adrenal gland, placenta.
[0085] A second aspect of the present disclosure provides a method for suggesting a tissue lesion risk, comprising: obtaining sequencing data of free DNA of a test sample and a reference sample; calculating a tissue contribution index of one or more tissues of the test sample and the reference sample according to the method for tracing tissue of free DNA of any one of the first aspect of the present disclosure; and based on the tissue contribution index of the test sample and the tissue contribution index of the reference sample having a significant difference for the same tissue, suggesting that the tissue of the sampling individual of the test sample is suspected or has a lesion.
[0086] In some embodiments, the one or more tissues are derived from one or more of liver, kidney, heart, pancreas, brain, lung, stomach, spleen, gallbladder, small salivary glands, colon, small intestine, large intestine, prostate, ovary, uterus, pituitary, whole blood, vagina, breast, thyroid, adrenal gland, placenta.
[0087] In some embodiments, wherein the tissue contribution index of the test sample and the tissue contribution index of the reference sample have a significant difference, the method comprises: prompting that the tissue lesion is in at least one of the pregnancy disease, the neurodegenerative disease, the cancer and the autoimmune disease.
[0088] A third aspect embodiment of the present disclosure proposes a cancer risk indication method based on free DNA, comprising: obtaining free DNA sequencing data of a test sample derived from a subject; calculating the tissue contribution index of one or more tissues of the test sample and a reference sample according to the method for tracing the tissue of free DNA of any one of the first aspect embodiments of the present disclosure; and for the same tissue, based on the tissue contribution index of the test sample and the tissue contribution index of the reference sample having a significant difference, prompting that the tissue of the subject is suspected or has a cancer lesion.
[0089] A fourth aspect embodiment of the present disclosure proposes a pregnancy disease risk indication method based on free DNA, comprising: obtaining free DNA sequencing data of a test sample derived from a subject; calculating the tissue contribution index of the placental tissue of the test sample and a reference sample according to the method for tracing the tissue of free DNA of any one of the first aspect embodiments of the present disclosure; and based on the tissue contribution index of the test sample and the tissue contribution index of the reference sample having a significant difference, prompting that the subject is suspected or has a pregnancy disease.
[0090] A fifth aspect embodiment of the present disclosure proposes a recipient tolerance risk prompting method based on free DNA, comprising: obtaining free DNA sequencing data of a test sample derived from a subject; calculating the tissue contribution index of the donor tissue of the test sample and a reference sample according to the method for tracing the tissue of free DNA of any one of the first aspect embodiments of the present disclosure; and based on the tissue contribution index of the test sample and the tissue contribution index of the reference sample having a significant difference, prompting that the subject is suspected or has a recipient intolerance.
[0091] A sixth aspect embodiment of the present disclosure proposes a disease treatment effect evaluation or prognosis method, comprising: obtaining free DNA sequencing data of a test sample at different time periods derived from a subject; calculating the tissue contribution index of the donor tissue of the test sample according to the method for tracing the tissue of free DNA of any one of the first aspect embodiments of the present disclosure; and based on the tissue contribution index of the test sample at different time periods, prompting the disease treatment effect or prognosis of the subject. It can be understood that based on different conditions, the change of the tissue contribution index of the test sample at different time periods can prompt the disease treatment effect or prognosis. The change can be significantly improved, decreased or basically stable.
[0092] In a seventh aspect, an embodiment of the present disclosure provides an apparatus for tissue provenance, comprising: a first obtaining module configured to obtain a set of tissue-specific expression genes of a tissue; a second obtaining module configured to obtain first distribution information of cell-free DNA originating from a sample on a specific region of each gene in the set of tissue-specific expression genes; and a third obtaining module configured to obtain a tissue contribution index of the tissue in the cell-free DNA to realize tissue provenance of the cell-free DNA.
[0093] In an eighth aspect, an embodiment of the present disclosure provides an apparatus for tissue provenance, comprising: a processor; a memory configured to store executable instructions; wherein the processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the method for tissue provenance of cell-free DNA according to any one of the embodiments of the first aspect, or the method for indicating a risk of a tissue lesion according to any one of the embodiments of the second aspect, or the method for indicating a risk of cancer based on cell-free DNA according to the embodiments of the third aspect, or the method for indicating a risk of a disease during pregnancy based on cell-free DNA according to the embodiments of the fourth aspect, or the method for indicating a risk of recipient tolerance based on cell-free DNA according to the embodiments of the fifth aspect, or the method for evaluating or prognosticating a treatment effect of a disease according to the embodiments of the sixth aspect.
[0094] In a ninth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the method for tissue provenance of cell-free DNA according to any one of the embodiments of the first aspect, or the method for indicating a risk of a tissue lesion according to any one of the embodiments of the second aspect, or the method for indicating a risk of cancer based on cell-free DNA according to the embodiments of the third aspect, or the method for indicating a risk of a disease during pregnancy based on cell-free DNA according to the embodiments of the fourth aspect, or the method for indicating a risk of recipient tolerance based on cell-free DNA according to the embodiments of the fifth aspect, or the method for evaluating or prognosticating a treatment effect of a disease according to the embodiments of the sixth aspect.
[0095] In a tenth aspect, the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, implements the method for tissue tracing of cell-free DNA according to any one of the first aspect, or the method for indicating risk of tissue lesion according to any one of the second aspect, or the method for indicating risk of cancer based on cell-free DNA according to any one of the third aspect, or the method for indicating risk of disease during pregnancy based on cell-free DNA according to any one of the fourth aspect, or the method for indicating risk of recipient tolerance based on cell-free DNA according to any one of the fifth aspect, or the method for evaluating treatment effect or prognosis of disease according to the sixth aspect.
[0096] In an eleventh aspect, the present disclosure provides a computer program comprising computer program code which, when run on a computer, causes the computer to perform the method for tissue tracing of cell-free DNA according to any one of the first aspect, or the method for indicating risk of tissue lesion according to any one of the second aspect, or the method for indicating risk of cancer based on cell-free DNA according to any one of the third aspect, or the method for indicating risk of disease during pregnancy based on cell-free DNA according to any one of the fourth aspect, or the method for indicating risk of recipient tolerance based on cell-free DNA according to any one of the fifth aspect, or the method for evaluating treatment effect or prognosis of disease according to the sixth aspect.
[0097] It should be noted that the above explanations of the method embodiments also apply to the above-mentioned apparatuses, electronic devices, vehicles, computer-readable storage media, computer program products and computer programs, and will not be repeated here.
[0098] The method for tissue tracing of cell-free DNA and its related applications according to the present disclosure will be explained in detail below with reference to the embodiments.
[0099] The experimental methods in the following embodiments are all conventional methods, and are performed according to the techniques or conditions described in the literature in the field or according to the product instructions, unless otherwise specified. The materials, reagents and the like used in the following embodiments can be obtained from commercial channels, unless otherwise specified.
[0100] The quantitative experiments in the following embodiments are all set up with three repeated experiments, and the results are averaged, unless otherwise specified.
[0101] Embodiment 1
[0102] The embodiment of the present disclosure uses 12 known tissue transcriptome data sets to screen tissue-specific expression gene data sets; and uses them to perform tissue tracing on pregnant women's plasma cfDNA samples and verify the reliability of the tracing.
[0103] 1.1 Obtain tissue-specific expression gene data set
[0104] (1) Download gene expression data of healthy human whole blood, lung, breast, aorta, heart, colon, stomach, pancreas, liver, ovary and kidney from GTEx database.
[0105] (2) Download the gene expression data of placental tissue disclosed in the literature, and the literature link is https: / / www.obgyn.cam.ac.uk / placentome / .
[0106] (3) Merge the gene expression data of the above 12 tissues and perform quality control and other preprocessing.
[0107] (4) Obtain the gene expression amount statistics TPM of the gene in each sample, and construct a gene expression matrix.
[0108] (5) Based on the above gene expression matrix, the gene expression level difference (V1) and the gene unique high expression difference (V2) of each gene in each tissue are calculated, and the calculation method is as follows:
[0109] V1=(TPM maximum value-TPM three quarters value) / (TPM three quarters value-TPM one quarter value);
[0110] V2=(TPM maximum value-TPM second largest value) / TPM maximum value. If V1>10 and V2>0.3, the gene can be considered as a specific high expression gene of the tissue corresponding to the TPM maximum value, and finally a candidate tissue-specific gene set is obtained.
[0111] (6) Sort the gene expression amount in the candidate tissue-specific gene set, and select the highly expressed genes as the final tissue-specific gene data set.
[0112] (7) Visualize the candidate tissue-specific gene data set.
[0113] The tissue-specific expression gene atlas of the embodiment is shown in FIG. 4, and it can be seen from the figure that the expression patterns of genes in different tissues have significant specific high expression. Finally, the top 500 specific high expression genes in whole blood tissue and the top 250 specific high expression genes in other tissues are selected as the final tissue-specific gene data set. The tissue-specific gene data set obtained in the embodiment can be used for subsequent tissue tracing of free DNA.
[0114] 1.2 Tissue tracing of pregnant women's plasma cfDNA
[0115] (1) Plasma cfDNA of 74 pregnant women with male fetus were collected and high-throughput sequenced (BGI sequencing platform, paired-end sequencing, read length 100bp, sequencing depth range: 8X-10X).
[0116] (2) The cfDNA sequencing data was pre-processed, and then the cfDNA fragments were aligned to the total of 38865 transcript start site (TSS) regions of autosomes, wherein the TSS region was 1000bp upstream and downstream of the TSS, and the aligned cfDNA fragments were limited to 150 to 198bp.
[0117] (3) Based on the number of covered bases of each cfDNA fragment aligned to the above-mentioned 38865 TSS regions and the total number of covered bases of the TSS regions, the normalized TSS region coverage score was calculated, wherein the calculation formula of the normalized TSS region coverage score was:
[0118] Wherein Base_count(i) was the number of free DNA covered bases of the ith TSS region, Cov_score(i) was the normalized TSS region coverage score of the ith TSS region, and Base_count(j) was the total number of free DNA covered bases of the jth TSS region in the hth TSS region.
[0119] (4) The TSS region coverage score calculated in the above-mentioned step (3) and the tissue-specific gene dataset obtained in Example 1.1 were input into the formula, and the tissue contribution index of 12 tissues to the pregnant women's plasma cfDNA was calculated, and the calculation formula was:
[0120] Wherein k was the number of tissue-specific expressed genes of the tissue, Cov_score(i,t) was the TSS region coverage score of the ith tissue-specific expressed gene of the tissue t, and TCI(t) was the normalized tissue contribution index.
[0121] In this example, the tissue contribution index (TCI) of 12 tissues in the plasma cfDNA of 74 pregnant women with male fetus was obtained.
[0122] 1.3 Verification of tissue tracing
[0123] (1) The reference fetal concentration (fetal fraction) of the above-mentioned 74 pregnant women with male fetus was calculated based on the Y chromosome method.
[0124] (2) The correlation between the fetal concentration of the sample and the placental tissue TCI was calculated.
[0125] (3) Visualize the relationship between fetal concentration and coverage of TSS region of placenta-specific highly expressed genes.
[0126] Figure 5 is a schematic diagram of the results of the organization tracing of free DNA in the plasma of pregnant women in the embodiments of the present disclosure.
[0127] The correlation between fetal concentration and TCI of placental tissue is shown in part A of Figure 5. The correlation coefficient between the contribution degree (TCI) of placental tissue obtained by the method for organization tracing of free DNA provided in the embodiments of the present disclosure and the estimated fetal concentration of Y chromosome reached 0.93 (P < 0.0001). Thus, it is shown that the method for organization tracing of free DNA provided in the present disclosure can also be used to assist non-invasive prenatal genetic testing (NIPT), and the high correlation between the contribution degree (TCI) of placental tissue obtained by the method and the estimated fetal concentration of Y chromosome also indicates the accuracy and reliability of the method for organization tracing of free DNA provided in the present disclosure.
[0128] In addition, the relationship between fetal concentration and coverage of TSS region of placenta-specific highly expressed genes is shown in part B of Figure 5. With the increase of fetal concentration, the coverage of TSS region of placenta-specific highly expressed genes gradually decreases. This indicates that with the increase of fetal concentration, the expression of placental tissue-specific genes is higher, which is consistent with the prior knowledge, thus again indicating the accuracy and reliability of the method for organization tracing of free DNA provided in the present disclosure.
[0129] Example 2 Plasma cfDNA organization tracing of bone marrow transplant and liver transplant recipients
[0130] In the embodiments of the present disclosure, the tissue transcriptome data sets of 12 known tissues are used to screen the data set of tissue-specific expressed genes, and the plasma cfDNA of bone marrow transplant and liver transplant recipients is subjected to organization tracing and verification of the reliability of organization tracing.
[0131] (1) Download the plasma cfDNA sequencing data of 48 bone marrow transplant recipients (Illumina sequencing platform, paired-end sequencing, read length 100 bp, sequencing depth: 0.75-2.5X, https: / / www.ncbi.nlm.nih.gov / bioproject / PRJNA352904) and 13 liver transplant recipients (Illumina sequencing platform, paired-end sequencing, read length 76 bp, sequencing depth: 0.9-1.5X, http: / / finaledb.research.cchmc.org).
[0132] (2) According to the steps (2)-(4) in embodiment 1.2, the tissue contribution index of 12 tissues to the plasma cfDNA of bone marrow transplantation and liver transplantation recipients was calculated.
[0133] (3) According to the reference, the reference blood cell concentration and the reference liver concentration of donor origin in the plasma cfDNA of the recipient were obtained, and the correlation with the tissue contribution index of bone marrow and liver provided in (2) was calculated (the reference is Sun K et al. Orientation-aware plasma cell-free DNA fragmentation analysis in open chromatin regions informs tissue of origin [J]. Genome Research, 2019, 29(3): 418-427. DOI: 10.1101 / gr.242719.118. and Sharon E et al. Quantification of transplant-derived circulating cell-free DNA in absence of a donor genotype [J]. Plos Computational Biology, 2017, 13(8). DOI: 10.1371 / journal.pcbi.1005629.).
[0134] FIGS. 6 and 7 are schematic diagrams of the results of the tissue tracing of the plasma free DNA of bone marrow transplantation and liver transplantation recipients in the embodiments of the present disclosure.
[0135] The correlation between the recipient blood cell TCI and the concentration of blood cells of donor origin is shown in part A of FIG. 6, and the correlation coefficient is 0.81, (P<0.0001), and the correlation between the recipient liver TCI and the concentration of liver of donor origin is shown in part A of FIG. 7, and the correlation coefficient is 0.90 (P<0.0001). Thus, it is shown that the method for tissue tracing of free DNA provided by the present disclosure is suitable for tissue tracing of recipients after organ transplantation, and the accuracy and reliability of the method for tissue tracing of free DNA provided by the present disclosure are verified.
[0136] In addition, as shown in part B of FIG. 6, with the increase of the concentration of blood cells derived from the donor, the coverage of the TSS region of the blood cell (developed from bone marrow) specific gene gradually decreases; similarly, as shown in part B of FIG. 7, with the increase of the concentration of liver derived from the donor, the coverage of the TSS region of the liver tissue specific gene gradually decreases. Again, the accuracy and reliability of the method for tracing the tissue of free DNA provided by the present disclosure are illustrated, and it is also shown that the method has a good application prospect in organ transplant patients.
[0137] Example 3 Tissue tracing of plasma free DNA of liver cancer patients
[0138] In the present embodiment, the tissue transcriptome data sets of 12 known tissues are used to screen the tissue-specific expression gene data set, and the plasma cfDNA of liver cancer patients is traced to the tissue and the reliability of the tracing is verified.
[0139] (1) The plasma cfDNA sequencing data (Illumina sequencing platform, double-end sequencing, read length 75 bp, sequencing depth: 1.5-2.5X) of 32 healthy references and 74 liver cancer patients were downloaded.
[0140] (2) According to the steps (2)-(4) in Example 1.2, the tissue contribution index of 12 tissues to the plasma cfDNA of healthy references and liver cancer patients was calculated.
[0141] (3) The tumor DNA concentration of liver cancer patients was calculated according to the copy number variation of plasma cfDNA of liver cancer patients, and the correlation between the tumor DNA concentration and the liver TCI of liver cancer patients provided in (2) was calculated.
[0142] (4) The liver TCI of healthy controls, low-load liver cancer patients and high-load liver cancer patients was subjected to variance analysis (high load: tumor concentration > 10%, low load: tumor concentration <= 10%).
[0143] (5) ROC analysis was performed according to the liver TCI provided in (2) of the present embodiment.
[0144] FIG. 8 is a schematic diagram of the results of the tissue tracing of plasma free DNA of liver cancer patients in the present embodiment.
[0145] The correlation between the tumor DNA concentration of liver cancer patients and the liver TCI is shown in part A of FIG. 8, and the correlation coefficient is 0.35 (P = 0.0003). Thus, it is shown that the method for tracing the tissue of free DNA provided by the present disclosure is suitable for cancer risk indication, especially liver cancer risk indication, indicating the accuracy and reliability of the method for tracing the tissue of free DNA provided by the present disclosure.
[0146] In addition, the liver TCI of the healthy control, low-tumor load liver cancer patient, and high-tumor load liver cancer patient is shown in part B of FIG. 8. The liver TCI of the liver cancer patient is significantly higher than that of the healthy control, and the liver TCI of the high-tumor load liver cancer patient is significantly higher than that of the low-tumor load liver cancer patient. Moreover, when the liver TCI provided by the present disclosure is used to diagnose the healthy control and the high-tumor load liver cancer patient, the area under the ROC curve (AUC) of the judgment method reaches 0.996, which can better correctly identify the high-tumor load liver cancer patient sample (part C of FIG. 8), again demonstrating the accuracy and reliability of the method for tracing the free DNA to the tissue provided by the present disclosure and the risk prompting method based thereon.
[0147] Example 4 Tissue tracing of pregnant woman plasma cfDNA at ultra-low sequencing depth
[0148] The present disclosure uses the tissue transcriptome data sets of 12 known tissues to screen the tissue-specific expression gene data set, and uses the same to perform tissue tracing of the pregnant woman plasma cfDNA sample at ultra-low sequencing depth and verify the reliability of the tissue tracing.
[0149] (1) Collect the cfDNA sequencing data at ultra-low sequencing depth (BGI sequencing platform sequencing, single-end sequencing, read length 35 bp, sequencing depth about 0.1X) of 40 plasma samples of 32 pregnant women at different gestational ages.
[0150] (2) According to the steps (2)-(4) in Example 1.2, the placental tissue TCI of the pregnant woman is calculated, wherein the length of the cfDNA fragment is not limited for comparison.
[0151] (3) Calculate the fetal concentration of the pregnant woman according to the seqFF method, and calculate the correlation with the placental tissue TCI provided in (2).
[0152] FIG. 9 is a schematic diagram of the results of the tissue tracing of the pregnant woman plasma free DNA at ultra-low sequencing depth according to the present disclosure. The correlation between the fetal concentration and the placental tissue TCI is shown in FIG. 9. Even in the sample of ultra-low depth single-end sequencing, the correlation coefficient between the contribution degree (TCI) of the placental tissue obtained by the method for tracing the free DNA to the tissue provided by the present disclosure and the fetal concentration can reach 0.7 (P<0.0001). This indicates that the method for tracing the free DNA to the tissue provided by the present disclosure is stable and suitable for different sequencing platforms, sequencing modes, and sequencing depths, and has wide adaptability.
[0153] In summary, the method for tracing free DNA provided by the embodiments of the present disclosure does not require high initial plasma volume and complex experimental processing, and can be realized by only conventional cfDNA whole genome sequencing, with low sequencing depth requirement, only 1-10X, and thus relatively low cost. The method is well applied in the tracing of various tissues such as placenta, liver and bone marrow, and has been proved to have high accuracy through multiple verifications. The method is not only suitable for paired-end sequencing data, but also suitable for single-end sequencing data, and is not limited by sequencing read length, and thus has a wide application prospect. In addition, the method provided by the embodiments of the present disclosure only needs to analyze the coverage of the TSS region, and the analysis region is small, and the computing resource occupation is small, so as to ensure the efficiency and rapidity of the analysis, and greatly reduce the required time.
[0154] In addition, the terms "first", "second", "third", etc. are used only for descriptive purposes and should not be construed as indicating or implying relative importance or an indicated number of technical features. Therefore, the features defined with "first", "second", etc. can explicitly or implicitly include at least one of the features. In the description of the present disclosure, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise explicitly specified.
[0155] In the present disclosure, the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In the present specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the skilled person in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.
[0156] Although the embodiments of the present disclosure have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A method for tissue provenance of cell-free DNA, comprising: obtaining a tissue-specific expression gene set of a tissue; obtaining first distribution information of cell-free DNA originated from a sample on a specific region of each gene in the tissue-specific expression gene set; obtaining a tissue contribution index of the tissue in the cell-free DNA based on the first distribution information, to achieve tissue provenance of the cell-free DNA, wherein the specific region comprises a transcription start site (TSS) region.
2. The method of claim 1, wherein the TSS region is selected from a region of 5 Kb upstream and downstream of the TSS, preferably from a region of 3 Kb upstream and downstream of the TSS, more preferably from a region of 1.5 Kb upstream and downstream of the TSS, most preferably from a region of 1 Kb upstream and downstream of the TSS.
3. The method of claim 1 or 2, wherein, the first distribution information comprises coverage and / or sequencing depth of the cell-free DNA on the TSS region of each gene in the tissue-specific expression gene set.
4. The method of claim 3, wherein the obtaining first distribution information of cell-free DNA originated from a sample on a specific region of each gene in the tissue-specific expression gene set comprises: obtaining second distribution information, wherein the second distribution information comprises coverage and / or sequencing depth of the cell-free DNA on the TSS region of each gene in the whole genome of the sample; obtaining the first distribution information based on the second distribution information and the tissue-specific expression gene set.
5. The method of claim 4, wherein the obtaining the first distribution information and / or the obtaining the second distribution information further comprises: screening the cell-free DNA mapped to the TSS region of each gene in the tissue-specific expression gene set according to a length criterion; and / or screening the cell-free DNA mapped to the specific region in the whole genome corresponding to the sample according to a length criterion, wherein the length criterion for screening the cell-free DNA is below 600 nt, preferably, the length criterion for screening the cell-free DNA is 150-210 nt.
6. The method of any one of claims 1-5, wherein the method further comprises, after the obtaining the first distribution information and / or the obtaining the second distribution information: normalizing the first distribution information to obtain normalized first distribution information; and / or, normalizing the second distribution information to obtain normalized second distribution information.
7. The method of any one of claims 1 to 6, wherein, the obtaining a tissue-specific expression gene set of a tissue comprises: performing tissue-specific gene differential expression analysis based on gene expression data of the tissue to obtain the tissue-specific expression gene set, optionally, the tissue-specific gene differential expression analysis comprises TAU value threshold method and / or clustering analysis method.
8. The method of claim 7, wherein, the tissue-specific gene differential expression analysis comprises: determining a tissue-specific expression gene candidate set based on gene expression data of the tissue; ranking each gene in the tissue-specific expression gene set based on expression level of each gene in the tissue-specific expression gene candidate set; and selecting the tissue-specific expression genes ranked N1-N2 in the sorting to obtain a set of tissue-specific expression genes, Preferably, the N1-N2 is 1-1000, 1-500, 1-250, more preferably, the N1-N2 is 1-500 or 1-250. The statistical index of the expression level of the gene is TPM, FPKM, RPKM, CPM or reads counts, preferably, the statistical index of the expression level of the gene is TPM, and the sorting is in descending order.
9. The method of claim 5, wherein, The normalization includes: In the first distribution information or the second distribution information, based on the number of covered bases of each of the cell-free DNA aligned to each gene in the set of tissue-specific expression genes or aligned to the TSS region of the autosome corresponding to the sample, and the total number of covered bases of the TSS region, the normalization is performed to obtain the normalized first distribution information, The normalization method includes constructing a mathematical model and / or a machine learning model. Preferably, wherein the normalized first distribution information is a TSS area coverage fraction, the mathematical model is: Wherein h is the total number of TSS regions in the whole genome, Base_count(i) is the number of covered bases of the i-th TSS region, Cov_score(i) is the normalized TSS coverage score of the i-th TSS region, and Base_count(j) is the total number of covered bases of the j-th TSS region in the h TSS regions.
10. The method of claim 9, wherein, The calculation formula of the tissue contribution index is: Wherein k is the number of genes in the set of tissue-specific expression genes, Cov_score(i,t) is the TSS coverage score of the i-th tissue-specific expression gene of the t-th tissue, TCI(t) is the normalized tissue contribution index, and x is any constant greater than 0.
11. The method of any one of claims 1-10, wherein the sample is derived from at least one of plasma, seminal plasma, saliva, urine, amniotic fluid, serum, pleural effusion, cerebrospinal fluid, synovial fluid, preferably, the sample is derived from plasma.
12. The method of any one of claims 1-11, wherein the tissue is one or more, and the tissue is derived from one or more of liver, kidney, heart, pancreas, brain, lung, stomach, spleen, gallbladder, small salivary glands, colon, small intestine, large intestine, prostate, ovary, uterus, pituitary, whole blood, vagina, breast, thyroid, adrenal gland, placenta.
13. A method for suggesting a risk of tissue lesion, comprising: obtaining sequencing data of cell-free DNA of a test sample and a reference sample; calculating the tissue contribution index of one or more tissues of the test sample and the reference sample according to the method for tracing the tissue of cell-free DNA of any one of claims 1-12; and for the same tissue, based on the significant difference between the tissue contribution index of the test sample and the tissue contribution index of the reference sample, suggesting that the tissue of the sampling individual of the test sample is suspected or has a lesion. 14.The method of claim 13, wherein the one or more tissues are derived from one or more of liver, kidney, heart, pancreas, brain, lung, stomach, spleen, gallbladder, minor salivary glands, colon, small intestine, large intestine, prostate, ovary, uterus, pituitary, whole blood, vagina, breast, thyroid, adrenal gland, placenta. 15.The method of claim 13 or 14, wherein a significant difference between the tissue contribution index of the test sample and the tissue contribution index of the reference sample indicates that the tissue lesion is in at least one of a pregnancy disorder, a neurodegenerative disease, a cancer, and an autoimmune disease. 16.A method for indicating cancer risk based on cell-free DNA, comprising: obtaining cell-free DNA sequencing data from a test sample derived from a subject; calculating the tissue contribution index of one or more tissues of the test sample and a reference sample according to the method for tissue tracing of cell-free DNA of any one of claims 1-12; and wherein a significant difference between the tissue contribution index of the test sample and the tissue contribution index of the reference sample indicates that the tissue of the subject is suspected or has a cancer lesion. 17.A method for indicating pregnancy disorder risk based on cell-free DNA, comprising: obtaining cell-free DNA sequencing data from a test sample derived from a subject; calculating the tissue contribution index of placental tissue of the test sample and a reference sample according to the method for tissue tracing of cell-free DNA of any one of claims 1-12; and wherein a significant difference between the tissue contribution index of the test sample and the tissue contribution index of the reference sample indicates that the subject is suspected or has a pregnancy disorder. 18.A method for indicating recipient tolerance risk based on cell-free DNA, comprising: obtaining cell-free DNA sequencing data from a test sample derived from a subject; calculating the tissue contribution index of donor tissue of the test sample and a reference sample according to the method for tissue tracing of cell-free DNA of any one of claims 1-12; and wherein a significant difference between the tissue contribution index of the test sample and the tissue contribution index of the reference sample indicates that the subject is suspected or has a recipient intolerance. 19.A method for evaluating or prognosing disease treatment effect, comprising: obtaining cell-free DNA sequencing data from test samples derived from a subject at different time periods; calculating the tissue contribution index of donor tissue of the test samples according to the method for tissue tracing of cell-free DNA of any one of claims 1-12; and wherein the tissue contribution index of the test samples at different time periods indicates the disease treatment effect or prognosis of the subject. 20.An apparatus for tissue tracing, comprising: a first obtaining module configured to obtain a tissue-specific expression gene set of a tissue; a second obtaining module configured to obtain first distribution information of cell-free DNA derived from a sample on a specific region of each gene in the tissue-specific expression gene set; and a calculating module configured to calculate a tissue contribution index of the tissue based on the first distribution information. a third obtaining module configured to obtain a tissue contribution index of the tissue in the cell-free DNA to enable tissue provenance of the cell-free DNA.
21. A tissue provenance device comprising: a processor; a memory for storing executable instructions; wherein the processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the method of tissue provenance of cell-free DNA according to any one of claims 1 to 12, or the method of risk indication of a tissue lesion according to any one of claims 13 to 15, or the method of risk indication of cancer based on cell-free DNA according to claim 16, or the method of risk indication of a disease during pregnancy based on cell-free DNA according to claim 17, or the method of risk indication of recipient tolerance based on cell-free DNA according to claim 18, or the method of evaluation or prognosis of treatment effect of a disease according to claim 19.
22. A computer readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the method of tissue provenance of cell-free DNA according to any one of claims 1 to 12, or the method of risk indication of a tissue lesion according to any one of claims 13 to 15, or the method of risk indication of cancer based on cell-free DNA according to claim 16, or the method of risk indication of a disease during pregnancy based on cell-free DNA according to claim 17, or the method of risk indication of recipient tolerance based on cell-free DNA according to claim 18, or the method of evaluation or prognosis of treatment effect of a disease according to claim 19.
23. A computer program product, wherein the computer program product comprises a computer program, and when the computer program is executed by a processor, the computer program implements the method of tissue provenance of cell-free DNA according to any one of claims 1 to 12, or the method of risk indication of a tissue lesion according to any one of claims 13 to 15, or the method of risk indication of cancer based on cell-free DNA according to claim 16, or the method of risk indication of a disease during pregnancy based on cell-free DNA according to claim 17, or the method of risk indication of recipient tolerance based on cell-free DNA according to claim 18, or the method of evaluation or prognosis of treatment effect of a disease according to claim 19.
24. A computer program, wherein the computer program comprises computer program code, and when the computer program code is run on a computer, the computer is caused to perform the method of tissue provenance of cell-free DNA according to any one of claims 1 to 12, or the method of risk indication of a tissue lesion according to any one of claims 13 to 15, or the method of risk indication of cancer based on cell-free DNA according to claim 16, or the method of risk indication of a disease during pregnancy based on cell-free DNA according to claim 17, or the method of risk indication of recipient tolerance based on cell-free DNA according to claim 18, or the method of evaluation or prognosis of treatment effect of a disease according to claim 19.
Citation Information
Patent Citations
Prediction model for early accurate detection of preeclampsia
CN110305954A
Model for predicting gestational diabetes mellitus by using peripheral blood free DNA
CN110387414A
Method of predicting gestational related diseases based on peripheral blood free DNA high-throughput sequencing
CN110580934A
Evaluation system for predicting tissue specific source and related disease probability of cfDNA and application
CN113539355A
Screening system of tissue-specific gene marker based on peripheral blood free DNA high-throughput sequencing and application of screening system
CN115019888A