Method and device for carrying out tissue tracing on free RNA (Ribonucleic Acid)

CN120752533APending Publication Date: 2025-10-03SHENZHEN HUADA GENE INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380094805.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-18
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The prior art is difficult to achieve accurate tissue traceability of free RNA at the overall level, especially for RNAs with weak non-tissue-specific expression or expression signals, which cannot effectively trace their source tissues. The analysis results are affected by the sequencing quality of the sample, and it is impossible to compare the contributions of multiple tissues in a single sample.

Method used

By constructing a tissue characteristic gene matrix, integrating information about known high-tissue specific genes and unknown high-weight tissue important genes, and calculating tissue contribution scores, thereby achieving overall level of tissue traceability of free RNA.

Benefits of technology

It improves the accuracy and globality of free RNA tissue traceability, and can be effectively applied to the traceability of free RNA from multiple sources to multiple tissues, which has important practical application value, especially in disease diagnosis and tissue lesion risk warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120752533A_ABST
    Figure CN120752533A_ABST
Patent Text Reader

Abstract

The invention provides a tissue tracing method and device for free RNA, equipment, a readable storage medium, a computer program product and a computer program. The method for performing tissue tracing on the free RNA comprises the following steps: constructing a tissue characteristic gene matrix by utilizing one or more transcriptome data sets; calculating a tissue contribution score according to a free RNA expression profile of a sample and the tissue characteristic gene matrix; and performing tissue tracing on the free RNA according to the tissue contribution score.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for tissue tracing of free RNA Technical Field

[0001] The present disclosure relates to the field of bioinformatics, and in particular to a method and apparatus, device, readable storage medium, computer program product, and computer program for tissue tracing via free RNA. Background Art

[0002] Cell-free RNA (cfRNA) is a mixture of endogenous and exogenous RNA present outside cells in body fluids. It is derived from transcripts from various tissues. Cell-free RNA in specific tissues of patients (e.g., those with neurodegenerative diseases, pregnancy disorders, and cancer) differs from that in healthy tissues, and thus can be used to reflect the health of the source tissue. Cell-free RNA transcriptome analysis provides a potential window into the health, phenotype, and development of various human organs.

[0003] In the related art, single-cell transcriptomes of specific or representative cells in tissues are usually used instead of tissues to trace the tissue origin of free RNA. However, this method can only complete the tissue traceability of free RNA based on the single-cell level (i.e., free RNA can only be traced back to the specific or representative single cells in the tissue), and cannot achieve tissue traceability of certain non-tissue-specific expression or weak expression signals in free RNA, making it difficult to know the contribution of each tissue to the free RNA at the overall level (i.e., tissue traceability of the overall level of free RNA). In addition, in the related art, even if the tissue traceability of free RNA is performed based on the total expression of specific genes in the tissue transcriptome data, the analysis results are affected by the sequencing quality of the samples, and only the contribution of one tissue can be analyzed and compared between samples, and the contribution of different tissues of a single sample cannot be compared as a whole, nor can the contribution of multiple tissues of a sample be compared; and the calculation depends on the inherent specific gene expression characteristics, and when the in vivo environment of organs and tissues changes (such as changes in health status), the specific expression of their genes may also change accordingly, so based on the original expression characteristics, accurate tissue traceability of free RNA cannot be achieved. At the same time, related technologies still have the problem of insufficient specificity of tissue-specific genes or the inability to be detected in body fluids. Related technologies are also unable to compare the contributions of multiple tissues in a single sample individually or as a whole.

[0004] Therefore, it is urgent to propose a method to accurately trace the tissue origin of free RNA at the tissue level.

[0005] Summary of the Invention

[0006] The present disclosure aims to solve one of the technical problems in the related art at least to a certain extent.

[0007] To this end, embodiments of the present disclosure provide a method, apparatus, device, readable storage medium, computer program product, and computer program for tissue tracing of cell-free RNA. The methods provided in embodiments of the present disclosure effectively integrate information from known highly tissue-specific genes and unknown high-weight tissue-important genes, yielding accurate cell-free RNA tissue tracing results that are more consistent with actual tissue levels, and thus have broad practical application prospects.

[0008] According to an embodiment of the first aspect of the present disclosure, a method for tissue tracing of free RNA is provided, comprising: constructing a tissue-characteristic gene matrix using one or more transcriptome data sets; calculating a tissue contribution score based on the free RNA expression profile of the sample and the tissue-characteristic gene matrix; and performing tissue tracing of the free RNA based on the tissue contribution score.

[0009] In some embodiments, constructing a tissue-characteristic gene matrix using one or more transcriptome datasets includes: generating a tissue-characteristic gene set using the one or more transcriptome datasets; and constructing the tissue-characteristic gene matrix based on the tissue-characteristic gene set.

[0010] In some embodiments, generating a tissue-characteristic gene set using the one or more transcriptome datasets includes: performing tissue-characteristic classification based on the genes in the transcriptome dataset using a classification algorithm, and obtaining the weights of the genes for the tissue-characteristic classification; weight-sorting the genes according to the weights; and generating a first characteristic gene set by taking the top N genes in the weight sorting, where N is a positive integer greater than 0.

[0011] In some embodiments, N is greater than 1000.

[0012] In some embodiments, N is greater than 1500.

[0013] In some embodiments, N is greater than 2500.

[0014] In some embodiments, N is 2500.

[0015] In some embodiments, the classification algorithm is a random forest or a gradient boosted decision tree.

[0016] In some embodiments, random forests are preferred.

[0017] In some embodiments, the one or more transcriptome datasets contain specific expression profiles of each tissue, and generating a tissue-characteristic gene set using the one or more transcriptome datasets further comprises: using the specific expression profiles of each tissue in the transcriptome dataset as a second characteristic gene set.

[0018] In some embodiments, generating a tissue-characteristic gene set using the one or more transcriptome datasets further includes: obtaining sample detection genes based on the free RNA expression profile of the sample; and taking the intersection of the specific expression profile of each tissue in the transcriptome dataset and the sample detection genes to generate the second characteristic gene set.

[0019] In some embodiments, generating a tissue-characteristic gene set using the one or more transcriptome datasets further comprises: taking a union of the first characteristic gene set and the second characteristic gene set as the tissue-characteristic gene set.

[0020] In some embodiments, constructing the tissue-characteristic gene matrix based on the tissue-characteristic gene set comprises: determining the expression value of each gene in the tissue-characteristic gene set based on the one or more transcriptome datasets; and constructing the tissue-characteristic gene matrix based on the expression value of each gene.

[0021] In some embodiments, the expression value is the median of the expression values ​​of the genes in the one or more transcriptome datasets.

[0022] In some embodiments, the transcriptome dataset comprises a tissue expression dataset and / or a single-cell transcriptome dataset, and the transcriptome dataset is derived from at least one of the GTEx database, the GEO database, the EMBL-EBI database, the scRNASeqDB, the HCA (Human Cell Atlas) database, and the HPA (Human Protein Atlas) database. The transcriptome dataset is preferably derived from the GTEx database and the HPA (Human Protein Atlas) database.

[0023] In some embodiments, calculating the tissue contribution score based on the free RNA expression profile of the sample and the tissue-characteristic gene matrix includes: generating a mixing matrix based on the intersection of the tissue-characteristic gene matrix and the free RNA expression profile of the sample; and calculating the tissue contribution score based on the mixing matrix and the tissue-characteristic gene matrix.

[0024] In some embodiments, wherein the calculation of the tissue contribution score is performed using deconvolution or quadratic programming,

[0025] In some embodiments, the deconvolution algorithm is at least one of support vector regression, non-negative least squares, and non-negative matrix factorization.

[0026] In some embodiments, calculating the tissue contribution score based on the mixing matrix and the tissue-characteristic gene matrix further includes: standardizing the mixing matrix and the tissue-characteristic gene matrix; and calculating the tissue contribution score based on the standardized mixing matrix and tissue-characteristic gene matrix.

[0027] In some embodiments, the normalization is at least one of Z-score normalization, log function normalization, and maximum-minimum scaling normalization.

[0028] In some embodiments, the tissue tracing of the free RNA according to the tissue contribution score includes: determining the tissue with the highest credibility value among the tissue contribution scores corresponding to each gene in the mixing matrix as the source tissue of the gene.

[0029] In some embodiments, the credible value is the size of the tissue contribution score, the p-value of the tissue contribution score and / or the quality value of the tissue contribution score. Preferably, the p-value of the tissue contribution score is a deconvolution quality measure p-value.

[0030] In some embodiments, the sample is derived from at least one of plasma, seminal plasma, saliva, urine, amniotic fluid, serum, pleural effusion, ascites, cerebrospinal fluid, and synovial fluid.

[0031] According to an embodiment of the second aspect of the present disclosure, a method for indicating tissue lesions is provided, comprising: obtaining free RNA expression profiles of a test sample and a control sample; obtaining tissue contribution scores of each tissue of the test sample and the control sample according to the method for tissue tracing of free RNA as described in any embodiment of the first aspect; and indicating the tissue lesion of the individual from which the test sample was sampled based on the significant difference between the tissue contribution score of the test sample and the tissue contribution score of the control sample for the same tissue.

[0032] In some embodiments, the tissue is derived from one or more of liver, kidney, heart, pancreas, brain, lung, stomach, spleen, gallbladder, minor salivary glands, colon, small intestine, large intestine, prostate, ovary, uterus, pituitary gland, whole blood, vagina, breast, thyroid gland, or adrenal gland.

[0033] In some embodiments, the tissue lesion occurs in at least one of cancer, pregnancy disorders, neurodegenerative diseases, and infectious diseases.

[0034] According to an embodiment of the third aspect of the present disclosure, a tissue tracing device is provided, comprising: a characteristic gene matrix construction module, the characteristic gene matrix construction module being configured to construct a tissue characteristic gene matrix using one or more transcriptome data sets; a tissue contribution score calculation module, the tissue contribution score calculation module being configured to calculate the tissue contribution score based on the free RNA expression spectrum of the sample and the tissue characteristic gene matrix; and a tissue tracing module, the tissue tracing module being configured to perform the tissue tracing on the free RNA based on the tissue contribution score.

[0035] According to an embodiment of the fourth aspect of the present disclosure, a tissue tracing device is provided, comprising: a processor; a memory for storing executable instructions; wherein the processor is used to read the executable instructions from the memory and execute the executable instructions to implement the method for tissue tracing of free RNA as described in any embodiment of the first aspect, or the method for indicating tissue lesion risk as described in any embodiment of the second aspect.

[0036] According to an embodiment of the fifth aspect of the present disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program. When the computer program is executed by a processor, the processor implements the method for tissue tracing of free RNA as described in any embodiment of the first aspect, or the method for indicating tissue lesion risk as described in any embodiment of the second aspect.

[0037] According to an embodiment of the sixth aspect of the present disclosure, a computer program product is provided, wherein the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method for tissue tracing of free RNA as described in any embodiment of the first aspect, or the method for indicating tissue lesion risk as described in any embodiment of the second aspect.

[0038] According to an embodiment of the seventh aspect of the present disclosure, a computer program is provided, wherein the computer program includes computer program code, and when the computer program code is run on a computer, the computer executes the method for tissue tracing of free RNA as described in any embodiment of the first aspect, or the method for indicating tissue lesion risk as described in any embodiment of the second aspect.

[0039] The method for tissue tracing of free RNA provided by the embodiments of the present disclosure has the following beneficial effects:

[0040] (1) The tissue tracing method of the embodiment of the present disclosure effectively integrates the information of high-specificity genes of known tissues and high-weight genes of unknown tissues, reduces the problem of tissue gene information loss, thereby optimizing the deconvolution effect and completing the global tissue contribution calculation, thereby improving tissue tracing from the single-cell level to the tissue level, and realizing the holistic tissue tracing of free RNA.

[0041] (2) The tissue tracing method of the embodiment of the present disclosure is based on global tissue contribution calculation and has extremely high accuracy. The obtained tissue contribution score is closer to the actual free transcript level contained in the free RNA originating from various organs and tissues, which is of great significance for practical applications.

[0042] (3) It can be effectively applied to the tissue tracing of free RNA from multiple sources in various tissues, and has excellent effects in indicating the risk of tissue lesions in various diseases such as pregnancy diseases, neurodegenerative diseases, cancer and infectious diseases. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] FIG1 is a flow chart of a method for tissue tracing of free RNA according to an embodiment of the present disclosure.

[0044] [Corrected on 27.09.2023 according to Rule 91] Figure 2(a) is a schematic diagram of a random forest with a partial enlargement of the method process for tissue tracing of free RNA in an embodiment of the present disclosure; Figure 2(b) is a schematic diagram of the clustering results of the tissue-characteristic gene matrix with a partial enlargement of the method process for tissue tracing of free RNA in an embodiment of the present disclosure.

[0045] FIG3 is a schematic diagram of the steps of the method for tissue tracing of free RNA according to an embodiment of the present disclosure.

[0046] FIG4 is a schematic structural diagram of a tissue tracing device according to an embodiment of the present disclosure.

[0047] FIG5 is a schematic diagram showing the results of tissue tracing of free RNA in various body fluids according to an embodiment of the present disclosure.

[0048] FIG6 is a schematic diagram of the clustering results of tissue tracing of free RNA in various body fluids according to an embodiment of the present disclosure.

[0049] FIG7 is a schematic diagram showing a comparison of tissue contribution scores of kidney tissue to free RNA in various body fluids according to an embodiment of the present disclosure.

[0050] FIG8 is a schematic diagram showing a comparison of tissue contribution scores of prostate tissue to free RNA in various body fluids according to an embodiment of the present disclosure.

[0051] FIG9 is a schematic diagram showing a comparison of the tissue contribution scores of spleen tissue to free RNA in various body fluids according to an embodiment of the present disclosure.

[0052] FIG10 is a schematic diagram showing a comparison of tissue contribution scores of liver tissue to free RNA in various body fluids according to an embodiment of the present disclosure.

[0053] Figures 11(a) and (b) are schematic diagrams showing the tissue contribution score of liver tissue calculated based on free RNA according to an embodiment of the present disclosure;

[0054] FIG12 is a schematic diagram of liver cancer tissue lesion risk results according to an embodiment of the present disclosure.

[0055] FIG13 is a schematic diagram of the results of tissue lesion risk indicated by infectious diseases according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0056] The embodiments of the present disclosure are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to be used to explain the present disclosure, but should not be understood as limiting the present disclosure.

[0057] Cell-free RNA is a mixture of endogenous and exogenous RNA present in body fluids and outside cells, and contains transcripts from a variety of tissues. In patients with neurodegenerative diseases, pregnancy disorders, and cancer, cell-free RNA in specific tissues differs from that in healthy tissues. Therefore, cell-free RNA can be used to reflect the health of the source tissue. Cell-free RNA transcriptome analysis provides a potential window into the health, phenotype, and development of various human organs.

[0058] It's important to note that cell-free RNA exists in body fluids, making it an ideal material for non-invasive diagnosis and a viable method for early screening of various diseases. However, cell-free RNA is a mixture of transcripts from multiple tissues, and diagnostic analysis requires reliable target tissue information from this complex transcriptome.

[0059] Those skilled in the art will be aware that related technologies usually use the expression of tissue-specific genes in free RNA as the main feature for analysis. However, tissue-specific genes refer to genes that are expressed or specifically highly expressed in only one or a few tissues, and some tissue-specific genes are difficult to detect in body fluids. On the contrary, when the specificity of tissue-specific genes is insufficient, it is impossible to accurately obtain information about the target tissue. Another commonly used method is to use single-cell transcriptome data of tissue-specific cells instead of tissues. This method will lose some cells or tissues that are known to contribute to the presence of free RNA.

[0060] Based on the above background, an embodiment of the first aspect of the present disclosure provides a method for tissue tracing of free RNA, including steps S1000 to S3000.

[0061] Step S1000 : constructing a tissue-characteristic gene matrix using one or more transcriptome datasets.

[0062] Step S2000 , calculating the tissue contribution score based on the free RNA expression profile of the sample and the tissue characteristic gene matrix.

[0063] Step S3000: performing tissue tracing on the free RNA according to the tissue contribution score.

[0064] The method for tissue tracing of free RNA provided by the embodiments of the present disclosure effectively integrates the information of highly specific genes of known tissues and high-weight genes of unknown tissues, reduces the problem of tissue gene information loss, thereby optimizing the deconvolution effect and improving the accuracy of the tissue tracing method. It can be effectively applied to the tissue tracing of free RNA from multiple sources to multiple tissues, and has excellent effects in indicating the risk of tissue lesions in various diseases such as pregnancy diseases, neurodegenerative diseases, cancers and infectious diseases.

[0065] The methods provided in the present disclosure can determine the contribution of a specific tissue to the free RNA in a sample and define a "tissue contribution score" to quantify the RNA derived from a specific tissue in the free RNA extracted from the sample. Calculating this tissue contribution score allows for rapid and effective tissue tracing.

[0066] In some embodiments, the cell-free RNA expression profile may be public data, or transcriptome data obtained by extracting cell-free RNA from body fluids and sequencing.

[0067] In some embodiments, the step S1000 of constructing a tissue-characteristic gene matrix using one or more transcriptome datasets includes steps S1100 to S1200.

[0068] Step S1100 : generating a tissue-characteristic gene set using the one or more transcriptome datasets.

[0069] Step S1200 : constructing the tissue-characteristic gene matrix based on the tissue-characteristic gene set.

[0070] Among them, it can be understood that the core of the construction of the tissue characteristic gene matrix is ​​to select a tissue characteristic gene set. In the related art, only tissue-specific genes are usually used to form a tissue characteristic gene set. In the disclosed embodiment, the tissue characteristic gene set not only includes reliable tissue-specific genes derived from multiple databases, but also includes important tissue high-weight genes in the tissue multi-classification task obtained by the classification algorithm. This effectively integrates the information of known tissue high-specificity genes and unknown tissue high-weight genes, reduces the problem of tissue gene information loss, and thus optimizes the subsequent tissue contribution score calculation effect.

[0071] In some embodiments, the step S1100 of generating a tissue-characteristic gene set using the one or more transcriptome datasets includes steps S1110 to S1112 .

[0072] Step S1110 , performing tissue characteristic classification based on the genes in the transcriptome dataset using a classification algorithm, and obtaining the weight of the gene for the tissue characteristic classification.

[0073] Step S1111 , sorting the genes by weight according to the weights.

[0074] Step S1112 , taking the top N genes in the weight ranking to generate a first characteristic gene set.

[0075] In some embodiments, N is a positive integer greater than 0. In some embodiments, N is greater than 1000. In some embodiments, N is greater than 1500. In some embodiments, N is greater than 2500.

[0076] It should be noted that if the number of genes included is too small, model performance will be seriously affected. Specifically, when N is a positive integer in the range of 1800-3000, the model performance is good. When N is the optimal value of 2500, the model performance is the best.

[0077] Those skilled in the art will appreciate that the classification algorithm, based on the mixed tissue gene expression signatures in the transcriptome data set, is used to successfully classify multiple independent tissues for the purpose of performing multi-classification tasks. The genes are sorted according to the weights of the various gene signatures under the default parameters, i.e., the high-weight genes of the tissues that are important in the multi-classification tasks of the tissues are obtained. Wherein, the default parameters are the default parameters of the RandomForestClassifier in the sklearn package.

[0078] It is understood that some tissues have very similar anatomical locations and transcriptome expression profiles, and one of these tissues can be used to represent other similar tissues. In some embodiments, the expression profiles of brain tissues in the cerebral cortex, hippocampus, cerebellum, and hypothalamus are very consistent, so the cerebral cortex can be used as a representative brain tissue.

[0079] In some embodiments, the classification algorithm is a random forest or a gradient boosted decision tree. In some embodiments, the classification algorithm is a random forest.

[0080] In some embodiments, tissue classification is performed based on the genes in the transcriptome dataset using a random forest model, and the weights of the genes for the tissue characteristic classification are obtained.

[0081] In some embodiments, the transcriptome dataset is processed into a sn*gn matrix E and a sn*1 matrix T, where sn is the number of tissue samples and gn is the number of genes. Matrix E represents the expression profiles of all tissue samples, and T represents the tissue type labels of the corresponding samples. In some embodiments, 27 tissue types are labeled 0-26. Tissue classification is then performed using the RandomForestClassifier method in the sklearn.ensemble package. In some embodiments, the function parameters are explained as follows:

[0082] rf = RandomForestClassifier(criterion = "entropy", class_weight = "balanced"); rf.fit(X_train, y_train), where X_train is the matrix E and y_train is the matrix T. Once tissue classification is complete, obtain the weights rf.feature_importances for the genes contributing to the tissue's characteristic classification. Sort the genes by their importance weights in rf.feature_importances_ and select the top 2500 genes with the highest weights. These are the most important characteristic genes for the tissue multi-classification task, generating the first characteristic gene set.

[0083] In some embodiments, the one or more transcriptome datasets include specific expression profiles of each tissue, and step S1100 of generating a tissue-specific gene set using the one or more transcriptome datasets further includes step S1120.

[0084] Step S1120 , using the specific expression profile of each tissue in the transcriptome dataset as a second characteristic gene set.

[0085] In some embodiments, the step S1100 of generating a tissue-characteristic gene set using the one or more transcriptome datasets further includes steps S1130 to S1140.

[0086] Step S1130, obtaining sample detected genes according to the free RNA expression profile of the sample; and

[0087] Step S1140 , obtaining the intersection of the specific expression profile of each tissue in the transcriptome dataset and the detected genes of the sample to generate the second characteristic gene set.

[0088] It should be noted that the sample-detected genes are genes that can be detected in free RNA extracted from bodily fluid samples. Specifically, when obtaining free RNA expression profiles, different cfRNA sequencing technologies lead to different sequencing quality. Those skilled in the art can choose whether to filter the sample-detected genes for stable detection genes based on the sequencing technology. In some embodiments, stable detection genes refer to genes that can be detected in at least 20% of the samples in the free RNA expression profile.

[0089] In some embodiments, the step S1100 of generating a tissue-characteristic gene set using the one or more transcriptome datasets further includes step S1150.

[0090] Step S1150 , taking the union of the first characteristic gene set and the second characteristic gene set as the tissue characteristic gene set.

[0091] In some embodiments, step S1200 of constructing the tissue-characteristic gene matrix based on the tissue-characteristic gene set includes steps S1210 to S1220.

[0092] Step S1210 , determining the expression value of each gene in the tissue-characteristic gene set based on the one or more transcriptome datasets.

[0093] Step S1220: constructing the tissue-characteristic gene matrix based on the expression values ​​of the genes.

[0094] In some embodiments, the expression value is the median of the expression values ​​of the genes in the one or more transcriptome datasets.

[0095] In some embodiments, the transcriptome dataset comprises a tissue expression dataset and / or a single-cell transcriptome dataset, and the transcriptome dataset is derived from at least one of the GTEx database, the GEO database, the EMBL-EBI database, the scRNASeqDB, the HCA (Human Cell Atlas) database, and the HPA (Human Protein Atlas) database.

[0096] In some embodiments, the transcriptome dataset is preferably derived from the GTEx database and the HPA (Human Protein Atlas) database.

[0097] In some embodiments, the step S2000 of calculating the tissue contribution score based on the free RNA expression profile of the sample and the tissue-characteristic gene matrix includes steps S2100 to S2200.

[0098] Step S2100 : generating a mixing matrix based on the intersection of the tissue-characteristic gene matrix and the cell-free RNA expression profile of the sample.

[0099] Step S2200 , calculating the tissue contribution score according to the mixing matrix and the tissue characteristic gene matrix.

[0100] In some embodiments, the method uses deconvolution or quadratic programming to calculate the tissue contribution score.

[0101] In some embodiments, the deconvolution algorithm is at least one of support vector regression, non-negative least squares, and non-negative matrix factorization.

[0102] In some embodiments, the step S2200 of calculating the tissue contribution score based on the mixing matrix and the tissue-characteristic gene matrix further includes steps S2210 to S2220.

[0103] Step S2210: standardize the mixing matrix and the tissue-characteristic gene matrix.

[0104] Step S2220 , calculating the tissue contribution score based on the standardized mixing matrix and the tissue-characteristic gene matrix.

[0105] In some embodiments, the normalization is at least one of Z-score normalization, log function normalization, and maximum-minimum scaling normalization.

[0106] Those skilled in the art will be able to adjust the normalization method based on the type of algorithm used and its effectiveness. In some embodiments, regression and matrix decomposition calculations are performed directly using the mixing matrix and the characteristic expression matrix based on non-negative least squares and non-negative matrix factorization; and Z-score normalization is performed on the characteristic expression matrix and the mixing matrix based on support vector regression.

[0107] In some embodiments, the step S3000 of performing tissue tracing on the free RNA according to the tissue contribution score includes step S3100.

[0108] Step S3100 : determining the tissue with the highest credibility value among the tissue contribution scores corresponding to each gene in the mixing matrix as the source tissue of the gene.

[0109] In some embodiments, the confidence value is the size of the tissue contribution score, the p-value of the tissue contribution score, and / or the quality value of the tissue contribution score.

[0110] In some specific embodiments, the specific process of the deconvolution is:

[0111] Assume that the free RNA expression profile is S, the tissue characteristic matrix generated above is A, and F is the contribution score of each tissue. Then the formula A*F=S is obtained.

[0112] Deconvolution is performed using support vector regression, with inputs A and S, to solve for F. Matrix A is a representative basis matrix of size gn × tc, where gn represents the number of genes and tc represents the number of tissue types. Matrix A represents the characteristic gene expression profiles for these tc tissue types. Vector F is a vector of size tc × 1, representing the contribution of each tissue. Vector S is the gene expression measurements observed in body fluids (cell-free RNA expression profiles) and is of size gn × 1.

[0113] In some specific embodiments, the NuSVR method in the sklearn.svm package of Python is used to perform deconvolution of support vector regression. The specific code is as follows:

[0114] The solver is initialized by the code clf=NuSVR(C=c,nu=nu,kernel="linear"), where c can take multiple values. The preset parameters are list_c=[0.5,0.8,1,1.5,2]; list_nu=[0.05,0.1,0.15,0.25,0.5,0.75,0.85];

[0115] The solver is solved by the code clf.fit(A,S'), and each combination of the above parameters will generate a corresponding preliminary tissue contribution score.

[0116] Based on the preliminary tissue contribution score, the predicted expression profile S_pred of the free RNA was calculated. The coefficients Coefficients solved by the solver and the predicted expression profile S_pred of the free RNA were used to calculate the mean square error. The preliminary tissue contribution score corresponding to the combination with the smallest mean square error was used as the tissue contribution score. The P value calculated by hypothesis testing was used as the confidence value for the tissue tracing.

[0117] In some embodiments, the sample is derived from at least one of plasma, seminal plasma, saliva, urine, amniotic fluid, serum, pleural effusion, ascites, cerebrospinal fluid, and synovial fluid.

[0118] A second aspect of the present disclosure provides a method for indicating tissue lesion risk, comprising:

[0119] Obtaining free RNA expression profiles of the test sample and the control sample;

[0120] Obtaining tissue contribution scores of each tissue of the test sample and the control sample according to the method for tissue tracing of free RNA according to any embodiment of the first aspect; and

[0121] For the same tissue, if the tissue contribution score of the test sample is significantly different from that of the control sample, it indicates that the tissue of the individual from which the test sample was sampled is lesional.

[0122] In some embodiments, the tissue is derived from one or more of liver, kidney, heart, pancreas, brain, lung, stomach, spleen, gallbladder, minor salivary glands, colon, small intestine, large intestine, prostate, ovary, uterus, pituitary gland, whole blood, vagina, breast, thyroid gland, or adrenal gland.

[0123] In some embodiments, the tissue lesion occurs in at least one of a pregnancy disorder, a neurodegenerative disease, a cancer, and an infectious disease.

[0124] In some embodiments, the cancer is liver cancer.

[0125] In some embodiments, the infectious disease is HBV infection.

[0126] The embodiment of the forty-third aspect of the present disclosure provides a tissue tracing device, including a characteristic gene matrix construction module, a tissue contribution score calculation module and a tissue tracing module.

[0127] The characteristic gene matrix construction module is configured to construct a tissue characteristic gene matrix using one or more transcriptome data sets.

[0128] A tissue contribution score calculation module is configured to calculate the tissue contribution score according to the free RNA expression profile of the sample and the tissue characteristic gene matrix.

[0129] A tissue tracing module is configured to perform tissue tracing on the free RNA according to the tissue contribution score.

[0130] An embodiment of the fourth aspect of the present disclosure provides a tissue tracing device, comprising: a processor; a memory for storing executable instructions; wherein the processor is used to read the executable instructions from the memory and execute the executable instructions to implement the method for tissue tracing of free RNA as described in any embodiment of the first aspect of the present disclosure, or the method for indicating tissue lesion risk as described in any embodiment of the second aspect.

[0131] An embodiment of the fifth aspect of the present disclosure provides a computer-readable storage medium, wherein the storage medium stores a computer program. When the computer program is executed by a processor, the processor implements the method for tissue tracing of free RNA as described in any embodiment of the first aspect, or the method for indicating tissue lesion risk as described in any embodiment of the second aspect.

[0132] An embodiment of the sixth aspect of the present disclosure provides a computer program product, wherein the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method for tissue tracing of free RNA as described in any embodiment of the first aspect, or the method for indicating tissue lesion risk as described in any embodiment of the second aspect.

[0133] The seventh aspect of the present disclosure provides a computer program, wherein the computer program includes computer program code, and when the computer program code is run on a computer, the computer executes the method for tissue tracing of free RNA as described in any embodiment of the first aspect, or the method for indicating tissue lesion risk as described in any embodiment of the second aspect.

[0134] In some embodiments, the present disclosure provides a method for tissue tracing of free RNA, a tree-based tissue tracing method, a tree model based Algorithm for CEll-free transcriptome origin analysis (TrACe). The TrACe method proposed in the present disclosure simultaneously considers tissue-specific genes and overall tissue expression to generate a characteristic gene set for tracing, and calculates the contribution scores of various tissues in the free RNA of body fluids by a deconvolution method. This solves the problems of the above-mentioned technology losing certain tissue signals and being unable to perform global tissue contribution calculations.

[0135] The following examples will explain in detail the methods for tissue tracing of free RNA proposed in this disclosure and their related applications. It should be noted that the experimental methods in the following examples, unless otherwise specified, are conventional methods and were performed according to the techniques or conditions described in literature in the field or according to product specifications. The materials and reagents used in the following examples, unless otherwise specified, can all be obtained from commercial sources.

[0136] Unless otherwise specified, the quantitative tests in the following examples were performed three times, and the results were averaged.

[0137] Example 1

[0138] In this example, the tissue transcriptome data of 27 known tissue types in the GTEX database were used as the training set. High-weight tissue genes were screened through a random forest model, and the tissue-specific genes in the HPA database were combined with the high-weight tissue genes to construct a tissue-characteristic gene matrix. Finally, the tissue contribution score in the free RNA was calculated through deconvolution to achieve convenient and accurate tissue tracing.

[0139] 1.1 Construction of tissue-specific gene matrix

[0140] (1) Download the GTEX_V8 dataset (https: / / www.gtexportal.org / home);

[0141] (2) Screening mRNA genes in the GTEX_V8 dataset as GTEX genes;

[0142] (3) The GTEX_V8 dataset was screened, and the screening criterion was that the number of tissue samples was greater than 80;

[0143] (4) The 27 tissues obtained in step (3), namely, kidney_cortex, lung, whole_blood stomach, vagina, nerve_tibial, prostate, uterus, pituitary, heart_left_ventricle, liver, testis, pancreas, minor_salivary_gland, artery_aorta, colon_sigmoid, ovary, esophagus_mucosa, adipose_subcutaneous, small_intestine_terminal_ileum, brain_cortex ,skin_sun_exposed_lower_leg,muscle_skeletal,breast_mammary_tissue,spleen,thyroid,adrenal_gland corresponding to the kidney, lung, whole blood, stomach, vagina, tibial nerve, prostate, uterus, pituitary, left ventricle, liver, testis, pancreas, minor salivary gland, aorta, colon, ovary, esophageal mucosa, subcutaneous fat, ileum, cerebral cortex, skin, skeletal muscle, breast tissue, spleen, thyroid and adrenal tissue transcriptome data were intersected with GTEX genes to obtain GTEX gene expression profiles corresponding to 27 tissues;

[0144] (5) The obtained GTEX gene expression profiles were processed into a sn*gn matrix E and a sn*1 matrix T, where sn is the number of tissue samples and gn is the number of genes. Matrix E is the expression profile of all tissue samples, and T is the tissue type label of the corresponding sample, with the 27 tissue types labeled 0-26.

[0145] (6) Using the random forest model to successfully classify 27 independent tissues, the GTEx gene expression profiles of the 27 tissues obtained in step (4) were subjected to a multi-classification task. The specific code is: rf = RandomForestClassifier(criterion = "entropy", class_weight = "balanced"); rf.fit(X_train, y_train),

[0146] Where X_train is the matrix E, y_train is the matrix T;

[0147] (7) Obtaining the weight rf.feature_importances of the gene for the tissue characteristic classification, and sorting the gene features according to the weight;

[0148] (8) Select the top 2500 gene features in the weight value ranking as tissue high-weight genes;

[0149] (9) constructing the first characteristic gene set by taking the median gene expression value of multiple samples in each tissue as the tissue high-weight gene expression value of the tissue;

[0150] (10) Download the tissue expression detect in single gene set from the HPA (Human Protein Atlas) database (https: / / www.proteinatlas.org / search / tissue_category_rna:any;Detected%20in%20single%20AND%20sort_by:tissue%20specific%20score);

[0151] (11) Screening tissue expression detect in single gene set for tissue-specific genes corresponding to the 27 tissue types obtained in step (3);

[0152] (12) taking the intersection of the tissue-specific genes obtained in step (11) and the genes detected from the free RNA of the sample as the second characteristic gene set;

[0153] (13) Merge the vectors of the first characteristic gene set and the second characteristic gene set to construct a tissue characteristic gene matrix.

[0154] The TrACe method of this embodiment constructs a tissue-characteristic gene matrix based on the unknown tissue high-weight genes and known tissue high-specificity genes screened by the random forest model, effectively avoiding the problem of tissue gene information loss and providing a good foundation for the next step of tissue contribution score calculation.

[0155] 1.2 Organization traceability

[0156] (1) Download the free RNA expression profiles of 20 samples (4 samples from each body fluid, including plasma, seminal plasma, saliva, urine, and amniotic fluid) (doi:10.1002 / ctm2.987);

[0157] (2) The intersection of the free RNA expression profile and the genes in the tissue-characteristic gene matrix is ​​taken to form a mixing matrix. Let the mixing matrix be S, which is the expression vector of the tissue-characteristic genes in the body fluid;

[0158] (3) Let the tissue characteristic gene matrix be A, and let S' be the expression vector of each individual in the expression spectrum. Use the nuSVR function to calculate each sample separately. The calculation formula is: A*F=S'

[0159] By using the above formula, deconvolution is performed based on the mixing matrix and the tissue characteristic gene matrix to solve the initial tissue contribution score F of each tissue. The specific implementation code is:

[0160] Here, matrix A is a representative matrix of size gn × tc, where gn represents the number of genes and tc represents the number of tissue types. Matrix A represents the characteristic gene expression profiles of the aforementioned tc tissue types. Vector F is a vector of size tc × 1, representing the contribution of each tissue. S is a matrix of gene expression measurements observed in body fluids (characteristic genes in the free RNA expression profile), of size gn × 1.

[0161] The c value can take multiple values. The preset parameters list_c = [0.5, 0.8, 1, 1.5, 2]; list_nu = [0.05, 0.1, 0.15, 0.25, 0.5, 0.75, 0.85]. Solved by clf.fit(A, S'), each set of parameter combinations will generate a corresponding preliminary tissue contribution score, expressed as coefficients;

[0162] (4) calculating the predicted expression profile S_pred of free RNA based on the preliminary tissue contribution score;

[0163] (5) The coefficients Coefficients solved by the solver and the free RNA predicted expression profile S_pred are used to calculate the mean square error, where the preliminary tissue contribution score corresponding to the combination with the smallest mean square error is taken as the absolute tissue contribution score;

[0164] (6) Based on the above calculation results, calculate the P value through hypothesis testing and use it as the credible value;

[0165] (7) normalizing the tissue contribution score to calculate the tissue relative contribution score, and completing tissue traceability based on the trustworthy value;

[0166] (8) Unsupervised clustering based on tissue contribution scores of free RNA.

[0167] The results of tissue tracing of free RNA from various body fluids in this embodiment are shown in Figure 5. The clustering results of tissue tracing of free RNA from various body fluids in this embodiment are shown in Figure 6. According to Figure 6, the types of body fluids can be distinguished, and the contribution scores of different tissues in various body fluids can be compared to determine the dominant contributing tissues of different body fluids. For example, the dominant contributing tissue of semen samples is prostate tissue. This result is consistent with prior common sense in this field and illustrates the reliability of the tissue tracing method provided in this embodiment.

[0168] The tissue contribution scores of kidney tissue to free RNA in various body fluids of this example are shown in Figure 7. The proportion of RNA in urine from kidney tissue is the highest, indicating that urine is suitable for non-invasive diagnosis of kidney disease.

[0169] The results of the tissue contribution of prostate tissue to free RNA in various body fluids in this example are shown in Figure 8. The prostate has the highest proportion of RNA in semen, suggesting that semen is suitable for non-invasive diagnosis of prostate diseases.

[0170] The results of the tissue contribution of spleen tissue to free RNA in various body fluids in this example are shown in Figure 9. The spleen has the highest proportion of RNA in plasma, suggesting that plasma is suitable for non-invasive diagnosis of spleen diseases.

[0171] The results of the tissue contribution of liver tissue to free RNA in various body fluids in this example are shown in Figure 10. The liver has the highest proportion of RNA in plasma, suggesting that plasma is suitable for non-invasive diagnosis of liver diseases.

[0172] In this embodiment, the TrACe method calculates the tissue contribution score through deconvolution, which can be effectively applied to tissue tracing of free RNA from multiple sources to multiple tissues.

[0173] Example 2

[0174] In this example, the tissue contribution fraction of liver tissue in plasma samples from healthy individuals and liver cancer patients was calculated using the TrACe method and compared.

[0175] (1) Download the free RNA expression profiles of 20 plasma samples from healthy individuals and 8 plasma samples from patients with liver cancer (https: / / doi.org / 10.1038 / s41698-022-00270-y);

[0176] (2) Obtaining the relative tissue contribution score of liver tissue using the TrACe tissue tracing method in Example 1;

[0177] (3) Statistical test was performed on the relative tissue contribution scores of liver tissues of healthy individuals and liver cancer patients.

[0178] Figure 12 shows a comparison of the tissue contribution scores for liver tissue from healthy individuals and liver cancer patients. Significant differences were observed between the tissue contribution scores of liver tissue from healthy individuals and lung tissue from liver cancer patients, demonstrating that tissue-derived RNA analysis of plasma can indicate the risk of tissue lesions. The tissue contribution scores calculated by TrACe in this example are highly effective in indicating the risk of tissue lesions in a variety of diseases.

[0179] Example 3

[0180] In this example, the tissue contribution fraction of liver tissue in plasma samples from healthy individuals and HBV-infected patients was calculated using the TrACe method and compared.

[0181] (1) Quantitative analysis of cell-free RNA in plasma samples from healthy individuals and HBV-infected patients was performed to obtain the cell-free RNA expression profile.

[0182] (2) Obtaining the relative tissue contribution score of liver tissue using the TrACe tissue tracing method in Example 1;

[0183] (3) Statistical test was performed on the relative tissue contribution scores of liver tissues of healthy individuals and HBV-infected patients.

[0184] The results of the comparison of the tissue contribution scores of liver tissue from healthy individuals and HBV-infected patients are shown in Figure 13. The tissue contribution score of liver tissue from healthy individuals to plasma is 0.030284967, while the tissue contribution score of liver tissue from HBV-infected patients to plasma is 0.118940404, which is a significant difference. This shows that tissue tracing analysis of plasma free RNA can indicate tissue lesions of infectious diseases. The tissue contribution score calculated by TrACe in this example has an excellent effect in indicating the risk of tissue lesions of infectious diseases.

[0185] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0186] In the present disclosure, the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0187] Although the embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are illustrative and are not to be construed as limitations on the present disclosure. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present disclosure.

Claims

1. A method for tissue tracing of free RNA, comprising: Using one or more transcriptome datasets to construct a tissue-specific gene matrix; Calculating tissue contribution scores according to the free RNA expression profile of the sample and the tissue characteristic gene matrix; and The tissue origin of the free RNA is traced according to the tissue contribution score.

2. The method according to claim 1, wherein: The method of constructing a tissue characteristic gene matrix using one or more transcriptome data sets includes: generating a tissue-characteristic gene set using the one or more transcriptome datasets; The tissue-characteristic gene matrix is ​​constructed based on the tissue-characteristic gene set.

3. The method according to claim 2, wherein: The step of generating a tissue characteristic gene set using the one or more transcriptome data sets comprises: Using a classification algorithm, classify the tissue characteristics according to the genes in the transcriptome data set, and obtain the weight of the gene for the tissue characteristics classification; Rank the genes according to the weights; and The top N genes in the weight ranking are taken to generate a first characteristic gene set, wherein N is a positive integer greater than 0, preferably, N is greater than 1000, more preferably greater than 1500, more preferably greater than 2500, and most preferably N is 2500. Wherein, the classification algorithm is random forest or gradient boosting decision tree, preferably random forest.

4. The method according to claim 3, wherein: The one or more transcriptome data sets contain specific expression profiles of each tissue, and the step of generating a tissue characteristic gene set using the one or more transcriptome data sets further comprises: The specific expression profile of each tissue in the transcriptome dataset is used as the second characteristic gene set.

5. The method according to claim 4, wherein: The step of generating a tissue characteristic gene set using the one or more transcriptome data sets further comprises: Obtaining sample detection genes according to the free RNA expression profile of the sample; and The second characteristic gene set is generated by taking the intersection of the specific expression profile of each tissue in the transcriptome data set and the detected genes of the sample.

6. The method according to claim 5, wherein: The step of generating a tissue characteristic gene set using the one or more transcriptome data sets further comprises: The union of the first characteristic gene set and the second characteristic gene set is taken as the tissue characteristic gene set.

7. The method according to claim 2, wherein: Constructing the tissue-characteristic gene matrix based on the tissue-characteristic gene set includes: Determining the expression value of each gene in the tissue-characteristic gene set based on the one or more transcriptome data sets; and The tissue characteristic gene matrix is ​​constructed based on the expression value of each gene. Optionally, the expression value is the median of the expression values ​​of each gene in the one or more transcriptome datasets.

8. The method according to claim 1, wherein: The transcriptome dataset comprises a tissue expression dataset and / or a single-cell transcriptome dataset, and the transcriptome dataset is derived from at least one of the GTEx database, the GEO database, the EMBL-EBI database, the scRNASeqDB, the HCA (Human Cell Atlas) database, and the HPA (Human Protein Atlas) database, and the transcriptome dataset is preferably derived from the GTEx database and the HPA (Human Protein Atlas) database.

9. The method according to claim 1, wherein: The step of calculating the tissue contribution score according to the free RNA expression profile of the sample and the tissue characteristic gene matrix comprises: generating a hybrid matrix based on the intersection of the tissue-characteristic gene matrix and the free RNA expression profile of the sample; and The tissue contribution score is calculated based on the mixing matrix and the tissue-characteristic gene matrix.

10. The method according to claim 9, wherein the calculation of the tissue contribution score is performed using deconvolution or quadratic programming, Optionally, the deconvolution algorithm is at least one of support vector regression, non-negative least squares and non-negative matrix decomposition.

11. The method according to claim 9, wherein: The step of calculating the tissue contribution score according to the mixing matrix and the tissue characteristic gene matrix further comprises: normalizing the mixing matrix and the tissue-characteristic gene matrix; and The tissue contribution score is calculated based on the standardized mixing matrix and the tissue characteristic gene matrix, Optionally, the normalization is at least one of Z-score normalization, log function normalization and maximum-minimum value scaling normalization.

12. The method according to any one of claims 1 to 11, wherein: The tissue tracing of the free RNA according to the tissue contribution score comprises: The tissue with the highest credibility value among the tissue contribution scores corresponding to each gene in the mixing matrix is ​​determined as the source tissue of the gene, Optionally, the credible value is the size of the tissue contribution score, the p-value of the tissue contribution score and / or the quality value of the tissue contribution score. Preferably, the p-value of the tissue contribution score is a deconvolution quality measure p-value.

13. The method according to claim 1, wherein the sample is derived from at least one of plasma, seminal plasma, saliva, urine, amniotic fluid, serum, pleural effusion, ascites, cerebrospinal fluid, and synovial fluid.

14. A method for indicating tissue lesion risk, comprising: Obtaining free RNA expression profiles of the test sample and the control sample; Obtaining tissue contribution scores of each tissue of the test sample and the control sample according to the method for tissue tracing of free RNA according to any one of claims 1 to 13; And for the same tissue, based on the fact that the tissue contribution score of the test sample is significantly different from that of the control sample, it is suggested that the tissue of the individual from which the test sample was sampled is lesion.

15. The method according to claim 14, wherein the tissue is derived from one or more of liver, kidney, heart, pancreas, brain, lung, stomach, spleen, gallbladder, minor salivary glands, colon, small intestine, large intestine, prostate, ovary, uterus, pituitary, whole blood, vagina, breast, thyroid, and adrenal glands.

16. The method of claim 14 or 15, wherein the tissue lesion occurs in at least one of a pregnancy disorder, a neurodegenerative disease, a cancer, and an infectious disease.

17. A tissue tracing device, comprising: A characteristic gene matrix construction module, wherein the characteristic gene matrix construction module is configured to construct a tissue characteristic gene matrix using one or more transcriptome data sets; A tissue contribution score calculation module, wherein the tissue contribution score calculation module is configured to calculate the tissue contribution score according to the free RNA expression profile of the sample and the tissue characteristic gene matrix; and A tissue tracing module is configured to perform tissue tracing on the free RNA according to the tissue contribution score.

18. A tissue tracing device, comprising: processor; A memory for storing executable instructions; Wherein, the processor is used to read the executable instructions from the memory and execute the executable instructions to implement the method for tissue tracing of free RNA as described in any one of claims 1 to 13, or the method for indicating tissue lesion risk as described in claim 14.

19. A computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the method for tissue tracing of free RNA as described in any one of claims 1 to 13, or the method for indicating tissue lesion risk as described in claim 14.

20. A computer program product, wherein the computer program product comprises a computer program, and when the computer program is executed by a processor, the computer program implements the method for tissue tracing of free RNA as described in any one of claims 1 to 13, or the method for indicating tissue lesion risk as described in claim 14.

21. A computer program, wherein the computer program comprises computer program code, and when the computer program code is run on a computer, the computer executes the method for tissue tracing of free RNA as described in any one of claims 1 to 13, or the method for indicating tissue lesion risk as described in claim 14.

Citation Information

Patent Citations

  • Methods for detecting disease using analysis of RNA

    CN113286883A

  • Methods for analysis of target molecules in biological fluids

    US20230086722A1