Construction and application of tissue source prediction model

By constructing a tissue source prediction model, using the third-generation sequencing data and RNA sequencing data processing methods, combined with machine learning models, we solve the accuracy and efficiency problems of source determination in tumor diagnosis at unknown primary sites, and achieve more efficient and accurate tumor source prediction.

CN119943440AActive Publication Date: 2025-05-06FUDAN UNIV SHANGHAI CANCER CENT
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510009181.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The prior art has problems with accuracy and efficiency in diagnosing tumors at unknown primary sites (CUPs). Especially in determining the source of the tumor, routine histological examinations have limitations and require the assistance of diagnosis with technical means such as molecular genetics.

Method used

By constructing a tissue source prediction model, the reference GTF files and tissue-specific RNA transcripts of full-length RNA transcripts are obtained using the third-generation sequencing data and RNA sequencing data processing, and training is combined with machine learning models (such as random forests, decision trees, XGBoost, etc.), and a sub-model is used to determine whether the sample to be tested is derived from a specific tissue, and finally the tissue source prediction model is constructed.

Benefits of technology

It improves the diagnostic accuracy and efficiency of tumors at unknown primary sites, can more accurately predict the primary tissue of the tumor, provides more targeted treatment strategies, and enhances the comparability and consistency of the diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005227848240000101
    Figure BDA0005227848240000101
  • Figure BDA0005227848240000111
    Figure BDA0005227848240000111
  • Figure BDA0005227848240000171
    Figure BDA0005227848240000171
Patent Text Reader

Abstract

The invention provides construction and application of a tissue source prediction model. Specifically, the invention develops a model for tissue source prediction or diagnosis, including primary tissue prediction of cancer; and a system based on the tissue source prediction model is developed. The tissue source prediction model can efficiently and accurately judge the source part of the to-be-detected tissue or the auxiliary diagnosis result of the primary part tissue of the to-be-detected tumor tissue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical diagnosis, and in particular, relates to the construction and application of a tissue origin prediction model. Background Art

[0002] Tumor pathology diagnosis is one of the gold standards for diagnosing many types of tumors and is highly accurate. Through tumor pathology diagnosis, the nature of the tumor can be clarified, the source of the tumor can be determined, the tumor can be staged and classified, and its degree of differentiation can be evaluated. By observing and analyzing the histological characteristics, grading, and staging of the tumor, the patient's survival rate and recurrence risk can be predicted, providing an important reference for treatment decisions and patient management, thereby guiding individualized treatment strategies.

[0003] However, different types of tumors may have similar histological features. Sometimes a patient may have multiple primary tumors or multiple independent malignant lesions at the same time. In addition, the origin of some metastatic tumors may not be determined by routine histological examination, such as Cancer of Unknown Primary (CUP). In some cases, pathology has certain limitations in diagnosing primary tumors and / or metastases, and molecular genetics and other technical means are needed to assist in diagnosis.

[0004] CUP refers to a metastatic malignant tumor that is histologically confirmed but not found after a comprehensive and detailed examination of the primary site. CUP is characterized by early metastasis, rapid progression, multiple organ involvement in more than half of the patients, unknown metastatic pattern, and high mortality. The clinical treatment of CUP patients is mainly based on traditional chemotherapy, but empirical chemotherapy is not effective, lacks specificity, and fails to improve the patient's survival prognosis.

[0005] The tracing of the primary lesion (including: tumor type and tissue origin) is of great significance for the selection of CUP populations that may benefit from targeted therapy. The pathological diagnosis of CUP is a complex task. For example, it is difficult to determine the location of the primary tumor in CUP. There may be multiple potential primary sites, including different organs and tissues, and the clinical manifestations are not specific. Faced with the diagnostic difficulties of CUP, doctors usually use a combination of multiple diagnostic methods, including detailed medical history, physical examination, imaging examination, pathological analysis, immunohistochemical staining and molecular biological testing. At the same time, work with a multidisciplinary team, including pathologists, radiologists, internists and oncologists, to work together to improve the diagnostic accuracy and treatment efficacy of tumors of unknown primary site.

[0006] In addition, tumor pathology diagnosis is somewhat subjective and requires high professional knowledge and experience of doctors. Different pathologists may have different interpretations and judgments on the same tissue specimen. This may lead to differences in consistency of diagnosis between different hospitals and different experts, thereby affecting the comparability of treatment decisions and results. For this reason, the technicians of the present invention recognize the need to obtain more specific markers or biological characteristics of tissues from the molecular biology level to improve the accuracy of diagnosis.

[0007] Therefore, there is an urgent need in this field to develop a new, more accurate and efficient cancer auxiliary diagnosis method or equipment, including the diagnosis of the origin of cancer tissue. Summary of the invention

[0008] The present invention provides a new method for predicting tissue origin.

[0009] In a first aspect of the present invention, a method for constructing a tissue source prediction model is provided, comprising the steps of:

[0010] (s1) Process the third-generation sequencing data and RNA sequencing (RNA-seq) data of cancer tissue samples, paracancerous tissue samples, and normal tissue samples from different tissue sources to obtain reference GTF files of full-length RNA transcripts and tissue-specific RNA transcripts (Ti-SRT, Tissue Specific RNA Transcript) of different tissues;

[0011] (s2) providing training set, test set and validation set data for model construction, wherein the data includes expression information of full-length RNA transcripts of RNA-seq data obtained according to the reference GTF file after RNA sequencing of tissue samples from different tissue sources; the expression information includes expression information of tissue-specific RNA transcripts of the different tissues;

[0012] Wherein, for a certain tissue, the expression information of the tissue-specific RNA transcripts of the tissue is the expression information of the tissue-specific RNA transcripts of the tissue-positive sample, and the expression information of the tissue-specific RNA transcripts of other tissues for the tissue is the expression information of the tissue-specific RNA transcripts of the tissue-negative sample;

[0013] (s3) Construction of tissue origin sub-model: constructing different sub-models for different tissues to determine whether the sample to be tested originates from the tissue; comprising the steps of: for any tissue, inputting the expression information of tissue-specific RNA transcripts of tissue-positive samples and tissue-negative samples for the tissue into different machine learning models for training; thereby obtaining a tissue origin prediction sub-model for determining whether the sample to be tested originates from the tissue; and using this step, obtaining tissue origin prediction sub-models for all tissues;

[0014] The machine learning models include: Random Forest model, Decision Tree model, XGBoost model, LightGBM model and CatBoost model;

[0015] (s4) Construction of a tissue origin prediction model: Based on the tissue origin prediction sub-models of all tissues obtained in (s3), the test set and validation set data are used for optimization selection to construct a tissue origin prediction model.

[0016] In another preferred example, for the tissue source prediction submodel for determining the source of adrenal tissue, the machine learning model is a random forest model. In another preferred example, for the tissue source prediction submodel for determining the source of bladder tissue, the machine learning model is an XGBoost model. In another preferred example, for the tissue source prediction submodel for determining the source of brain tissue, the machine learning model is a random forest model. In another preferred example, for the tissue source prediction submodel for determining the source of breast tissue, the machine learning model is a random forest model. In another preferred example, for the tissue source prediction submodel for determining the source of cervical tissue, the machine learning model is a LightGBM model. In another preferred example, for the tissue source prediction submodel for determining the source of intestinal tissue, the machine learning model is a CatBoost model. In another preferred example, for the tissue source prediction submodel for determining the source of B cells, the machine learning model is a LightGBM model. In another preferred example, for the tissue source prediction submodel for determining the source of esophageal tissue, the machine learning model is a random forest model. In another preferred example, for the tissue source prediction submodel for determining the source of head and neck tissue, the machine learning model is an XGBoost model. In another preferred example, for the tissue source prediction submodel for determining the source of kidney tissue, the machine learning model is a CatBoost model. In another preferred example, for the tissue source prediction submodel for determining the source of bone marrow tissue, the machine learning model is a CatBoost model. In another preferred example, for the tissue source prediction submodel for determining the source of liver tissue, the machine learning model is a CatBoost model. In another preferred example, for the tissue source prediction submodel for determining the source of lung tissue, the machine learning model is a random forest model. In another preferred example, for the tissue source prediction submodel for determining the source of ovarian tissue, the machine learning model is a CatBoost model. In another preferred example, for the tissue source prediction submodel for determining the source of pancreatic tissue, the machine learning model is a LightGBM model. In another preferred example, for the tissue source prediction submodel for determining the source of ganglion tissue, the machine learning model is a CatBoost model. In another preferred example, for the tissue source prediction submodel for determining the source of prostate tissue, the machine learning model is a CatBoost model. In another preferred example, for the tissue source prediction submodel for determining the source of bone and / or muscle tissue, the machine learning model is a random forest model. In another preferred example, for the tissue source prediction submodel for determining the source of skin tissue, the machine learning model is a LightGBM model. In another preferred example, for the tissue source prediction submodel for determining the source of gastric tissue, the machine learning model is a random forest model.In another preferred example, for the tissue source prediction submodel for determining the source of testicular tissue, the machine learning model is a CatBoost model. In another preferred example, for the tissue source prediction submodel for determining the source of thyroid tissue, the machine learning model is a random forest model. In another preferred example, for the tissue source prediction submodel for determining the source of thymus tissue, the machine learning model is an XGBoost model. In another preferred example, for the tissue source prediction submodel for determining the source of uterine tissue, the machine learning model is a random forest model. In another preferred example, for the tissue source prediction submodel for determining the source of eye tissue, the machine learning model is a random forest model.

[0017] In another preferred embodiment, the expression information of the transcript includes the expression level of the transcript measured in the form of TPM (Transcript per million).

[0018] In another preferred embodiment, in step (s3), the following sub-steps are included:

[0019] (s3a) constructing different tissue origin prediction sub-models for different tissues to determine whether the sample to be tested originates from the tissue; for any tissue, inputting the expression information of tissue-specific RNA transcripts of tissue-positive samples and tissue-negative samples of the tissue into the following five machine learning models for training: Random Forest model, Decision Tree model, XGBoost model, LightGBM model and CatBoost model; thereby obtaining the hyperparameters of the five machine learning models corresponding to any one of the tissues; and based on this, obtaining five trained machine learning models (or the first tissue origin prediction sub-model or the first sub-model) for determining the origin of any one of the tissues;

[0020] (s3b) using the test set and validation set data, scoring the five first tissue origin prediction sub-models for any tissue origin judgment using the following scoring indicators: AUC value, accuracy, sensitivity and specificity; for any tissue, using the first tissue origin prediction sub-model with the highest score as the final tissue origin prediction sub-model (or the second tissue origin prediction sub-model or the second sub-model) for judging the tissue;

[0021] (s3c) obtaining tissue source prediction sub-models for all tissues, that is, all second tissue source prediction sub-models for tissue source prediction;

[0022] In another preferred embodiment, in step (s4), the following steps are included:

[0023] (S4a) Based on the tissue origin prediction sub-model of all tissues obtained in (S3), using the validation set data, the tissue origin sub-model of all tissues is scored using the ROC AUC function;

[0024] (s4b) All tissue origin prediction sub-models with AUC ≥ 0.9 were selected to construct a tissue origin prediction model.

[0025] In another preferred embodiment, the tissue samples include tissue samples from m different tissue sources, wherein m is a positive integer ≥10.

[0026] In another preferred embodiment, m≥15; preferably, n is 15-50; more preferably, n is 18-40.

[0027] In another preferred embodiment, the m different tissue sources include: adrenal gland, B cells, bladder, brain, breast, bone & muscle, bone marrow, cervix, intestine, esophagus, eye, ganglion, kidney, liver, lung, head and neck, ovary, pancreas, prostate, skin, stomach, testicle, thymus, thyroid and uterus.

[0028] In another preferred example, the tissue origin prediction submodel is used to determine whether the tissue of the sample to be tested originates from the following parts: adrenal gland, bladder, bone marrow, brain, breast, intestine, eye, ganglion, liver, pancreas, prostate, thyroid, B cells, kidney, head and neck, skin, stomach and uterus.

[0029] In another preferred embodiment, the tissue-specific RNA transcripts include tissue-specific RNA transcripts for the following tissues: adrenal gland, B cells, bladder, brain, breast, bone & muscle, bone marrow, cervix, intestine, esophagus, eye, ganglion, kidney, liver, lung, head and neck, ovary, pancreas, prostate, skin, stomach, testis, thymus, thyroid and uterus.

[0030] In another preferred embodiment, in step (s1), it includes: establishing a full-length transcript reference GTF file by processing third-generation sequencing data and filtering transcript annotations; the GTF file is used to quantify RNAseq transcripts obtained by sample RNA sequencing; and

[0031] Compare RNAseq transcript expression information of different tissue samples to obtain tissue-specific RNA transcripts of different tissues.

[0032] In another preferred embodiment, in step (s1), the following steps are included:

[0033] (a) Collect the third-generation sequencing data and RNA sequencing data of cancer tissue samples, paracancerous tissue samples, and normal tissue samples from different tissue sources, and generate non-redundant transcriptome GTF files;

[0034] (b) optimizing the non-redundant transcriptome GTF file, wherein, for a certain tissue RNA-seq, if all splicing sites of a transcript are detected in ≥m samples, the transcript is retained, where m is 1-90%, preferably 2-50%;

[0035] (c) Each transcript was compared with the reference transcriptome (GENCODE v.35), and the transcripts were mainly divided into four groups, including full splicing matches, incomplete splicing matches, novel in the catalog, and novel out of the catalog. After filtering, a reference GTF file for full-length RNA transcripts was established:

[0036] Among them, for all incomplete splicing matching transcripts, such transcripts are filtered out from the entire transcriptome;

[0037] Filter out unqualified novel transcripts in the catalog and outside the catalog (for example, cage peaks and polyA sites were annotated using SQANTI3, and novel transcripts in the catalog and outside the catalog were filtered when there were 16 or more adenines within 20 bases downstream of the annotated transcription termination site);

[0038] For similar transcripts, the longest transcript is retained and other transcripts are filtered out;

[0039] (d) quantifying the full-length transcripts of the tissue sample RNA-seq data based on the reference GTF file for the full-length RNA transcripts;

[0040] (e) By comparison, RNA transcripts with significant differences between each tissue sample and other tissue samples are obtained, thereby obtaining tissue-specific RNA transcripts of different tissues.

[0041] In another preferred embodiment, the significant difference is that the gene is expressed in a specific tissue but not expressed or very low expressed in other tissues.

[0042] In another preferred embodiment, the tissue-specific RNA transcript is obtained by the following steps:

[0043] (a) Collect the third-generation sequencing data and RNA sequencing data of cancer tissues, paracancerous tissues, and normal tissue samples from different tissue sources;

[0044] (b) establishing a full-length transcript reference GTF file by processing third-generation sequencing data and filtering transcript annotations; the GTF file is used to quantify RNAseq transcripts obtained by sample RNA-seq sequencing; and

[0045] (c) Using Salmon software, the RNA-seq short read sequences obtained by RNA sequencing are aligned to the transcript library corresponding to the GTF file, and the expression level of the transcripts in each sample is measured in the form of TPM (Transcript per million) to generate transcript expression information of the tissue to be tested; the transcript expression information includes the expression information of tissue-specific RNA transcripts of different tissues.

[0046] In another preferred embodiment, when the tissue to be tested is tumor tissue, the tissue origin prediction model is used to evaluate the primary tissue of the tumor.

[0047] In a second aspect of the present invention, a reference GTF file for a full-length RNA transcript is provided, obtained by the following steps:

[0048] (i) Collect the third-generation sequencing data and RNA-seq sequencing data of cancer tissue samples, paracancerous tissue samples, and normal tissue samples from different tissue sources, and generate non-redundant transcriptome GTF files;

[0049] (ii) optimizing the non-redundant transcriptome GTF file, wherein, for a certain tissue RNA-seq, if all splicing sites of a transcript are detected in ≥m samples, the transcript is retained, where m is 1-90%, preferably 2-50%;

[0050] (iii) Each transcript was compared with the reference transcriptome (GENCODE v.35), and the transcripts were mainly divided into four groups, including full splicing matches, incomplete splicing matches, novel in the catalog, and novel out of the catalog. After filtering, a reference GTF file for full-length RNA transcripts was established:

[0051] Among them, for all incomplete splicing matching transcripts, such transcripts are filtered out from the entire transcriptome;

[0052] Filter out unqualified novel transcripts in the catalog and outside the catalog (for example, cage peaks and polyA sites were annotated using SQANTI3, and novel transcripts in the catalog and outside the catalog were filtered when there were 16 or more adenines within 20 bases downstream of the annotated transcription termination site);

[0053] For similar transcripts, the longest transcript is retained and other transcripts are filtered out.

[0054] In another preferred embodiment, the full splicing match, incomplete splicing match, novel in the catalog and novel out of the catalog are, respectively, transcripts that are completely consistent with known isoforms in the reference transcriptome database compared to the reference transcriptome; transcripts that are completely contained by known transcripts but have fewer exons than known transcripts, and whose intron sequences completely correspond to known transcripts; transcripts that display known splicing sites but present new splicing patterns or exon combinations; and transcripts that have at least one splicing site that has not been annotated in existing databases.

[0055] In a third aspect of the present invention, a method for determining expression information of a full-length RNA transcript is provided, comprising the steps of:

[0056] (i) Based on the reference GTF file for full-length RNA transcripts described in the second aspect of the present invention, full-length transcripts are quantified on the RNA-seq sequencing data of the tissue sample to obtain full-length RNA transcript expression information.

[0057] In another preferred example, the full-length RNA transcript expression information includes the quantitative results of the full-length RNA transcript expression spectrum.

[0058] In a fourth aspect of the present invention, a tissue source prediction system is provided, the system comprising:

[0059] (a) an input module, wherein the input module is configured to input the full-length RNA transcript expression information of the sample to be tested;

[0060] (b) an evaluation module (or prediction module): the evaluation module (or prediction module) is configured to input the full-length RNA transcript expression information into the tissue origin prediction model constructed by the method described in the first aspect of the present invention, thereby obtaining a tissue origin prediction result of the sample to be tested; and

[0061] (d) An output module, wherein the output module is configured to output the prediction result.

[0062] In another preferred embodiment, the system further comprises (a0) a preprocessing module, wherein the preprocessing module is configured to analyze and determine the full-length RNA transcript expression information of the sample to be tested based on the RNA-seq sequencing information of the sample to be tested.

[0063] In another preferred example, the system further includes (a0) a preprocessing module, which is configured to analyze the RNA-seq sequencing information of the sample to be tested using the method described in the third aspect of the present invention to determine the full-length RNA transcript expression information of the sample to be tested.

[0064] In another preferred example, the system further includes (d) a control module, wherein the control module is configured to control the operation of each module.

[0065] In another preferred example, the system further includes (e) a storage module, which is configured to store the following data: hyperparameters, prediction probabilities, and preset reference thresholds of each tissue source prediction sub-model in the tissue source prediction model.

[0066] In another preferred example, the preset reference threshold includes a preset reference threshold of each sub-model in the tissue source prediction model.

[0067] In another preferred embodiment, the determination in the evaluation module includes:

[0068] (i) inputting the full-length RNA transcript expression information into a tissue origin prediction model, and then inputting it into the sub-models of each tissue origin prediction model in turn, thereby obtaining the prediction values ​​of all sub-models; wherein each of the sub-models is used to independently determine whether the sample to be tested is derived from a specific tissue;

[0069] (ii) determining whether the prediction values ​​of all sub-models are greater than or equal to the preset reference threshold C1 of the corresponding tissue source prediction; the determination includes: if the prediction value of the sub-model is greater than or equal to the preset reference threshold C1 of the corresponding tissue source prediction, then determining that the tissue is a candidate tissue of the sample to be tested; thereby obtaining all candidate tissues;

[0070] (iii) All candidate tissues are sorted from large to small according to their corresponding prediction values, thereby obtaining a sorting result of the candidate tissues, and obtaining the candidate tissue with the largest prediction value as the tissue source of the sample to be tested; that is, the tissue of the sample to be tested is derived from the candidate tissue with the largest prediction value.

[0071] In another preferred example, when the sample to be tested is a tumor sample, the tissue origin prediction is the primary tissue origin prediction result of the tumor tissue.

[0072] In another preferred example, the tissue origin prediction model has a total of 18 sub-models, and the 18 sub-models are used to determine whether the tissue of the sample to be tested originates from the following parts: adrenal gland, bladder, bone marrow, brain, breast, intestine, eye, ganglion, liver, pancreas, prostate, thyroid, B cells, kidney, head and neck, skin, stomach and uterus.

[0073] In a fifth aspect of the present invention, a computer storage medium is provided, wherein the storage medium is used to store a computer program corresponding to the algorithm of the tissue origin prediction model constructed by the method described in the first aspect of the present invention.

[0074] In a sixth aspect of the present invention, a method for predicting data organization sources is provided, comprising the steps of:

[0075] (a) Providing data: providing RNA sequencing information of the sample to be tested;

[0076] (b) Data preprocessing: for the sequencing information, analyzing and determining the full-length RNA transcript expression information of the sample to be tested, or using the method described in the third aspect of the present invention to analyze and determine the full-length RNA transcript expression information of the sample to be tested;

[0077] (c) analysis and evaluation: inputting the full-length RNA transcript expression information into the cancer prediction model constructed by the method described in the first aspect of the present invention, thereby obtaining an evaluation result of the sample to be tested; and

[0078] (d) Output results: Output the evaluation results.

[0079] In another preferred embodiment, the evaluation process in the analysis and evaluation is as follows:

[0080] (i) inputting the full-length RNA transcript expression information into a tissue origin prediction model, and then inputting it into each sub-model of the tissue origin prediction model in turn, thereby obtaining the prediction values ​​of all sub-models; wherein each of the sub-models is used to independently determine whether the sample to be tested is derived from a specific tissue;

[0081] (ii) determining whether the prediction values ​​of all sub-models are greater than or equal to the preset reference threshold C1 of the corresponding tissue source prediction; the determination includes: if the prediction value of the sub-model is greater than or equal to the preset reference threshold C1 of the corresponding tissue source prediction, then determining that the tissue is a candidate tissue of the sample to be tested; thereby obtaining all candidate tissues;

[0082] (iii) All candidate tissues are sorted from large to small according to their corresponding prediction values, thereby obtaining a sorting result of the candidate tissues, and obtaining the candidate tissue with the largest prediction value as the tissue source of the sample to be tested; that is, the tissue of the sample to be tested is derived from the candidate tissue with the largest prediction value.

[0083] In another preferred example, the preset reference threshold includes a preset reference threshold of each sub-model in the tissue source prediction model.

[0084] In another preferred embodiment, the method for obtaining the expression information of the full-length RNA transcript of the tissue to be tested according to the RNA sequencing information of the tissue to be tested is:

[0085] The RNA-seq short read sequence of the tissue to be tested is aligned to the transcript library corresponding to the GTF file described in the second aspect of the present invention, and the expression level of each transcript in the sample is measured in the form of TPM (Transcript per million), thereby generating the full-length RNA transcript expression information of the tissue to be tested.

[0086] In another preferred example, when the sample to be tested is a tumor sample, the tissue origin prediction is the primary tissue origin prediction result of the tumor tissue.

[0087] In another preferred embodiment, the method is an in vitro method.

[0088] In another preferred embodiment, the method is non-diagnostic and non-therapeutic.

[0089] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features specifically described below (such as embodiments) can be combined with each other to form a new or preferred technical solution. Due to space limitations, they will not be described one by one here. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] Figure 1A -B shows the ROC curve of the tissue origin sub-model in Example 3 of the present invention when predicting different tissue origins.

[0091] Figure 2 An example schematic diagram of the tissue origin prediction model of the present invention is shown (the result obtained in the figure is that the tissue originates from B cells), and Feat in the figure refers to feature.

[0092] Figure 3 A schematic diagram of a prediction module according to an example of the present invention is shown.

[0093] Figure 4 A schematic diagram of the tissue origin prediction system of the present invention is shown, where Feat refers to feature.

[0094] Figure 5 A schematic diagram of a process for predicting tissue origin of a population to be tested according to an example of the present invention is shown. DETAILED DESCRIPTION

[0095] The inventors have obtained tissue-specific RNA transcripts of 25 tissues based on the sequencing information of normal different tissues through extensive and in-depth research, and developed a model (tissue source prediction model) for predicting the tissue source of the sample to be tested (including prediction of the primary site of the cancer sample) for the first time. Specifically, the model of the present invention can be used for diagnosing or predicting 18 tissue sources in total, namely: adrenal gland, bladder, bone marrow, brain, breast, intestine, eye, ganglion, liver, pancreas, prostate, thyroid, B cell, kidney, head and neck, skin, stomach and uterus. And, when the sample to be tested is cancer tissue, the predicted tissue source is the primary site or tissue of the cancer. The tissue source prediction model of the present invention has a higher prediction accuracy. It can be used for early clinical auxiliary diagnosis. On this basis, the present invention has been completed.

[0096] the term

[0097] In order to more easily understand the present disclosure, some terms are first defined. As used in this application, unless otherwise expressly provided herein, each of the following terms should have the meaning given below. Other definitions are set forth throughout the application.

[0098] The term "about" can refer to a value or composition that is within an acceptable error range for a particular value or composition as determined by one of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined. For example, as used herein, the expression "about 100" includes all values ​​between 99 and 101 (e.g., 99.1, 99.2, 99.3, 99.4, etc.).

[0099] As used herein, the term "comprising" or "including (comprising)" may be open, semi-closed and closed. In other words, the term also includes "consisting essentially of" or "consisting of".

[0100] As used herein, unless otherwise indicated, any concentration range, percentage range, ratio range or integer range should be understood to include any integer value within the range and fractional values ​​thereof (e.g., one tenth and one hundredth of an integer) where appropriate.

[0101] As used herein, the term "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0102] As used herein, "transcript expression information", "transcript quantitative results", "transcript expression level" and "transcript expression value" have the same meaning and can be used interchangeably, all referring to the alignment of the RNA-seq short read sequence of the tissue sample to the transcript library corresponding to the GTF file constructed by the present invention, and the expression level of each transcript in the sample measured in the form of TPM (Transcript per million).

[0103] As used herein, "full-length RNA transcript" and "full-length transcript" have the same meaning and can be used interchangeably, and both refer to the complete RNA molecule produced from a gene, including all parts of the coding region and non-coding region. It contains a series of information such as the 5' untranslated region (5'UTR), exon region, intron region (which is cut out in mature mRNA but retained in the full-length transcript) and 3' untranslated region (3'UTR).

[0104] Tissue-specific RNA transcripts

[0105] The present invention calculates normal tissue-specific RNA transcripts (Tissue-Specific RNA transcript, Tissue-SRT), thereby using it to specifically identify the tissue origin of tumor tissue or normal tissue.

[0106] Normal tissue-specific RNA transcripts are defined as specific RNA transcripts that are highly expressed only in a single tissue and not expressed or very low expressed in other normal tissues. These RNA transcripts are only expressed in a specific tissue.

[0107] The data sources for the calculation of tissue-specific RNA transcripts include: next-generation sequencing RNAseq data of 28 normal tissue samples in the public database GTEx and GEO datasets, including adrenal gland, B cell, bladder, brain, breast, bone & muscle, bone marrow, cervix, intestine, esophagus, eye, ganglion, kidney, liver, lung, head and neck, ovary, pancreas, prostate, skin, stomach, testis, thymus, thyroid, uterus, heart, spleen and small intestine.

[0108] In order to obtain tissue-specific RNA transcripts, the expression profiles of multiple tissue types were used to calculate the specificity score of each transcript. The Shannon entropy was used to measure tissue specificity, and its formula is:

[0109]

[0110] Among them, H t represents the Shannon entropy of transcript t, N is the total number of tissue types, and p it represents the expression ratio of transcript t in tissue type i. The expression ratio of each transcript in all tissue types is calculated as follows:

[0111]

[0112] where x it Represents the expression value of transcript t in tissue type i, which is the median value of transcript t in each type i. tThe maximum value of log2(N) is t =log2(N), it means that the transcript t is expressed exactly the same in all tissue types. t The minimum value of is 0, indicating that transcript t is specifically expressed only in a certain tissue type. Therefore, the Shannon entropy is converted into a specificity score, and its calculation formula is:

[0113] S t =log2(N)-H t

[0114] After transformation, S t The maximum value is log2(N), and the minimum value is 0. Finally, the maximum expression ratio p of the transcript in tissue / cancer / cell type is calculated. mt The second largest expression ratio p nt More than twice and specificity score S t When it is greater than 1, it is defined as a transcript specifically expressed in a certain tissue type.

[0115] Through the analytical method, 25 normal tissue-specific RNA transcripts were identified (3 normal tissues had no associated tumors, and the small intestine, heart and spleen tissue-specific RNA transcripts were not calculated), including adrenal, B cell, bladder, brain, breast, bone & muscle, bone marrow, cervix, intestine, esophagus, eye, ganglion, kidney, liver, lung, head and neck, ovary, pancreas, prostate, skin, stomach, testis, thymus, thyroid and uterus tissue type-specific RNA transcripts (a list of 25 tissue-specific RNA transcripts expressed only in various normal tissues, which is unique to the present invention).

[0116] Tissue sample full-length RNA transcript quantification method of the present invention

[0117] The method for quantifying full-length RNA transcripts in tissue samples of the present invention comprises the following steps:

[0118] (St1): 177 tumor tissue specimens from patients with hepatocellular carcinoma, colorectal cancer, ovarian cancer, breast cancer, nasopharyngeal carcinoma, neuroendocrine tumors, gastric cancer, gastrointestinal stromal tumors, renal cancer and cervical cancer were collected. RNA was extracted from the collected tissues and then subjected to PacBio Iso-seq third-generation sequencing. The original PacBio sequencing data was processed using the Iso-Seq workflow.

[0119] (St2): Collect third-generation PacBio and ONT transcriptome sequencing data of 21 cancer tissue samples (including esophageal cancer, lung cancer, glioblastoma, leukemia, lymphoma, melanoma, myeloma and sarcoma) and 153 normal tissue samples (including adipose tissue, adrenal gland, brain, breast, heart, liver, lung, muscle, ovary, pancreas, testis, fibroblasts, GM12878, H1, H9, HEK293T and WTCll cells) from public databases.

[0120] (St3): FASTQ sequences from PacBio and ONT data were aligned using minimap2 and processed using the standard pipeline to obtain a merged GTF transcriptome. Full-length transcripts from all samples were merged into a non-redundant transcriptome GTF file using gffcompare software.

[0121] (St4): In the merged GTF file, if all splice sites of a transcript are detected in at least 5 samples in TCGA and GTEx cancer / tissue RNA-seq, the transcript is retained.

[0122] (St5): These transcripts were compared with the reference transcriptome (GENCODEv.35) using SQANTI3 and gffcompare software, and the transcripts were mainly divided into four groups, including full splicing matches, incomplete splicing matches, novel in the catalog, and novel outside the catalog. All incomplete splicing match transcripts may be the result of RNA degradation and incomplete reverse transcription, and such transcripts were filtered out of the entire transcriptome.

[0123] (St6): Cage peaks and polyA sites were annotated using SQANTI3, and novel transcripts both in- and outside the catalog were filtered when there were 16 or more adenines within 20 bases downstream of the annotated transcription stop site.

[0124] (St7): The full-length sequence of the transcript was obtained using gffread software according to the transcript position provided by GTF. ORFs in the transcripts of the third-generation sequencing data were predicted using GeneMark. When the stop codon was located more than 50 nucleotides upstream of the last exon-exon splice site, the transcript was predicted to induce nonsense-mediated decay.

[0125] (St8): The whole genome of human hg38 transposable elements was downloaded from the UCSC genome browser, and the BEDTools software was used to determine whether the transcripts overlapped with TE within TSS ± ​​100bp. Similar transcripts were merged according to the following criteria: i) up to three exon differences were allowed. ii) each exon did not differ by more than four bases. iii) the upper limit of the dissimilarity allowed for transcription start and stop sites was 100. iv) for coding transcripts, it was necessary to ensure that the ORFs between transcripts did not differ.

[0126] (St9): For similar transcripts, the longest one was retained and other transcripts were filtered out. A reference GTF file containing 1,069,895 full-length RNA transcripts was established for full-length transcript quantification of tissue sample RNAseq data.

[0127] (St10): During data processing, Salmon software can be used to create an index for the transcript GTF file containing 1,069,895 full-length RNA transcripts obtained above.

[0128] (St11): The RNA-seq short read sequences were aligned to the transcript library corresponding to the GTF file, and the expression level of the transcripts in each sample was measured in the form of TPM (Transcript per million). A quantitative method was developed for quantifying the expression spectrum of the full-length RNA transcripts of samples through RNAseq data.

[0129] Tissue source prediction model based on machine learning

[0130] The tissue origin prediction model based on machine learning of the present invention relies on the specific biological characteristics of different normal tissues (tissue-specific transcripts, Tissue-SRT), uses machine learning to set positive and negative thresholds, and constructs a prediction model. According to the expression characteristics of each tissue-specific transcript, a tissue origin prediction model is established to determine the tissue origin of the cancer. The tissue origin prediction system constructed based on the tissue origin prediction model predicts based on the specific transcript (SRT) data expressed by the tissue of the object to be tested, diagnoses or predicts what kind of tissue the tissue is from; when the sample or tissue to be tested is a tumor tissue, the tissue origin prediction model of the present invention can be used to predict the primary site of the sample to be tested.

[0131] As a branch of artificial intelligence, machine learning allows computer systems to learn from data and automatically improve by using statistical and computer science methods, and to make predictions and decisions by analyzing and identifying patterns and rules in data. In the field of oncology, it can analyze data such as patients' clinical symptoms, medical images, laboratory tests, and histopathology to discover new patterns and correlations between variables and generate predictive models. By applying machine learning technology, the field of oncology diagnosis can extract valuable information from large amounts of complex data, helping doctors provide more accurate and rapid diagnostic results, and improving the treatment effects and survival rates of cancer patients.

[0132] In the model construction process of the present invention, a machine learning model is used to construct the final model of the present invention. The machine learning model includes: a random forest model, a decision tree model, an XGBoost model, a LightGBM model and a CatBoost model.

[0133] Tissue Origin Prediction of the Present Invention

[0134] The present invention provides multiple applications of the tissue origin prediction model of the present invention, including the development of software, equipment or system for auxiliary diagnosis of tissue origin. Figure 5 A schematic diagram of a process for predicting the tissue origin of a population to be tested based on the tissue origin auxiliary diagnosis model of the present invention is shown.

[0135] like Figure 5 As shown, first obtain the full-length transcript expression information of the object to be tested, the steps include: after obtaining the biopsy tissue sample from the population to be tested, the full-length transcript reference GTF file established by processing the third-generation sequencing data and filtering the transcript annotation is used to quantify the RNAseq transcript of the tissue sample; the RNA-seq short read sequence is aligned to the transcript library obtained by long read sequencing, and the expression amount of the transcript in each sample is measured in the form of TPM (Transcript per million), and the expression information of the full-length RNA transcript of the tissue to be tested is generated. The transcript expression information includes the tissue-specific RNA transcript of the tissue to be tested.

[0136] After the expression information of the full-length RNA transcript is input into the tissue source prediction model, a sub-model corresponding to the tissue source is determined from the tissue source prediction model; the expression information of the full-length RNA transcript is input into a sub-model corresponding to the tissue source to obtain the output of the sub-model, and the output of the sub-model is the predicted value that the tissue to be tested belongs to the tissue source corresponding to the sub-model; the sub-model corresponding to the next tissue source is determined, and the expression information of the full-length RNA transcript is input into the sub-model corresponding to the next tissue source to obtain the output of the next sub-model, until the outputs of all sub-models are determined. If it is determined that the tissue to be tested includes cancer tissue, the tissue source of the tissue to be tested is determined according to the output of the tissue source prediction model.

[0137] In another preferred example, when predicting the tissue source, each time a judgment is made as to whether it is one of the tissues, for example, when judging whether the tissue source is the adrenal gland, a sub-model whose tissue source is the adrenal gland is obtained, and a set of adrenal gland feature names is determined according to the sub-model, and corresponding features are obtained from tissue-specific transcripts, and input into the classifier of the corresponding sub-model. When determining the corresponding classifier, the corresponding classifier can be obtained according to the classifier name. The classifier is preset with a classifier name and a classifier parameter, and the classifier name and the classifier parameter can be determined according to the corresponding sub-model in the preset model of the tissue source. The classifier parameters include a predetermined hyperparameter to achieve the determination of whether the tissue source is adrenal tissue. The output result of the sub-model is the predicted value of the tissue source of the tissue to be tested corresponding to the sub-model; if the predicted value of the tissue source corresponding to the sub-model of the tissue to be tested is greater than the preset tissue source prediction threshold, it is determined that the tissue source of the tissue to be tested includes the tissue source; if the predicted value of the tissue source corresponding to the sub-model of the tissue to be tested is less than the preset tissue source prediction threshold, it is determined that the tissue source of the tissue to be tested does not include the tissue source.

[0138] Continue to judge whether the source is from other tissues in the above manner to obtain the output of all sub-models. Finally, the output of the tissue source prediction model is generated according to the tissue source corresponding to one or more sub-models whose output prediction values ​​are greater than the tissue source prediction threshold. If the output result of one of the sub-models shows that the tissue source of the tissue to be tested includes the adrenal gland, and the output of another sub-model shows that the tissue source includes B cells, then the output of the tissue source prediction model includes the adrenal gland and B cells.

[0139] In another preferred embodiment, the tissue origin prediction device, software or system of the present invention includes modules selected from the following groups: an input module, a preprocessing module, a prediction module (or an evaluation module), and an output module.

[0140] In another preferred embodiment, the tissue origin prediction system of the present invention makes predictions based on the full-length RNA transcript expression information of the patient's tissue (including tissue-specific transcript (SRT) data) for diagnosis or prediction of the tissue origin of the sample to be tested, and the prediction indicators are objective.

[0141] The technical solution of the present invention predicts the probability of the current sample tissue being of a certain tissue origin (including the probability of the primary site of the cancer tissue being of a certain tissue origin) based on the gene information of the sample tissue based on a pre-trained tissue origin prediction model. The gene feature information includes ti-SRT (tissue-specific transcript). All these features are obtained by analyzing tissue samples using long-read and short-read RNA sequencing technologies, and integrating RNA sequencing third-generation sequencing data from a variety of normal and cancerous human tissues and cells. Tissue-specific RNA transcripts are analyzed and obtained from a large amount of case data. Five models, namely, the random forest model, the decision tree model, the XGBoost model, the LightGBM model and the CatBoost model, use grid search according to each feature to obtain the hyperparameters of each model, compare the roc_auc of the five models under their hyperparameters, select the model with the highest roc_auc and its hyperparameters, and obtain five prediction models for all types of tumors that have been trained, thereby realizing the probability of the sample being diagnosed as cancer and the probability of predicting the source of cancer as outcome indicators, which is very objective. The tissue origin prediction system can assist in the diagnosis of the primary and / or metastatic status of the tumor, and can also improve the diagnostic accuracy of tumors of unknown primary site.

[0142] In addition, the application of this system is independent of pathologists, requires less human intervention, can provide standardization and consistency, and offer pathologists highly stable auxiliary tumor diagnosis, thereby reducing the burden on physicians and improving work efficiency. It has high clinical transformation value and market prospects.

[0143] It should be noted that in the device or apparatus of the present application, the storage medium (computer-readable medium) used to store the model or algorithm of the present invention may be a computer-readable signal medium or a non-transitory computer-readable storage medium or any combination of the above two. The non-transitory computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of non-transitory computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0144] In the present application, a non-transitory computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a non-transitory computer-readable storage medium, which may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0145] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0146] The main advantages of the present invention include:

[0147] (a) The tissue origin prediction system of the present invention can assist in the diagnosis of the primary and / or metastatic status of a tumor, and can also improve the diagnostic accuracy of tumors with unknown primary sites and multiple primary tumors, helping doctors provide more accurate and rapid diagnostic results, improving diagnostic efficiency, and having good clinical application value and market prospects.

[0148] The present invention will be further described below in conjunction with specific examples. It should be understood that these examples are intended to illustrate the present invention only and are not intended to limit the scope of the present invention. The experimental methods in the following examples where specific conditions are not specified are usually performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or under conditions recommended by the manufacturer. Unless otherwise indicated, percentages and parts are weight percentages and weight parts.

[0149] Example 1 Construction of GTF file

[0150] 1.1 Data Acquisition

[0151] 177 tumor tissue specimens from patients with hepatocellular carcinoma, colorectal cancer, ovarian cancer, breast cancer, nasopharyngeal carcinoma, neuroendocrine tumors, gastric cancer, gastrointestinal stromal tumors, renal cancer and cervical cancer were collected. RNA was extracted from the collected tissues and PacBio Iso-seq third-generation sequencing and second-generation sequencing were performed (this part of the data was generated by the self-tissue sequencing of the present invention). The original PacBio sequencing data was processed using the Iso-Seq workflow.

[0152] The third-generation PacBio and ONT transcriptome sequencing data and second-generation sequencing data of 21 cancer tissue samples (including esophageal cancer, lung cancer, glioblastoma, leukemia, lymphoma, melanoma, myeloma and sarcoma) and 153 normal tissue samples (including adipose tissue, adrenal gland, brain, breast, heart, liver, lung, muscle, ovary, pancreas, testis, fibroblasts, GM12878, H1, H9, HEK293T and WTC11 cells) from public databases were collected.

[0153] 1.2 Construction of GTF file

[0154] FASTQ sequences from PacBio and ONT data were aligned using minimap2 and processed using a standard pipeline to obtain a merged GTF transcriptome. Full-length transcripts from all samples were merged into a non-redundant transcriptome GTF file using gffcompare software.

[0155] The non-redundant transcriptome GTF files were optimized using a variety of methods. In the merged GTF files, if all splicing sites of a transcript were detected in at least 5 samples in the TCGA and GTEx cancer / tissue RNA-seq, the transcript was retained.

[0156] These transcripts were compared with the reference transcriptome (GENCODE v.35) using SQANTI3 and gffcompare software, and the transcripts were mainly divided into four groups, including full splicing matches, incomplete splicing matches, novel in the catalog, and novel out of the catalog.

[0157] All incompletely spliced ​​transcripts, which may be the result of RNA degradation and incomplete reverse transcription, were filtered out from the entire transcriptome.

[0158] Cage peaks and polyA sites were annotated using SQANTI3, and both novel in- and out-of-catalog transcripts were filtered when there were 16 or more adenines within 20 bases downstream of the annotated transcription stop site.

[0159] The full-length sequence of the transcript was obtained using gffread software according to the transcript position provided by GTF. ORFs in the transcripts of the third-generation sequencing data were predicted using GeneMark. When the stop codon is located more than 50 nucleotides upstream of the last exon-exon splice site, the transcript is predicted to induce nonsense-mediated decay.

[0160] The whole genome transposable elements of human hg38 were downloaded from the UCSC genome browser, and the BEDTools software was used to determine whether the transcripts overlapped with TE within TSS ± ​​100 bp. Similar transcripts were merged according to the following criteria: i) up to three exon differences were allowed. ii) each exon did not differ by more than four bases. iii) the upper limit of the dissimilarity allowed for the transcription start and end sites was 100. iv) for coding transcripts, it was necessary to ensure that the ORFs between transcripts did not differ.

[0161] For similar transcripts, the longest one was retained and other transcripts were filtered out. A reference GTF file containing 1,069,895 full-length RNA transcripts was established for full-length transcript quantification of tissue sample RNAseq data (this newly generated GTF file was developed by the inventors based on their own sequencing data and public data analysis).

[0162] During data processing, the Salmon software can be used to create an index for the GTF file containing 1,069,895 full-length RNA transcripts obtained above (creating an index based on the newly generated GTF file is unique to the present invention).

[0163] The RNA-seq short-read sequences of the above samples (data sources include public databases and the second-generation sequencing data obtained by the above self-test) were aligned to the transcript library corresponding to the GTF file, and the expression level of the transcripts in each sample was measured in the form of TPM (Transcript per million). A quantitative method was developed for quantifying the expression spectrum of the full-length RNA transcripts of samples through RNAseq data.

[0164] Transcript quantification was performed on RNAseq data of 33 TCGA tumors and 28 normal tissues in public databases (TCGA, GTEx, and GEO), and the FLIBase database based on full-length RNA transcripts was developed, creating a comprehensive resource library of full-length transcript information of human cancer tissues and normal tissues.

[0165] The database covers a variety of cancer types and different normal tissue types, identifies a large number of unannotated tumor-specific RNA transcripts, and provides panoramic information on tissue-specific RNA transcripts in different normal tissues. FLIBase has significant advantages in discovering a large number of previously unannotated isomers and tumor-specific RNA transcripts. Using this database to conduct in-depth mining and research on the cancer transcriptome at the transcript level is expected to gain new insights into the precise diagnosis and treatment of cancer.

[0166] Example 2 Calculation of tissue-specific RNA transcripts

[0167] Transcript quantification and analysis were performed on RNA-seq data of 28 normal tissues in public databases (GTEx and GEO) to determine the tissue-specific RNA transcripts of each tissue sample.

[0168] Among them, the number of each normal tissue sample in the public database GTEx and GEO dataset is shown in Table 1.

[0169] Table 1

[0170]

[0171]

[0172] In order to obtain tissue-specific RNA transcripts, the expression profiles of multiple tissue types were used to calculate the specificity score of each transcript. The Shannon entropy was used to measure tissue specificity, and its formula is:

[0173]

[0174] Among them, H t represents the Shannon entropy of transcript t, N is the total number of tissue types, and p itrepresents the expression ratio of transcript t in tissue type i. The expression ratio of each transcript in all tissue types is calculated as follows:

[0175]

[0176] where x it Represents the expression value of transcript t in tissue type i, which is the median value of transcript t in each type i. t The maximum value of log2(N) is t =log2(N), it means that the transcript t is expressed exactly the same in all tissue types. t The minimum value of is 0, indicating that transcript t is specifically expressed only in a certain tissue type. Therefore, the Shannon entropy is converted into a specificity score, and its calculation formula is:

[0177] S t -log2(N)-H t

[0178] After transformation, S t The maximum value is log2(N), and the minimum value is 0. Finally, the maximum expression ratio p of the transcript in tissue / cancer / cell type is calculated. mt The second largest expression ratio p nt More than twice and specificity score S t When it is greater than 1, it is defined as a transcript specifically expressed in a certain tissue type.

[0179] Through the analysis method, specific RNA transcripts of 25 normal tissues (including 3 normal tissues without related tumors, small intestine, heart and spleen tissue-specific RNA transcripts were not calculated) were identified, including adrenal, B cell, bladder, brain, breast, bone & muscle, bone marrow, cervix, intestine, esophagus, eye, ganglion, kidney, liver, lung, head and neck, ovary, pancreas, prostate, skin, stomach, testis, thymus, thyroid and uterus tissue type-specific RNA transcripts (Tissue Specific RNA transcript, Tissu e -SRT) (a list of tissue-specific RNA transcripts of 25 tissues that are only expressed in various normal tissues, calculated by the inventors themselves, and developed for the first time and unique to the present invention).

[0180] Example 3 Construction of tissue source prediction model

[0181] The tissue origin prediction model of the present invention is used to distinguish which tissue a given tissue sample originates from, or which tissue type of cancer a cancer patient suffers from (including the type of carcinoma in situ that has metastasized).

[0182] 3.1 Sample Data

[0183] Sample data source: The sample data includes RNA-seq data of 33 cancer tissues in the TCGA dataset in the public database, as shown in Table 2.

[0184] Table 2 Number of cancer tissue samples of 33 cancers in the TCGA dataset in public databases

[0185]

[0186]

[0187] 3.2 Model construction

[0188] 3.2.1 Construction of sub-model

[0189] In the process of model construction, the cancer samples corresponding to each specific tissue source in the TCGA dataset were set as positive samples in the training set, and the cancer samples from other tissue sources were set as negative samples in the training set. In the process of this model construction, first, tissue source prediction sub-models were constructed for 25 different tissues, that is, for any tissue, a corresponding sub-model was constructed to determine whether the sample to be tested originated from the tissue. Then, the tissue source prediction model was constructed by comprehensively selecting the optimal sub-models.

[0190] Therefore, when constructing the model, it includes the construction of a sub-model for predicting each tissue. That is, this construction constructs sub-models for predicting the following 25 tissue sources: adrenal gland, B cell, bladder, brain, breast, bone & muscle, bone marrow, cervix, intestine, esophagus, eye, ganglion, kidney, liver, lung, head and neck, ovary, pancreas, prostate, skin, stomach, testis, thymus, thyroid and uterus.

[0191] According to the organ type, 32 TCGA cancers (mesothelioma (MESO) lacks corresponding normal tissue RNAseq data of peritoneum and pleura, so it is not included) are sorted into 25 organ-derived cancers (here only classified according to the tissue source of the organ, a total of 25 operations are performed, and each operation distinguishes one tissue source), as shown in Table 3. Among them, the feature number in Table 3 is the number of tissue-specific RNA transcripts of each tissue calculated in Example 2.

[0192] Table 3 List of tissue sources and corresponding cancer types

[0193]

[0194]

[0195] All cancer tissue samples in the TCGA dataset were classified according to the list of tissue origin and cancer type, namely: adrenal gland, B cell, bladder, brain, breast, bone & muscle, bone marrow, cervix, intestine, esophagus, eye, ganglion, kidney, liver, lung, head and neck, ovary, pancreas, prostate, skin, stomach, testis, thymus, thyroid uterus, and MESO, a total of 26 types.

[0196] The classification method shown in Table 3 is as follows: Based on 25 normal tissue-specific RNA transcripts, a binary classification method is used, and one feature is used for classification each time, for a total of 25 operations. (After 25 trainings, all features corresponding to the feature numbers of different tissues are the specific RNA transcript numbers of different tissues)

[0197] In the first test, 1947 adrenal gland features were used to distinguish adrenal cancer from the other 25 cancers.

[0198] Second, 1,292 B-cell signatures were used to distinguish lymphoma from 25 other cancers, with positive B-cell signatures for DLBC and negative B-cell signatures for 25 other cancers.

[0199] The third time, 573 bladder features were used to distinguish bladder cancer from the other 25 cancers, with positive results for BLCA and negative results for the other 25 cancers.

[0200] In the fourth round, 5466 brain features were used, with positive results for GBM and LGG cancer tissues and negative results for the other 25 cancer tissues, to distinguish brain tumors from the other 25 cancers;

[0201] The fifth time, 786 breast features were used to distinguish breast cancer from 25 other cancers, with positive results for BRCA and negative results for 25 other cancers.

[0202] The sixth time, using 17,668 bone & muscle features, positive for SARC cancer tissues, negative for the other 25 cancer tissues, distinguished sarcomas from the other 25 cancers;

[0203] In the seventh round, 24,759 bone marrow features were used to distinguish leukemia from 25 other cancers, with positive results for LAML and negative results for the other 25 cancers.

[0204] In the eighth time, cervical characteristics were used to distinguish cervical cancer from the other 25 cancers, with positive results for CESC cancer tissue and negative results for the other 25 cancer tissues;

[0205] In the 9th time, 209 intestinal features were used, with positive results for COAD and READ cancer tissues and negative results for the other 25 cancer tissues, to distinguish intestinal cancer from the other 25 cancers;

[0206] In the 10th time, 103 esophageal features were used to distinguish esophageal cancer from the other 25 cancers, with positive results for ESCA and negative results for the other 25 cancers.

[0207] In the 11th round, 13,535 eye features were used, with positive results for UVM cancer tissue and negative results for the other 25 cancer tissues, to distinguish uveal melanoma from the other 25 cancers;

[0208] In the 12th round, 69,670 ganglion features were used to distinguish paraganglioma from 25 other cancers, with positive PCPG cancer tissues and negative other 25 cancer tissues.

[0209] In the 13th round, 1482 kidney features were used to distinguish renal cancer from the other 25 cancers, with positive results for KICH, KIRC, and KIRP cancer tissues and negative results for the other 25 cancer tissues;

[0210] In the 14th round, 21,988 liver features were used to distinguish liver cancer from 25 other cancers, with positive results for LIHC and CHOL cancer tissues and negative results for the other 25 cancer tissues;

[0211] In the 15th round, 1449 lung features were used to distinguish lung cancer from 25 other cancers, with positive results for LUAD and LUSC cancer tissues and negative results for the other 25 cancer tissues.

[0212] In the 16th round, 1357 head and neck features were used to distinguish head and neck squamous cell carcinoma from 25 other cancers, with positive results for HNSC and negative results for the other 25 cancers.

[0213] In the 17th round, 2615 ovarian features were used to distinguish ovarian cancer from the other 25 cancers, with positive results for OV cancer tissues and negative results for the other 25 cancer tissues.

[0214] In the 18th round, 1,135 pancreatic features were used to distinguish pancreatic cancer from 25 other cancers, with positive results for PAAD and negative results for the other 25 cancers.

[0215] In the 19th round, 2146 prostate features were used to distinguish prostate cancer from 25 other cancers, with positive results for PRAD and negative results for the other 25 types of cancer tissues.

[0216] In the 20th round, 2273 skin features were used to distinguish skin melanoma from 25 other cancers, with positive results for SKCM and negative results for the other 25 cancers.

[0217] In the 21st round, 384 gastric features were used to distinguish gastric cancer from the other 25 cancers, with positive results for STAD cancer tissues and negative results for the other 25 cancer tissues;

[0218] In the 22nd test, 53,097 testicular features were used to distinguish testicular cancer from 25 other cancers.

[0219] In the 23rd round, 1,685 thymic features were used to distinguish thymic carcinoma from 25 other cancers, with positive results for THYM cancer tissues and negative results for the other 25 cancer tissues.

[0220] In the 24th round, 2821 thyroid features were used to distinguish thyroid cancer from the other 25 cancers, with positive results for THCA in cancer tissues and negative results for the other 25 cancer tissues.

[0221] In the 25th time, 989 uterine features were used, with positive results for UCEC and UCS cancer tissues and negative results for the other 25 cancer tissues, to distinguish uterine tumors from the other 25 cancers.

[0222] Taking the tissue origin prediction sub-model (first sub-model) for determining whether the sample to be tested is derived from the adrenal gland as an example, the construction process is as follows:

[0223] (a) Dataset distribution: The RNAseq data of each source tissue in the TCGA dataset were divided into training set and test set in a 3:1 ratio.

[0224] (b) Sample classification: The samples in the TCGA dataset are classified into positive samples with adrenal cancer (such as adrenocortical carcinoma) as the source, and negative samples with non-adrenal cancer (i.e. cancer from other tissues).

[0225] (c) obtaining the expression information (or quantitative results) of the full-length RNA transcripts of the positive and negative samples according to the tissue sample full-length RNA transcript quantification method; specifically, aligning the RNA-seq short read sequence of the sample to the transcript library corresponding to the GTF file constructed in Example 1, and measuring the expression amount of the transcript in each sample in the form of TPM (Transcript per million), thereby obtaining the expression information (or quantitative results) of the full-length RNA transcripts of the tissue samples;

[0226] (d) using the expression information of the full-length RNA transcripts of the tissue sample to perform a grid search on the machine learning model to be trained to determine the hyperparameters of the model, and when performing the grid search, each parameter combination is cross-validated, such as using a 5-fold cross validation and using roc auc as a scoring function; and finally determining the optimal hyperparameters;

[0227] The machine learning models include: Random Forest model, Decision Tree model, XGBoost model, LightGBM model and CatBoost model; when determining the hyperparameters, the expression information of the full-length RNA transcripts of the tissue samples are respectively input into the above five machine learning models to obtain the optimal hyperparameters corresponding to each machine learning model; and based on this, five tissue source prediction sub-models (or first sub-models) for determining whether the sample to be tested is derived from the adrenal gland are obtained;

[0228] (e) Using the test set data, the five tissue origin prediction sub-models trained above are verified and evaluated, including by scoring using the ROC AUC function.

[0229] (f) Using the validation set data to validate and evaluate the five tissue origin prediction sub-models trained above, such as by scoring using the ROC AUC function, and finally determining a model with the best performance from the five sub-models as the tissue origin prediction sub-model (second sub-model) for ultimately determining whether the sample to be tested is derived from the adrenal gland. The validation set is RANseq data of cancer tissues of 33 cancers obtained from the public database GEO and the dbGAP independent data set, as shown in Table 4.

[0230] Table 4 Cancer tissue sample data of 33 cancers in the public database GEO and dbGAP dataset

[0231]

[0232]

[0233]

[0234] The other 24 sub-models for predicting different tissue origins all adopt the same construction process as the tissue origin prediction sub-model used to determine whether the sample to be tested is derived from the adrenal gland, and finally obtain the tissue origin prediction sub-models for determining whether the sample to be tested is derived from any other tissue.

[0235] Finally, a total of 25 tissue origin prediction sub-models were obtained, which were used to determine whether the sample to be tested originated from the following 25 tissues: adrenal gland, B cell, bladder, brain, breast, bone & muscle, bone marrow, cervix, intestine, esophagus, eye, ganglion, kidney, liver, lung, head and neck, ovary, pancreas, prostate, skin, stomach, testis, thymus, thyroid and uterus.

[0236] Among them, the five machine learning models corresponding to each tissue origin prediction, that is, the five tissue origin prediction sub-models (or first sub-models) corresponding to each tissue prediction, have AUC values ​​on the training set, test set and validation set, respectively, as shown in Table 2 below.

[0237] Table 5

[0238]

[0239]

[0240]

[0241]

[0242]

[0243]

[0244]

[0245]

[0246] As shown in Table 5, in both the training set and the test set, the five first sub-models performed well in all evaluation indicators. The accuracy of the model can be further verified by the validation set, and other data sets can be used to verify whether the model has the same predictive ability in an independent data set.

[0247] The results showed that in the validation set, among the five first models, each tissue had its own best performing algorithm model in predicting various evaluation indicators of 25 tissue origins (covering 32 types of cancer); the best algorithm model would serve as the final tissue origin prediction sub-model (second sub-model) for the tissue.

[0248] Among them, the AUC of the best algorithm of each prediction sub-model for 18 cancer origins (covering 24 types of cancer) was greater than 0.9, including adrenal gland, bladder, bone marrow, brain, breast, intestine, eye, ganglion, liver, pancreas, prostate, thyroid, B cell, kidney, head and neck, skin, stomach and uterus, proving the effectiveness of this tissue origin prediction model in predicting the tissue origin of most cancers.

[0249] Figure 1A and Figure 1B The figure shows the ROC curve of the tissue origin prediction model according to an example of the present invention when predicting different tissues.

[0250] 3.2.2 Model construction and verification

[0251] In 3.2.1, the final selection of each tissue source prediction sub-model or second sub-model is shown in Table 6.

[0252] Table 6 Selection of each sub-model in the tissue source prediction model

[0253]

[0254] According to Table 6, the present invention selected 18 sub-models with excellent performance of AUC ≥ 0.9 for the construction of tissue source prediction model. The present invention constructed the tissue source prediction model by using specific transcripts in the training set samples for modeling. The algorithm model was determined after verification in the test set and / or validation set.

[0255] The tissue origin prediction model of the present invention is used to diagnose the tissue origin of the patient's cancer based on the full-length RNA transcript expression information of the patient's tumor tissue (including the specific transcript (Specific RNA transcript, SRT) data expressed by the tissue). The AUC of the best algorithm of the prediction model of the tissue origin of 18 cancers (covering at least 24 cancers) is greater than 0.9, including adrenal gland, bladder, bone marrow, brain, breast, intestine, eye, ganglion, liver, pancreas, prostate, thyroid, B cells, kidney, head and neck, skin, stomach and uterus. Therefore, in the application of this model, the tissue origin prediction only involves predicting the origin of 18 tissues (including cancer tissues) (covering at least 24 cancers verified by the present invention).

[0256] The schematic diagram of the tissue source prediction model of the present invention is as follows Figure 2 shown.

[0257] When predicting tissue origin, each time a judgment is made as to whether it is one of the tissues, such as when judging whether the tissue origin is the adrenal gland, the corresponding features are obtained from the tissue-specific transcripts according to the adrenal gland feature name set, and input into the classifier of the corresponding sub-model. When determining the corresponding classifier, the corresponding classifier can be obtained according to the classifier name. The classifier is preset with a classifier name and classifier parameters, and the classifier name and classifier parameters can be determined according to the corresponding sub-model in the preset model of tissue origin. The classifier parameters include predetermined hyperparameters to determine whether the tissue source is adrenal tissue. If the tissue source is not the adrenal gland, continue to judge whether the source is a B cell in the above manner, if the source is a B cell. As an example of the present invention, Figure 2 The results of prediction that the tissue origin is B cells are shown.

[0258] Example 4 Tissue Origin Diagnosis System and Device

[0259] Based on the tissue source prediction model constructed by the present invention, the present invention develops a tissue source prediction system or device. In the development process, software for tissue source prediction is developed.

[0260] The above tissue source prediction system includes, in addition to an input module and an output module, a preprocessing module and a prediction module (or an evaluation module).

[0261] The preprocessing module is used to quantify the expression value of the full-length RNA transcript of the sample to be tested.

[0262] Figure 3 A schematic diagram of a prediction module of an example of the present invention is shown. The module extracts the normal tissue-specific RNA transcript values ​​of the sample to be tested based on the feature name set and inputs them into the corresponding classifier. The classifier (submodel) and its parameters of each tissue category are determined by a fixed configuration file. Each classifier predicts the probability and threshold of whether the input sample belongs to this category.

[0263] The prediction module is configured to respond to receiving the full-length RNA transcript expression information of the tissue to be tested and input it into a pre-trained tissue origin prediction model; and determine the tissue origin of the tissue to be tested according to the output of the tissue origin prediction model.

[0264] Taking the adrenal gland as an example, after the expression information of the full-length RNA transcript of the tissue to be tested is input into the tissue origin prediction model, a candidate prediction result is obtained, including a prediction probability and a preset reference threshold; when the prediction probability in the prediction result is greater than or equal to the preset probability threshold, it is judged as positive, which is used to characterize that the target sample (or the sample to be tested) is derived from the adrenal gland.

[0265] Figure 4 A schematic diagram showing a tissue source prediction system according to an example of the present invention.

[0266] like Figure 4 As shown in the figure, when a sample needs to be diagnosed, the prediction module reads the sample file and divides it into 18 test sets according to the features. Tissue source prediction is performed.

[0267] After inputting the expression information of the full-length RNA transcript of the tissue to be tested into the tissue source prediction model, a sub-model corresponding to a tissue source is determined from the tissue source prediction model; the expression information of the full-length RNA transcript is input into the sub-model corresponding to the one tissue source to obtain the output of the sub-model, and the output of the sub-model is the predicted value that the tissue to be tested belongs to the tissue source corresponding to the sub-model. Determine the sub-model corresponding to the next tissue source, input the expression information of the full-length RNA transcript into the sub-model corresponding to the next tissue source, and obtain the output of the next sub-model until the outputs of all sub-models are determined.

[0268] The output of the tissue source prediction model is generated according to the tissue sources corresponding to one or more sub-models whose output prediction values ​​are greater than the tissue source prediction threshold, and the prediction probabilities corresponding to multiple candidate primary sites (or tissue sources) are obtained. The prediction probabilities in the candidate prediction results are sorted to obtain a prediction probability sorting result.

[0269] The tissue source prediction result is determined according to the candidate primary site corresponding to the maximum prediction probability in the prediction probability sorting result. When the maximum prediction probability is greater than or equal to the preset probability threshold, the candidate primary site corresponding to the maximum prediction probability is used as the tissue source prediction result. Or, when the maximum prediction probability is less than the preset probability threshold, the candidate primary sites corresponding to the first three prediction probabilities in the prediction probability sorting result are used as the tissue source prediction result. The prediction probability sorting results are arranged in descending order according to the probability value.

[0270] The following is relevant information for developing software:

[0271] Software name: Machine learning-based specific RNA transcript tissue source diagnosis system

[0272] Development hardware environment

[0273] CPU AMD Ryzen 7 6800U with Radeon Graphics

[0274] Memory 64GB RAM (2100MHz)

[0275] The operating system for which the software was developed

[0276] Distributor ID: Ubuntu

[0277] Description: Ubuntu 22.04.3LTS

[0278] Release:22.04

[0279] Codename: jammy

[0280] Software development environment / development tools: Vim 8.2

[0281] Programming language: Python 3.

[0282] The output module of the tissue source prediction device, software or system of the present invention is used to output the tissue source of the sample to be tested based on the prediction result of the tissue source prediction model. The tissue source of the tissue to be tested may include one or more. If the tissue to be tested is cancer tissue, the tissue source of the cancer may include one or more.

[0283] Example 5 External Validation

[0284] The present invention also obtains a first external detection set and a second external detection set, and performs a first external verification and a second external verification respectively.

[0285] 5.1 First External Validation

[0286] The first external validation set is used to evaluate the accuracy of the tissue origin prediction model or device or system of the present invention in diagnosing primary cancer and metastatic cancer. The data of the first external detection set is the detection data of the present invention, specifically including: 128 cases of hepatocellular carcinoma, 16 cases of colorectal cancer, 1 case of renal cancer, 1 case of gastric cancer, 34 cases of liver metastasis (12 cases of breast cancer liver metastasis, 19 cases of colorectal cancer liver metastasis, 3 cases of gastric cancer liver metastasis).

[0287] The tissue-specific RNA transcripts of the cancer samples from different tissues or different tissues of the primary site are input into the tissue origin prediction model, device or system of the present invention to obtain the prediction or diagnosis results shown in Tables 7 and 8 below.

[0288] Table 7 Diagnostic compliance rate of tumor tracing (tissue origin)

[0289]

[0290] Table 8 Diagnostic or prediction results of some patients with multiple cancer types

[0291]

[0292]

[0293] Note: NA is below the threshold. The score (or predicted probability) is calculated by the machine model selected.

[0294] As shown in Table 7, from the performance of predicting the origin of tumors, the overall correct first-place coincidence rate of the tissue origin prediction model is 80.6%, that is, in more than 80.6% of cases, the cancer prediction model can correctly predict the origin of cancer in the first attempt. In addition, its overall three-place coincidence rate is as high as 98.4%, indicating that even if the first choice prediction is not correct, there is a 98.4% chance that the first three most likely options listed contain the correct answer.

[0295] As shown in Table 8, predictions are made for the primary tumor and metastatic tissue samples of patients with liver cancer, colorectal cancer, breast cancer, gastric cancer, breast cancer liver metastasis, colorectal cancer liver metastasis, and gastric cancer liver metastasis. The top three most likely options for the prediction results include the patient's tissue pathology examination results.

[0296] 5.2 Second External Verification

[0297] The second external validation set was used to evaluate the accuracy of the tissue-derived prediction model of the present invention in diagnosing the primary site of metastatic cancer. From the second external validation data set (transcriptome sequencing of 500 tumor patients and 22 tissue and organ metastatic foci), 8 types of cancer with known sources (65 cases of prostate cancer, 12 cases of head and neck cancer, 6 cases of adrenal cancer, 5 cases of brain tumor, 4 cases of bladder cancer, 2 cases of pancreatic cancer, 2 cases of melanoma and 1 case of thyroid cancer) were selected for prediction. The results are shown in Tables 9 and 10.

[0298] Table 9 Tumor tracing diagnosis compliance rate

[0299]

[0300] Table 10 Diagnostic or prediction results of some patients with multiple cancer types

[0301]

[0302]

[0303] Note: NA means below the threshold.

[0304] As shown in Table 9, among the 8 cancer metastasis tissues with known primary tumor sources, the overall correct first-place coincidence rate of the tissue source prediction model was 81.3%, that is, in more than 81.3% of cases, the cancer prediction model was able to correctly predict the source site of the cancer in the first attempt. In addition, its overall three-place coincidence rate was as high as 96.9%, indicating that even if the first choice prediction was not correct (or the result with the highest prediction probability was not the correct result), the first three most likely options listed had a 96.9% chance of containing the correct answer.

[0305] As shown in Table 10, predictions were made for metastatic tissue samples of patients with prostate cancer metastasis, head and neck cancer metastasis, adrenal cancer metastasis, brain glioma metastasis, melanoma metastasis, and thyroid cancer metastasis, and the top three most likely options for the prediction results all included the patient's tissue pathology examination results.

[0306] The present invention further collects more primary and metastatic cancer tissue samples, including more cancer types (covering 18 types) and more samples, and performs RNA sequencing to verify the accuracy of the model.

[0307] The above test verifications demonstrate the potential of the present invention as a valuable auxiliary tool in clinical practice, which may guide the best treatment strategy for CUP patients and further prolong overall survival. It is worthy of in-depth study in prospective randomized trials in the future.

[0308] discuss

[0309] Tumor pathological diagnosis is one of the gold standards for diagnosing many types of tumors, with a high degree of accuracy. Through tumor pathological diagnosis, the nature of the tumor can be clarified, the origin of the tumor can be determined, the tumor can be staged and classified, and its degree of differentiation can be evaluated. By observing and analyzing the histological characteristics, grading, and staging of the tumor, the patient's survival rate and recurrence risk can be predicted, providing an important reference for treatment decisions and patient management, thereby guiding individualized treatment strategies. However, different types of tumors may have similar histological characteristics, and sometimes patients may have multiple primary tumors or multiple independent malignant lesions at the same time. In addition, some metastatic tumors may not be able to determine their origin through routine histological examinations, such as cancer of unknown primary site (CUP). In some cases, pathology has certain limitations in diagnosing primary and / or metastatic tumors, and requires the use of molecular genetics and other technical means to assist in diagnosis.

[0310] CUP refers to a metastatic malignant tumor that is histologically confirmed but not found after a comprehensive and detailed examination of the primary site. CUP has the characteristics of early metastasis, rapid progression, multiple organ involvement in more than half of the patients, unknown metastasis pattern, and high mortality rate. The clinical treatment of CUP patients is mainly based on traditional chemotherapy, but the effect of empirical chemotherapy is poor, lacks specificity, and fails to improve the survival prognosis of patients. The tracing of the primary lesion (including: tumor type and tissue origin) is of great significance for the selection of CUP populations that may benefit from targeted therapy. The pathological diagnosis of CUP is a complex task. For example, it is difficult to determine the location of the primary tumor in CUP. There may be multiple potential primary sites, including different organs and tissues, and the clinical manifestations are not specific. Faced with the diagnostic difficulties of CUP, doctors usually use a combination of multiple diagnostic methods, including detailed medical history, physical examination, imaging examination, pathological analysis, immunohistochemical staining, and molecular biological testing. At the same time, they work with multidisciplinary teams, including pathologists, radiologists, internists, and oncologists, to work together to improve the diagnostic accuracy and treatment effect of tumors of unknown primary site. In addition, tumor pathology diagnosis is somewhat subjective and requires high professional knowledge and experience of doctors. Different pathologists may have different interpretations and judgments on the same tissue specimen. This may lead to differences in consistency of diagnosis between different hospitals and different experts, thereby affecting the comparability of treatment decisions and results. For this reason, the technicians of the present invention recognize the need to obtain more specific markers or biological characteristics of tissues from the molecular biology level to improve the accuracy of diagnosis.

[0311] In the human genome, more than 95% of genes have more than one transcript, and the same gene usually expresses different transcripts in different tissues or disease states. These different transcripts differ in structure and function. This diversity of transcripts is particularly evident in different tissues, cell types, and disease states. In particular, in the study of human cancer, the expression pattern of transcripts shows significant tissue and pathological state preferences. Specific RNA transcripts (SRT) are transcripts that are specifically expressed in a certain tissue or disease.

[0312] With the continued development of high-throughput RNA sequencing technology and the significant improvement of computational analysis methods, researchers are now able to conduct a deeper analysis of transcriptional regulatory mechanisms under normal physiological and pathological conditions, and are expected to identify new transcripts. Although short-read RNA-seq technology has been widely used in transcriptomics research, it has inherent limitations in reconstructing transcripts, especially when multiple transcript isoforms of the same gene share exons, which makes it difficult to accurately distinguish and identify. Given the high complexity of the transcriptome and the above challenges, the task of fully reconstructing all transcripts relying solely on short-read data is particularly arduous. In contrast, long-read sequencing technology (LR-Seq), also known as third-generation sequencing technology, includes Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT), which has the advantage of producing longer reads (>10kb) and can accurately capture full-length isoforms. Third-generation sequencing technology facilitates comprehensive and accurate characterization of transcript structure by eliminating the need for transcript assembly. Thanks to its high accuracy in capturing complete transcript isoforms, the third-generation sequencing technology provides a detailed molecular basis for studying the transcriptome characteristics of tissues and cancers. There are also many studies, including our own, that have performed long-read RNA sequencing on tissues and samples and discovered unannotated tissue-specific transcripts.

[0313] The expression of different transcripts of the same gene in different tissues is closely related to the specific functions of the tissue. The definition of tissue-specific RNA transcripts is that a specific transcript of a gene is only expressed in a specific tissue and is not expressed or very low in other tissues. The specific RNA transcripts expressed in normal tissues are called tissue SRT (ti-SRT), that is, tissue-specific RNA transcripts.

[0314] Therefore, the present invention uses long-read and short-read RNA sequencing technologies to analyze tumor and normal tissue samples, and integrates RNA sequencing third-generation sequencing data from a variety of normal and cancerous human tissues and cells, identifying a large number of previously unannotated tissue-specific RNA transcripts in large-scale sample sets. These results have laid a solid scientific foundation for further research and implementation of precise RNA diagnosis and treatment strategies.

[0315] All documents mentioned in the present invention are cited as references in this application, just as each document is cited as reference individually. In addition, it should be understood that after reading the above teachings of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the claims attached to this application.

Claims

1. A method for constructing a tissue source prediction model, characterized in that: Includes steps: (s1) Process the third-generation sequencing data and RNA sequencing (RNA-seq) data of cancer tissue samples, paracancerous tissue samples, and normal tissue samples from different tissue sources to obtain reference GTF files of full-length RNA transcripts and tissue-specific RNA transcripts (Ti-SRT) of different tissues; (s2) providing training set, test set and validation set data for model construction, wherein the data includes expression information of full-length RNA transcripts of RNA-seq data obtained according to the reference GTF file after RNA sequencing of tissue samples from different tissue sources; the expression information includes expression information of tissue-specific RNA transcripts of the different tissues; Wherein, for a certain tissue, the expression information of the tissue-specific RNA transcripts of the tissue is the expression information of the tissue-specific RNA transcripts of the tissue-positive sample, and the expression information of the tissue-specific RNA transcripts of other tissues for the tissue is the expression information of the tissue-specific RNA transcripts of the tissue-negative sample; (s3) Construction of tissue origin sub-model: constructing different sub-models for different tissues to determine whether the sample to be tested originates from the tissue; comprising the steps of: for any tissue, inputting the expression information of tissue-specific RNA transcripts of tissue-positive samples and tissue-negative samples for the tissue into different machine learning models for training; thereby obtaining a tissue origin prediction sub-model for determining whether the sample to be tested originates from the tissue; and using this step, obtaining tissue origin prediction sub-models for all tissues; The machine learning models include: Random Forest model, Decision Tree model, XGBoost model, LightGBM model and CatBoost model; and (s4) Construction of a tissue origin prediction model: Based on the tissue origin prediction sub-models of all tissues obtained in (s3), the test set and validation set data are used for optimization selection to construct a tissue origin prediction model.

2. The method according to claim 1, characterized in that The expression information of the transcript includes the expression amount of the transcript measured in the form of TPM (Transcript per million).

3. The method according to claim 1, characterized in that In step (s1), the following steps are included: (a) Collect the third-generation sequencing data and RNA sequencing data of cancer tissue samples, paracancerous tissue samples, and normal tissue samples from different tissue sources, and generate non-redundant transcriptome GTF files; (b) optimizing the non-redundant transcriptome GTF file, wherein, for a certain tissue RNA-seq, if all splicing sites of a transcript are detected in ≥m samples, the transcript is retained, where m is 1-90%, preferably 2-50%; (c) Each transcript was compared with the reference transcriptome (GENCODE v.35), and the transcripts were mainly divided into four groups, including full splicing matches, incomplete splicing matches, novel in the catalog, and novel out of the catalog. After filtering, a reference GTF file for full-length RNA transcripts was established: Among them, for all incomplete splicing matching transcripts, such transcripts are filtered out from the entire transcriptome; Filter out unqualified novel transcripts in the catalog and outside the catalog (for example, cage peaks and polyA sites were annotated using SQANTI3, and novel transcripts in the catalog and outside the catalog were filtered when there were 16 or more adenines within 20 bases downstream of the annotated transcription termination site); For similar transcripts, the longest transcript is retained and other transcripts are filtered out; (d) quantifying the full-length transcripts of the tissue sample RNA-seq data based on the reference GTF file for the full-length RNA transcripts, thereby obtaining expression information of the full-length RNA transcripts; and (e) By comparison, RNA transcripts with significant differences in expression information between each tissue sample and other tissue samples are obtained, thereby obtaining tissue-specific RNA transcripts of different tissues.

4. The method according to claim 1, characterized in that In step (s3), the following sub-steps are included: (s3a) constructing different tissue origin prediction sub-models for different tissues to determine whether the sample to be tested originates from the tissue; for any tissue, inputting the expression information of tissue-specific RNA transcripts of tissue-positive samples and tissue-negative samples of the tissue into the following five machine learning models for training: Random Forest model, Decision Tree model, XGBoost model, LightGBM model and CatBoost model; thereby obtaining the hyperparameters of the five machine learning models corresponding to any one of the tissues; and based on this, obtaining five trained machine learning models (or the first tissue origin prediction sub-model) for determining the origin of any one of the tissues; (s3b) using the test set and validation set data to score the five first tissue origin prediction sub-models for any tissue origin judgment using the following scoring indicators: AUC value, accuracy, sensitivity and specificity; For any type of tissue, the first tissue source prediction sub-model with the highest score is used as the final tissue source prediction sub-model (or the second tissue source prediction sub-model) for determining the tissue; and (s3c) Obtaining tissue origin prediction sub-models for all tissues, that is, all second tissue origin prediction sub-models used for tissue origin prediction.

5. The method according to claim 1, characterized in that In step (s4), the following steps are included: (s4a) based on the tissue origin prediction sub-model of all tissues obtained in (s3), using the validation set data, and using the rocauc function to score the tissue origin sub-model of all tissues; (s4b) All tissue origin prediction sub-models with AUC ≥ 0.9 were selected to construct a tissue origin prediction model.

6. The method according to claim 1, characterized in that The tissue samples include tissue samples from the following tissues: adrenal gland, B cell, bladder, brain, breast, bone and / or muscle, bone marrow, cervix, intestine, esophagus, eye, ganglion, kidney, liver, lung, head and neck, ovary, pancreas, prostate, skin, stomach, testicle, thymus, thyroid and uterus.

7. A reference GTF file for full-length RNA transcripts, characterized in that The steps include the following: (i) Collect the third-generation sequencing data and RNA-seq sequencing data of cancer tissue samples, paracancerous tissue samples, and normal tissue samples from different tissue sources, and generate non-redundant transcriptome GTF files; (ii) optimizing the non-redundant transcriptome GTF file, wherein, for a certain tissue RNA-seq, if all splicing sites of a transcript are detected in ≥m samples, the transcript is retained, where m is 1-90%, preferably 2-50%; (iii) Each transcript was compared with the reference transcriptome (GENCODE v.35), and the transcripts were mainly divided into four groups, including full splicing matches, incomplete splicing matches, novel in the catalog, and novel out of the catalog. After filtering, a reference GTF file for full-length RNA transcripts was established: Among them, for all incomplete splicing matching transcripts, such transcripts are filtered out from the entire transcriptome; Filter out unqualified novel transcripts in the catalog and outside the catalog (for example, cage peaks and polyA sites were annotated using SQANTI3, and novel transcripts in the catalog and outside the catalog were filtered when there were 16 or more adenines within 20 bases downstream of the annotated transcription termination site); For similar transcripts, the longest transcript is retained and other transcripts are filtered out.

8. A tissue source prediction system, characterized in that: The system comprises: (a) an input module, wherein the input module is configured to input the full-length RNA transcript expression information of the sample to be tested; (b) an evaluation module: the evaluation module is configured to input the full-length RNA transcript expression information into the tissue origin prediction model constructed by the method according to claim 1, thereby obtaining a tissue origin prediction result of the sample to be tested; and (d) An output module, wherein the output module is configured to output the prediction result.

9. The prediction system according to claim 8, characterized in that The assessments in the assessment module include: (i) inputting the full-length RNA transcript expression information into a tissue origin prediction model, and then inputting it into the sub-models of each tissue origin prediction model in turn, thereby obtaining the prediction values ​​of all sub-models; wherein each of the sub-models is used to independently determine whether the sample to be tested is derived from a specific tissue; (ii) determining whether the prediction values ​​of all sub-models are greater than or equal to the preset reference threshold C1 of the corresponding tissue source prediction; the determination includes: if the prediction value of the sub-model is greater than or equal to the preset reference threshold C1 of the corresponding tissue source prediction, then determining that the tissue is a candidate tissue of the sample to be tested; thereby obtaining all candidate tissues; and (iii) All candidate tissues are sorted from large to small according to their corresponding prediction values, thereby obtaining a sorting result of the candidate tissues, and obtaining the candidate tissue with the largest prediction value as the tissue source of the sample to be tested; that is, the tissue of the sample to be tested is derived from the candidate tissue with the largest prediction value.

10. A computer storage medium, characterized in that: The storage medium is used to store a computer program corresponding to the algorithm of the tissue origin prediction model constructed by the method described in the first aspect of the present invention.

Citation Information

Patent Citations

  • Methods and materials for identifying the origin of a carcinoma of unknown primary origin

    CN101365950A

  • Transcript classification method

    CN111916147A

  • Tumor specific circular RNA neoantigen identification method and device, equipment and medium

    CN115240773A

  • Method for identifying tissue source of intermediate mesenchymal stem cells of sample and application thereof

    CN115565608A

  • Gene transcript marker combination for typing diagnosis of non-small cell lung cancer and typing diagnosis device

    CN115595370A