Construction method and system of cross-platform organization traceability model and storage medium

By constructing a cross-platform tissue tracing model and using machine learning methods to screen and train gene expression profile data from different platforms, the problem of decreased accuracy caused by cross-platform data differences was solved, and the prediction accuracy and recall of the primary tumor location were improved.

CN115620898BActive Publication Date: 2025-12-19GENEIS TECH BEIJING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211368223.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2025-12-19
Estimated Expiration
2042-11-03

AI Technical Summary

Technical Problem

Existing cross-platform tumor tracing methods often show decreased test results on different platforms, leading to reduced model accuracy and making them unsuitable for use on multiple data platforms.

Method used

By constructing a cross-platform organization tracing model, machine learning methods are used to select and screen genes from gene expression profile data from different platforms. Variable parameter thresholds are set to screen the training set, and model conditions are adjusted to improve accuracy. Multi-omics databases are used for training and validation.

Benefits of technology

It improves the accuracy and recall of predicting the location of the primary tumor site, and is particularly suitable for patients with metastatic cancer whose primary site is unknown. It also improves the applicability and accuracy of the model on multiple data platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620898B_ABST
    Figure CN115620898B_ABST
Patent Text Reader

Abstract

The application discloses a construction method and system of a cross-platform organization traceability model and a storage medium. Most current tumor traceability methods are based on modeling on one platform data and then testing on another platform data. This often leads to good test results on the same platform and a decline in test results on another platform. The method of the application solves the problem of differences between platforms well, can make the organization traceability model applicable to multiple sets of data platforms, and improves the accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of tissue tracing, in particular to a cross-platform tissue tracing model construction method, system and storage medium. BACKGROUND

[0002] Cancer of unknown primary site (CUP) refers to a malignant tumor that is histologically diagnosed as metastatic cancer, but the primary site cannot be determined. This type of tumor accounts for about 5% of all tumors. The treatment of CUP is mainly based on empirical chemotherapy, and the prognosis of patients is generally poor, with a median survival time of only 8-11 months. Identifying the primary site of the tumor helps doctors develop targeted treatment plans and improve patient survival rates. However, about 20%-50% of CUP patients cannot find the primary site. With the publication of more and more cancer-related research data, tumor tissue tracing through tissue or blood sequencing combined with machine learning algorithms has become a new solution.

[0003] Research has found that tumors always retain the gene expression characteristics of their tissue origin during their occurrence, development, and metastasis. Based on this principle, several tumor tracing products based on nucleic acid expression have been developed and approved by the US FDA, such as CancerTYPE ID based on RT-PCR technology and Tissue Of Origin based on microarray technology.

[0004] Most current tumor tracing methods are based on modeling data from one platform and then testing on another platform. This often leads to good test results on the same platform and reduced results on another platform. The main reason is that different methods of detecting data from different platforms result in differences between the two types of data. Solving this difference problem can make the tissue tracing model applicable to multiple data platforms and improve the accuracy of the model.

[0005] The information in the background art is merely intended to explain the general context of the present application and should not be considered as admitting or in any form implying that this information constitutes prior art commonly known to those skilled in the art. SUMMARY

[0006] To solve at least some of the technical problems in the prior art, the present application provides a cross-platform tissue tracing method that greatly improves the accuracy of predicting the primary site of cancer tissue by selecting genes from gene expression profile data from different platforms and using machine learning methods. Specifically, the present application includes the following.

[0007] In a first aspect of the present application, a cross-platform tissue tracing model construction method is provided, which includes the following steps:

[0008] (1) providing a first multi-omics database and a second multi-omics database from different platforms, wherein the first multi-omics database and the second multi-omics database each respectively comprises a plurality of biological information data related to at least one cancer, and at least part of the biological information data of the first multi-omics database and the second multi-omics database are the same;

[0009] (2) dividing the second multi-omics database into a test set and a standby set, and combining the standby set and the first multi-omics database into a candidate training set;

[0010] (3) screening a training set from the candidate training set;

[0011] (4) training the training set to obtain an accuracy rate, changing the conditions for multiple times of training, and selecting an optimal model according to the accuracy rate obtained by each training; and

[0012] (5) verifying the optimal model using the test set.

[0013] In some embodiments, the method for constructing a cross-platform organization traceability model according to the present application, wherein step (3) screening a training set from the candidate training set comprises:

[0014] (3-1) setting a threshold value T1 of a first variable parameter corresponding to the data in the first multi-omics database, a threshold value T2 of a second variable parameter corresponding to the data in the standby set, and screening the corresponding data according to the threshold values T1 and T2 to obtain a preliminary screening data set, wherein the threshold values T1 and T2 are respectively a deviation value of the amount of biological information;

[0015] (3-2) setting a threshold value T3 of a third variable parameter, and further screening the data in the preliminary screening data set according to the threshold value T3 to obtain a training set, wherein the third variable parameter is the consistency of the data in the first multi-omics database with the corresponding data in the standby set.

[0016] In some embodiments, the method for constructing a cross-platform organization traceability model according to the present application, wherein the first variable parameter and the second variable parameter are respectively a coefficient of variation of the amount of biological information.

[0017] In some embodiments, the method for constructing a cross-platform organization traceability model according to the present application, wherein the calculation of the consistency of the data in the first multi-omics database with the corresponding data in the standby set comprises forming a one-dimensional vector for each biological information in the data in the first multi-omics database and the data in the standby set, respectively, and calculating the correlation and / or rank sum test value between the one-dimensional vectors.

[0018] In some embodiments, the method for constructing a cross-platform tissue tracing model according to the present application, wherein the accuracy of the model corresponding to the training set is trained by adjusting the threshold T1, T2 and T3 size, and the optimal model is selected according to the accuracy.

[0019] In some embodiments, the method for constructing a cross-platform tissue tracing model according to the present application, wherein the tissue tracing refers to the tracing of cancerous tissue with unknown primary tumor.

[0020] In some embodiments, the method for constructing a cross-platform tissue tracing model according to the present application, wherein the biological information data is selected from gene expression profile data, gene methylation data, gene mutation data and proteomics data.

[0021] In a second aspect of the present application, a cross-platform tissue tracing system is provided, comprising:

[0022] a. a data acquisition module configured to obtain biological information data from two different platforms or biological information data in the same form as that from two different platforms;

[0023] b. a prediction module comprising a prediction model, and configured to output a tissue tracing result when the biological information data from different platforms is input into the prediction model;

[0024] wherein the prediction model is a cross-platform tissue tracing model constructed according to any of the above methods.

[0025] In some embodiments, the cross-platform tissue tracing system according to the present application, further comprising a display module for displaying the tissue tracing result.

[0026] In a third aspect of the present application, a computer storage medium or cloud is provided, which stores a computer program, and the computer program is executed by a computer to implement the above-mentioned construction method or the cross-platform tissue tracing system.

[0027] Most current tumor tracing methods are based on modeling on one platform data and then testing on another platform data. This often leads to good test results on the same platform, but the test results may decrease on another platform. The method of the present application well solves the problem of differences between platforms, and can make the tissue tracing model applicable to multiple data platforms and improve the accuracy of the model.

[0028] The method of the present application is particularly suitable for patients with metastatic cancer of unknown primary origin, patients with lesions that cannot be determined as primary or recurrence of cancer, patients with rare malignant tumors, patients with limited tumor biopsy specimens that cannot be detected by conventional pathology, patients with no obvious treatment effect, patients with a history of multiple cancers, patients with different clinical histories and histological diagnoses, and the like.

[0029] In exemplary embodiments, the accuracy of the present application is verified by using 10-fold cross-validation on TCGA samples and actual clinical sample testing. 10-fold cross-validation randomly divides the sample data into 10 parts, and sequentially selects one part as the test set and the remaining 9 parts as the training set. After training the model with the 9 training sets, the 1 test set is tested. After completing the 10 training and testing processes, each sample is predicted once.

[0030] The present application compares the predicted primary tumor tissue with the actually known primary tumor tissue to calculate evaluation indexes commonly used in statistics, including accuracy, recall rate, and F1 score, and the like. The results show that the average accuracy and recall rate of the present application for tracing 20 kinds of cancers are 96%. The model built on TCGA is used to predict the data of GEO, and the results are all around 50%. BRIEF DESCRIPTION OF DRAWINGS

[0031] FIG. 1 tSNE plot of TCGA and GEO data. The dotted line circle is GEO, and the rest is TCGA.

[0032] FIG. 2 Schematic flowchart of the present application. DETAILED DESCRIPTION

[0033] A number of exemplary embodiments of the present application will now be described in detail with reference to the drawings, which are not to be construed as limiting the present application. Rather, the specific embodiments are to be understood as being illustrative of certain aspects, features and embodiments of the present application.

[0034] It should be understood that the terms used in the present application are merely for the purpose of describing particular embodiments and are not intended to limit the present application. In addition, for numerical ranges in the present application, it should be understood that the upper limit and the lower limit of the range and every intermediate value between them are specifically disclosed. Every smaller range within the stated value or within the stated range and every intermediate value between the stated values or within the stated range is also included in the present application. The upper limit and the lower limit of these smaller ranges can be included or excluded independently from the range.

[0035] All technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains unless clearly indicated otherwise. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present application, the preferred methods and materials are described. All publications mentioned in this specification are herein incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any reference is not an admission that it is prior art with respect to the present application. Unless otherwise indicated, all percentages are on a weight basis.

[0036] As used herein, the terms "first", "second", "third", etc. do not necessarily mean a sequential order or a positional order, and are not intended to limit the present application. They are merely used to distinguish features or operations described by the same technical terms.

[0037] As used herein, the term "tissue tracing" has the same meaning as tumor tissue tracing, cancer tissue tracing or cancerous tissue tracing, which refers to a process of analyzing the tissue origin of a tumor with an unknown primary site.

[0038] As used herein, the term "cross-platform" refers to being applicable to different platforms or databases. Different platforms or databases include different data, and cross-platform can reduce, decrease or avoid the impact of data differences between different platforms or databases. Therefore, as used herein, cross-platform can be understood as being applicable to various platforms or being applicable to various databases.

[0039] As used herein, the terms "first multi-omics database" and "second multi-omics database" are sometimes collectively referred to as "multi-omics databases". Each of the first multi-omics database and the second multi-omics database comprises a plurality of biological information data related to at least one cancer, preferably each multi-omics database comprises 2 or more, preferably 5 or more, more preferably 10 or more, such as 15 or more, 20 or more, 25 or more, 30 or more, 50 or more, 100 or more or even more types of cancer data. The types of cancer included in each of the first multi-omics database and the second multi-omics database can be the same or different. In different cases, the first multi-omics database and the second multi-omics database comprise at least one type of cancer in common, preferably 2 or more, 5 or more, more preferably 10 or more, such as 15 or more, 20 or more, 30 or more or 50 or more types of cancer.

[0040] In the present disclosure, the term "biological information data" refers to data related to biological molecules, examples of which include, but are not limited to, gene expression profile data, gene methylation data, gene mutation data, and proteomic data. Any one of the data, or a combination of two or more of the data, can be used in the present disclosure. The biological information data can also include other data information, such as data related to the source of the biological molecules, for example, information from different parts or tissues of the body, such as lung, liver, brain, kidney, ovary, etc. The biological information data at least includes information related to the amount of biological molecules, such as relative amount, absolute amount, expression amount, content, etc.

[0041] Model construction

[0042] In a first aspect of the present disclosure, a method for constructing a cross-platform organization traceability model is provided, which comprises at least the following steps:

[0043] (1) providing a first multi-omics database and a second multi-omics database from different platforms, wherein each of the first multi-omics database and the second multi-omics database comprises a plurality of biological information data related to at least one cancer, and the first multi-omics database and the second multi-omics database comprise the same type of biological information data;

[0044] (2) dividing the second multi-omics database into a test set and a standby set, and combining the standby set and the first multi-omics database into a candidate training set;

[0045] (3) screening a training set from the candidate training set;

[0046] (4) training the training set to obtain an accuracy rate, and selecting an optimal model according to the accuracy rate; and

[0047] (5) verifying the optimal model using the test set.

[0048] In the present disclosure, the source of the multi-omics database in step (1) is not limited, which can be a known public database or a self-established local database. Preferably, the first multi-omics database and the second multi-omics database have the same type of data, but there are differences between the same type of data.

[0049] In the present application, step (2) is a step of dividing the second multi-omics database, including dividing the second multi-omics database into a test set and a standby set, and combining the standby set and the first multi-omics database into a candidate training set. The candidate training set mixes the first multi-omics data and the second multi-omics data. Wherein, when the second multi-omics database is divided into a test set and a standby set, the amount of data in the test set and the standby set is not limited, generally speaking, the amount of data in the standby set is greater than that in the test set. An exemplary division method includes dividing the data in the second multi-omics database into multiple equal parts, taking one of the equal parts as the test set, and combining the remaining equal parts as the standby set.

[0050] In the present application, step (3) is a step of screening data in the candidate training set to obtain the training set. The general screening step includes a step of setting a variable parameter threshold and selecting data in different data sets or databases as valid data according to the variable parameter threshold. An exemplary screening step includes:

[0051] (3-1) setting a threshold T1 corresponding to a first variable parameter of the data in the first multi-omics database, a threshold T2 corresponding to a second variable parameter of the data in the standby set, and screening the corresponding data according to the thresholds T1 and T2 to obtain a preliminary screening data set, wherein the thresholds T1 and T2 are respectively a deviation value of the amount of biological information; and

[0052] (3-2) setting a threshold T3 of a third variable parameter, and further screening the data in the preliminary screening data set according to the threshold T3 to obtain the training set, wherein the third variable parameter is the consistency of the data in the first multi-omics database with the corresponding data in the standby set.

[0053] In some embodiments, the first variable parameter and the second variable parameter are respectively a coefficient of variation of the amount of biological information, such as a coefficient of variation of gene expression amount. Such as standard deviation / mean (std / mean) or variance / mean. In a certain embodiment, when the first variable parameter or the second variable parameter is greater than the corresponding threshold value, it indicates that the amount of the biological molecule is unstable in the corresponding cancer type, and therefore needs to be deleted from the first multi-omics database or the corresponding standby set. Similarly, each biological molecule in the candidate training set is screened to obtain a preliminary screening data set.

[0054] In some embodiments, the third variable parameter is the consistency degree of the data in the first multi-omics database in the candidate training set with the corresponding data in the standby set. The consistency degree is preferably calculated by forming a one-dimensional vector according to the amount of each biological molecule in the data in the first multi-omics database and the data in the standby set, and calculating the correlation and / or rank sum test value between the one-dimensional vectors.

[0055] In the present application, steps (2) and (3) can be performed sequentially and can be preferably repeated multiple times, for example, repeated 2 times or more, 5 times or more, 8 times or more, or even more.

[0056] In the present application, step (4) is a model training step, which includes a step of training the accuracy rate of the training set. The machine learning model known in the art can be used, examples of which include but are not limited to linear regression algorithm, logistic regression algorithm, naive Bayes algorithm, k- nearest neighbor algorithm, support vector machine algorithm, decision tree algorithm, random forest algorithm, k-means algorithm and gradient boosting algorithm, etc.

[0057] In an exemplary embodiment, when training the model, the model is trained to obtain the accuracy rate of the model corresponding to the training set by adjusting the size of the threshold T1, T2 and T3, and the optimal model is selected according to the size of the accuracy rate.

[0058] Those skilled in the art should understand that the numbers (1), (2) and the like are only for the purpose of distinguishing different steps and do not have the meaning of indicating the sequence of steps. The sequence of the above steps is not particularly limited as long as the purpose of the present application can be achieved. In addition, those skilled in the art should also understand that other steps or operations can be included before and after the above steps (1)-(4) or between any of these steps, for example, further optimizing and / or improving the method described in the present application.

[0059] Cross-platform organization traceability system

[0060] In a second aspect of the present application, a cross-platform organization traceability system is provided, which at least includes:

[0061] a. A data acquisition module is configured to obtain biological information data from two different platforms;

[0062] b. A prediction module includes a prediction model, and is configured to output an organization traceability result when biological information data from different platforms is input into the prediction model, wherein the prediction model is a cross-platform organization traceability model constructed by the method of the first aspect;

[0063] c. Optionally, further comprising a display module for displaying the organization traceability result.

[0064] In the present application, the data acquisition module is configured to obtain bioinformatics data from two different platforms, or to obtain bioinformatics data from other channels. An exemplary data acquisition module includes a multi-mode input port. The multi-mode input port is capable of receiving bioinformatics data from different sources, including online and offline sources. Another exemplary data acquisition module includes a port connected to a sequencing instrument, and further includes a data processing unit configured to process data obtained from the test instrument into the same form of bioinformatics data as in the two platforms.

[0065] In the present application, the form of the prediction module is not limited, as long as it includes the prediction model obtained by the method of the first aspect of the present application. Exemplarily, the prediction module is a processor. Preferably, the prediction module is communicatively connected with the data acquisition module, so as to be able to call the data in the data acquisition module and input the data into the prediction model for operation.

[0066] Embodiment

[0067] This embodiment is illustrated by taking the two groups of data of TCGA and GEO as examples. Specifically as follows:

[0068] I. Sample information

[0069] A PCA is made for the data of TCGA and GEO, as shown in the following figure. FIG. 1 It can be seen from the figure that the two groups of data have great differences.

[0070] RNAseq expression profile data of 7633 patients respectively suffering from 20 kinds of cancers from the TCGA database, and RNAseq expression profile data of 459 patients respectively suffering from the same 20 kinds of cancers from the GEO database are obtained. According to the clinical information of the patients, they are divided into GEO frozen primary data, a total of 269 cases, GEO frozen metastatic data, a total of 42 cases, GEO FFPE primary data, a total of 142 cases, TCGA selected frozen primary data, a total of 7579 cases.

[0071] II. Modeling steps

[0072] 1. Data division and preprocessing

[0073] 1.1 A database is established by using the expression data of 20 cancer types of bladder cancer, breast cancer, cervical squamous cell carcinoma and adenocarcinoma, colon cancer, glioblastoma multiforme, head and neck squamous cell carcinoma, renal clear cell carcinoma, renal papillary cell carcinoma, acute myeloid leukemia, brain low-grade glioma, hepatocellular liver cancer, lung adenocarcinoma, lung squamous carcinoma, ovarian serous cystadenocarcinoma, pancreatic cancer, prostate cancer, rectal adenocarcinoma, gastric cancer, thyroid cancer and endometrial cancer in TCGA and GEO projects. By screening specific expression data in TCGA and GEO as features, and taking cancer classification as labels, a database is established.

[0074] 1.2 Stratified sampling of GEO frozen primary data into three parts by cancer type.

[0075] 2. Model building

[0076] Model building is performed using all TCGA data and two out of three parts of GEO frozen primary data (referred to as GEOpart). In the model evaluation phase, the three parts of data are used in rotation.

[0077] First, three definitions are made for each gene:

[0078] (1) CV1 is the coefficient of variation CV of the expression of the gene in each cancer type in TCGA, i.e. (std / mean), which is 20 values, corresponding to 20 cancer types. (2) CV2 is the coefficient of variation CV of the expression of the gene in each cancer type in GEOpart data, which is also 20 values. (3) CORR is the correlation coefficient (such as Pearson correlation coefficient) between the average expression of the 20 cancer types in TCGA and the average expression of the 20 cancer types in GEOpart, which is 1 value.

[0079] Next, three hyperparameters T1, T2, T3 are defined, corresponding to the thresholds of CV1, CV2, CORR respectively. For a gene, (CV1 of all cancer types) < T1 and (CV2 of all cancer types) < T2 and CORR > T3 are simultaneously true, the gene is included in the subsequent consideration. The purpose of filtering is: (1) when the coefficient of variation of a gene in a cancer type is large, it represents that the gene is unstable in that cancer type, which can be filtered by T1, T2. (2) When the correlation coefficient of the average expression of a gene in each cancer type in TCGA and GEOpart data is low, it represents that the gene shows different sequencing tendencies on different platforms, which can be filtered by T3.

[0080] After determining the genes, random forest is used to train the training set data and test the model on the test set, and the prediction accuracy is used as the selection standard of hyperparameters T1, T2, T3. Grid search is used to combine multiple T1, T2, T3, and finally T1, T2, T3 are determined.

[0081] 3. Model evaluation

[0082] Then the random forest is used to model the TCGA frozen primary data and the GEO frozen primary data, the TCGA frozen primary data is always placed in the training data, and two of the three GEO frozen primary data are used as GEOpart to select genes, and the TCGA frozen primary data is added to the training data to select genes and train the model, and the remaining one of the GEO frozen primary data is tested to obtain the accuracy of the model, the accuracy obtained three times is averaged to obtain the final cross-validation accuracy of the model, according to the accuracy, T1, T2 and T3 are selected, and independent verification is performed on other data sets of GEO.

[0083] III. Result statistics

[0084] The accuracy and results of each data set at each threshold value are shown in Table 1, which is only modeled using TCGA frozen primary data without gene selection. After gene selection using TCGA frozen primary data and GEO frozen primary data and parameters T1, T2 and T3, the accuracy and results of the model at each data set at each threshold value are shown in Table 2.

[0085] Table 1: TCGA frozen primary data model accuracy statistics table

[0086]

[0087] Table 2: TCGA and GEO frozen primary data plus three-parameter gene selection model accuracy statistics table

[0088]

[0089] Comparing the results of Table 1 and Table 2, it is found that the accuracy of the model of the present application on the GEO frozen primary, frozen metastatic and FFPE primary data has been greatly improved compared to before, effectively solving the problem of poor prediction effect of cross-platform data sets.

[0090] Although the present application has been described with reference to exemplary embodiments, it should be understood that the present application is not limited to the disclosed exemplary embodiments. Various adjustments or changes can be made to the exemplary embodiments of the present application without departing from the scope or spirit of the present application. The scope of the claims should be based on the broadest interpretation to encompass all modifications and equivalent structures and functions.

Claims

1. A method for constructing a cross-platform organization provenance model, characterized in that, The method comprises the following steps: (1) providing a first multi-omics database and a second multi-omics database from different platforms, wherein the first multi-omics database and the second multi-omics database each respectively comprises a plurality of biological information data related to at least one cancer, and at least part of the biological information data in the data of the first multi-omics database and the second multi-omics database is the same; (2) dividing the second multi-omics database into a test set and a standby set, and combining the standby set and the first multi-omics database into a candidate training set; (3) screening a training set from the candidate training set, comprising: (3-1) setting a threshold T1 of a first variable parameter corresponding to the data in the first multi-omics database, a threshold T2 of a second variable parameter corresponding to the data in the standby set, and screening the corresponding data according to the thresholds T1 and T2 to obtain a preliminary screening data set, wherein the thresholds T1 and T2 are each a deviation value of the amount of biological information, and the first variable parameter and the second variable parameter are each a coefficient of variation of the amount of biological information; (3-2) setting a threshold T3 of a third variable parameter, and further screening the data in the preliminary screening data set according to the threshold T3 to obtain the training set, wherein the third variable parameter is the consistency of the data in the first multi-omics database with the corresponding data in the standby set; (4) training the training set to obtain an accuracy rate, changing the conditions for multiple times of training, and selecting an optimal model according to the accuracy rate obtained by each training; and (5) verifying the optimal model using the test set.

2. The method of claim 1, wherein, The calculation of the consistency of the data in the first multi-omics database with the corresponding data in the standby set comprises forming a one-dimensional vector for each biological information in the data in the first multi-omics database and the data in the standby set, respectively, and calculating the correlation and / or rank sum test value between the one-dimensional vectors.

3. The method of claim 1, wherein, The accuracy rate of the model corresponding to the training set is obtained by adjusting the sizes of the thresholds T1, T2 and T3 to train the model, and the optimal model is selected according to the accuracy rate.

4. The method of claim 1, wherein, The tissue tracing refers to the tracing of cancer tissue with unknown primary focus.

5. The method for constructing a cross-platform organization provenance model according to claim 1, wherein, The biological information data is selected from gene expression profile data, gene methylation data, gene mutation data and proteomics data.

6. A cross-platform organization traceability system, comprising: The method comprises: a. a data acquisition module configured to obtain biological information data from two different platforms or biological information data in the same form as the biological information data from the two different platforms; b. a prediction module comprising a prediction model and configured to output a tissue tracing result when the biological information data from different platforms is input into the prediction model; wherein the prediction model is a cross-platform tissue tracing model constructed according to the method of any one of claims 1-5.

7. The cross-platform organization provenance system of claim 6, wherein, Further comprising a display module for displaying the tissue tracing result.

8. A computer storage medium or cloud, characterized in that, A computer program is stored therein, and the computer program is executed by a computer to implement the construction method according to any one of claims 1-5, or the cross-platform tissue tracing system of claim 6 or 7.