Method for reconstructing a continuously varying transcriptome on the basis of image processing and apparatus
By reconstructing continuous changes in the transcriptome using a generative artificial intelligence model, the static nature of tumor omics data has been addressed, enabling accurate reconstruction and prediction of tumorigenesis and development, and providing traceability of tumorigenesis mechanisms.
Patent Information
- Application Number
- PCT/CN2024/094760
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-22
- Publication Date
- 2025-11-27
AI Technical Summary
Existing tumor omics data are static and cannot accurately identify key functional changes and future trends in the process of tumor development, nor can they trace the mechanisms of tumorigenesis.
Using a generative artificial intelligence approach, a transcriptome image reconstruction model was used to perform N-step average sampling to generate N+2 continuous dynamic transcriptome data. Principal component analysis was then performed to reconstruct key functional changes and their sequence during tumor development and progression, and to predict future trends.
The key functional changes, their sequence, and future trends during tumorigenesis were accurately reconstructed, solving the static nature of tumor omics data and enabling traceability of tumorigenesis mechanisms.
Smart Images

Figure CN2024094760_27112025_PF_FP_ABST
Abstract
Description
Method and device for reconstructing continuously changing transcriptome based on picture processing TECHNICAL FIELD
[0001] The present application relates to the field of generative artificial intelligence technology, in particular to a method and device for reconstructing continuously changing transcriptome based on picture processing. BACKGROUND
[0002] This section is intended to provide background or context to the embodiments of the application recited in the claims. The description herein does not constitute admission that the prior art is prior art nor does it constitute an admission of any description in this section as prior art to an application described herein and / or in another application also owned by the applicant of the present application.
[0003] Existing tumor omics data are static, that is, by collecting the static data of pathological specimens of tumor patient individuals to identify the static difference between the transcriptome of tumor tissue (abnormal tissue) samples and the transcriptome of paracancer normal tissue samples. Relying on static data to identify transcriptome changes cannot determine how the tumor actually occurs and how it changes. There are three reasons: 1) when the tumor is discovered, the tumor tissue has already undergone many changes, and the original state of the tumor has already disappeared; 2) whether through exon sequencing or through differential gene comparison, given that each tumor patient / sample has many mutant genes and many differentially expressed genes, it is impossible to determine which is the real cause and it is unknown what new mutations and new differentially expressed genes will be generated next; 3) although single-cell sequencing can construct the relationship between tumor sample cells, there is no evidence to support the existence of a continuous evolution relationship between tumor sample cells.
[0004] SUMMARY
[0005] The embodiments of the present application provide a method for reconstructing continuously changing transcriptome based on picture processing, which accurately reconstructs the transcriptome continuous change process by using transcriptome continuity dynamic data reconstructed based on transcriptome pictures, and further accurately reconstructs the key functional change events, event occurrence order and future transcriptome change trend in the tumor occurrence and development process. The method comprises:
[0006] Obtaining pictures of normal tissue transcriptome and its corresponding abnormal tissue transcriptome to be reconstructed;
[0007] Inputting the pictures of the normal tissue transcriptome and its corresponding abnormal tissue transcriptome to be reconstructed into a pre-trained generative transcriptome change process reconstruction model, performing N-step uniform sampling between the normal tissue transcriptome and its corresponding abnormal tissue transcriptome, and generating N+2 transcriptome continuity dynamic data; N is a positive integer greater than 1, and the N+2 transcriptome continuity dynamic data constitutes an individualized transcriptome evolution path, and the generative transcriptome change process reconstruction model is obtained by pre-training a generative network using a transcriptome picture sample data set;
[0008] The principal component analysis is performed on the individualized transcriptome evolution path to obtain key functional change events and event occurrence sequences in the tumor occurrence and development process represented by the principal components, and a tangent direction of the individualized transcriptome evolution path is used to indicate a transcriptome change trend to be occurred after a preset period of the tumor.
[0009] The embodiment of the present application provides a device for reconstructing continuously changing transcriptome based on picture processing, which is used for accurately reconstructing a transcriptome continuous change process by using transcriptome continuity dynamic data reconstructed based on transcriptome pictures, and further accurately reconstructing key functional change events, event occurrence sequences and transcriptome change trends to be occurred in the future in the tumor occurrence and development process.
[0010] The acquisition unit is configured to acquire pictures of a normal tissue transcriptome to be reconstructed and a corresponding non-normal tissue transcriptome.
[0011] The reconstruction unit is configured to input the pictures of the normal tissue transcriptome to be reconstructed and the corresponding non-normal tissue transcriptome into a pre-trained generative transcriptome change process reconstruction model, perform N-step uniform sampling between the normal tissue transcriptome and the corresponding non-normal tissue transcriptome, and generate N+2 transcriptome continuity dynamic data; N is a positive integer greater than 1, and the N+2 transcriptome continuity dynamic data constitute an individualized transcriptome evolution path; the generative transcriptome change process reconstruction model is obtained by pre-training a generative network using a transcriptome picture sample data set.
[0012] The processing unit is configured to perform principal component analysis on the individualized transcriptome evolution path to obtain key functional change events and event occurrence sequences in the tumor occurrence and development process represented by the principal components, and a tangent direction of the individualized transcriptome evolution path is used to indicate a transcriptome change trend to be occurred after a preset period of the tumor.
[0013] The embodiment of the present application further provides a computer device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the method for reconstructing continuously changing transcriptome based on picture processing when executing the computer program.
[0014] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the method for reconstructing continuously changing transcriptome based on picture processing.
[0015] The embodiment of the present application further provides a computer program product, which comprises a computer program, and the computer program is executable on a processor to implement the method for reconstructing continuously changing transcriptome based on picture processing.
[0016] Compared with the scheme in the prior art that recognizes the static difference between the non-normal tissue transcriptome and the normal tissue transcriptome by collecting pathological specimens of tumor patients, the scheme for reconstructing the continuously changing transcriptome based on picture processing provided in the embodiment of the present application inputs the pictures of the normal tissue transcriptome to be recognized and the corresponding non-normal tissue transcriptome into a pre-trained generative transcriptome change process reconstruction model, performs N-step uniform sampling between the normal tissue transcriptome and the corresponding non-normal tissue transcriptome, generates N+2 transcriptome continuity dynamic data, and equivalently reconstructs the transcriptome continuity dynamic data; N is a positive integer greater than 1, and the N+2 transcriptome continuity dynamic data constitutes an individualized transcriptome evolution path; principal component analysis is performed on the individualized transcriptome evolution path to obtain key functional change events and event occurrence sequences in the tumor occurrence and development process represented by principal components, and a tangent direction of the individualized transcriptome evolution path is used to indicate the transcriptome change trend that will occur after a preset period of time of the tumor, so that the transcriptome continuity dynamic data reconstructed based on the transcriptome pictures is used to accurately reconstruct the transcriptome continuous change process, and then the key functional change events, the event occurrence sequences and the transcriptome change trend that will occur in the future in the tumor occurrence and development process are accurately reconstructed. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings, which are included to provide a further understanding of the present application and constitute a part of this application, illustrate embodiments of the present application and do not limit the present application. In the drawings:
[0018] Fig. 1 is a flowchart of the method for reconstructing the continuously changing transcriptome based on picture processing in the embodiment of the present application;
[0019] Fig. 2 is a flowchart of the process of pre-training the transcriptome change process reconstruction model in the embodiment of the present application;
[0020] Fig. 3A is a schematic diagram of the development trajectory of thirty kinds of clinically common tumors from the normal cancer-adjacent tissue transcriptome to the corresponding cancer tissue transcriptome in the principal component I (PC1) and the principal component II (PC2) space; Fig. 3B is a graph of the transcriptome level change of representative early neural development related genes in the artificial intelligence deduced neuroblastoma occurrence path; Fig. 3C is a schematic diagram of the expression levels of PLP1 and SOX10 in the Bulk-seq clinical sample; Fig. 3D is a UMAP cell clustering diagram of adrenal single cell sequencing of embryos and neuroblastoma patients; Fig. 3E is a schematic diagram of the main composition of neuroblastoma found by UMAP clustering analysis; Figs. 3F to 3G are schematic diagrams of SOX10 (F) and PLP1 (G) tables;
[0021] Figs. 4A to 4H are schematic diagrams of the reconstruction principle of the individualized occurrence path of lung cancer in the embodiment of the present application; Figs. 4I to 4K are schematic diagrams of the principle of existing single cell sequencing;
[0022] Figure 5 is a structural schematic diagram of an apparatus for reconstructing a continuously changing transcriptome based on picture processing according to an embodiment of the present application. DETAILED DESCRIPTION
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, further detailed descriptions of the embodiments of the present application are given below with reference to the drawings. Here, the illustrative embodiments of the present application and their descriptions are used to explain the present application but are not intended to limit the present application.
[0024] The information collected in the technical solution of the present application is information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with relevant laws, regulations, and standards of relevant countries and regions, necessary security measures are taken, public order and good customs are not violated, and corresponding operation portals are provided for the user to choose authorization or refusal. The present application provides corresponding operation portals for the user to choose to agree or refuse the automatic decision result; if the user chooses to refuse, the expert decision process is entered.
[0025] The method for reconstructing a continuously changing transcriptome based on picture processing claimed in the present application is an information processing method in which all steps are implemented by a computer or other apparatus.
[0026] In view of the technical problems existing in the prior art, the embodiments of the present application provide a scheme for reconstructing a continuously changing transcriptome based on picture processing, which is a scheme for reconstructing a continuously changing transcriptome in a tumor occurrence process based on generative artificial intelligence. The scheme can accurately reconstruct a continuously changing transcriptome in a tumor occurrence and development process based on transcriptome picture deep learning, and further accurately reconstruct key functional change events, event occurrence sequences, and future transcriptome change trends in the tumor occurrence and development process. The scheme for reconstructing a continuously changing transcriptome based on picture processing is described in detail below.
[0027] Figure 1 is a flow schematic diagram of a method for reconstructing a continuously changing transcriptome based on picture processing according to an embodiment of the present application, as shown in Figure 1, the method includes the following steps:
[0028] Step 101: obtaining pictures of a normal tissue transcriptome to be reconstructed and a corresponding non-normal tissue transcriptome;
[0029] Step 102: input the picture of the normal tissue transcriptome to be reconstructed and the corresponding non-normal tissue transcriptome into the pre-trained generative transcriptome change process reconstruction model, perform N-step uniform sampling between the normal tissue transcriptome and the corresponding non-normal tissue transcriptome, and generate N+2 transcriptome continuity dynamic data; N is a positive integer greater than 1, and the N+2 transcriptome continuity dynamic data constitutes an individualized transcriptome evolution path; the generative transcriptome change process reconstruction model is obtained by pre-training a generative network using a transcriptome picture sample data set;
[0030] Step 103: perform principal component analysis on the individualized transcriptome evolution path to obtain key functional change events and event occurrence sequences in the tumor occurrence and development process represented by principal components, and a tangent direction of the individualized transcriptome evolution path for indicating a transcriptome change trend to be occurred after a preset period of the tumor.
[0031] The method for reconstructing continuously changed transcriptomes based on picture processing provided by the embodiment of the application works as follows: obtaining pictures of a normal tissue transcriptome to be reconstructed and a corresponding non-normal tissue transcriptome; inputting the pictures of the normal tissue transcriptome to be reconstructed and the corresponding non-normal tissue transcriptome into a pre-trained generative transcriptome change process reconstruction model (a transcriptome generative artificial intelligence model), performing N-step uniform sampling between the normal tissue transcriptome and the corresponding non-normal tissue transcriptome, and generating N+2 transcriptome continuity dynamic data, which is equivalent to reconstructing the transcriptome continuity dynamic data; N is a positive integer greater than 1, and the N+2 transcriptome continuity dynamic data constitutes an individualized transcriptome evolution path; performing principal component analysis on the individualized transcriptome evolution path to obtain key functional change (transformation) events and event occurrence sequences in the tumor occurrence and development process represented by principal components, that is, key functional changes and change sequences, which can include origin cells and genes in tumor occurrence, intermediate transitional events in the tumor occurrence process (for example, intermediate transitional event information corresponding to the rising and falling of the expression level of a nerve cell characteristic gene in FIG. 3B, which proves that a cell gene mutation is likely to occur), and the like, that is, for example, some tumors originate from inflammation, some tumors originate from mutation of a tumor suppressor gene, some tumors originate from dedifferentiation, and some tumors have all these events but have different sequences, and a tangent direction of the individualized path (the individualized transcriptome evolution path) (for example, the tangent direction represents a “future trend” in FIG. 4) indicates a transcriptome change trend to be occurred after a preset period of the tumor, that is, a future evolution trend.
[0032] In the embodiment of the application, the meaning of “reconstruction” can be: constructing transcriptome continuity dynamic data based on the obtained pictures of the normal tissue transcriptome and the corresponding non-normal tissue transcriptome.
[0033] Compared with the scheme in the prior art that recognizes static differences between non-normal tissue transcriptomes and normal tissue transcriptomes by collecting pathological specimens of tumor patients, the method for reconstructing continuously changing transcriptomes based on picture processing provided in the embodiments of the present application realizes continuously dynamic data of transcriptomes reconstructed by interpolation sampling based on transcriptome pictures, accurately reconstructs the continuously changing process of transcriptomes, i.e., reconstructs the continuously changing process between the cancer-adjacent normal tissue transcriptome and the cancer tissue transcriptome, and further accurately reconstructs the key functional change events, the order of events, and the transcriptome change trend (gene expression change or expression change gene) that will occur in the future in the process of tumor occurrence and development, i.e., can accurately reconstruct 1) key functional change events, 2) the order of events, 3) transitional events in the process of tumor occurrence, and 4) the subsequent trend of tumor development in the process of tumor occurrence. The 2) to 4) above cannot be obtained by comparison of static transcriptome data. The method for reconstructing continuously changing transcriptomes based on picture processing is described in detail below.
[0034] The method for reconstructing continuously changing transcriptomes based on picture processing provided in the embodiments of the present application is a method for reconstructing transcriptome changes in the process of tumor occurrence based on generative artificial intelligence, which solves the problem that the tumor pathogenesis cannot be reconstructed and traced, and discovers the general evolution of transcriptomes in the process of tumor occurrence and the individualized evolution process of transcriptomes in the process of tumor occurrence of each patient and each tumor sample. The method is described in detail below.
[0035] The method for reconstructing continuously changing transcriptomes based on picture processing provided in the embodiments of the present application involves a method for constructing a tumor pathogenesis path and subsequent evolution, which includes the following steps:
[0036] 1. A generative transcriptome change process reconstruction model (transcriptome generative artificial intelligence model) is obtained by generative artificial intelligence training;
[0037] 2. The trained model is used for interpolation sampling to generate intermediate states between any two points. For example, according to the tumor tissue transcriptome (non-normal tissue transcriptome) and the cancer-adjacent normal tissue (normal tissue within the preset range of non-normal tissue) transcriptome of a tumor patient, a series of intermediate mutation states between the normal to cancer tissue mutation spectrum are generated by interpolation sampling, and the transcriptome evolution process (transcriptome evolution path) is calculated;
[0038] 3. Based on the reconstruction of the transcriptome series (transcriptome continuity dynamic data), the universal rules of tumor occurrence and evolution (such as the key functional change events in the development of a certain tumor, the order of the events, and the transcriptome change trend after a preset period) and the individual-specific rules (such as the key events of a tumor in a certain patient and the order of the events, and the gene expression change trend, i.e., the transcriptome change trend, after a preset period) are discovered.
[0039] The method of reconstructing the continuously changing transcriptome based on image processing will be described in detail below in combination with FIGS. 2 to 4.
[0040] 1. Transcriptome image generation
[0041] During the generation of the transcriptome image, each gene is fixed at the corresponding position in the 512x512 coordinate system, and the expression amount of the gene corresponds to the 512x512 array pixel value. The conversion method from the expression amount to the pixel intensity is log(FPKM+1)x14, and the value range is limited between 0-255. The transcriptome image can be a single-channel grayscale image.
[0042] As can be seen from the above, in an embodiment, during the generation of the transcriptome image, each gene is fixed at the corresponding position in the 512x512 coordinate system, and the expression amount of the gene corresponds to the 512x512 array pixel value. The conversion method from the expression amount to the pixel intensity is log(FPKM+1)x14, and the value range is limited between 0-255. This implementation can improve the generation effect of the transcriptome image, further improve the accuracy of the reconstruction of the continuously changing process of the transcriptome, and further improve the reconstruction accuracy of the key functional change events, the order of the events, and the transcriptome change trend that will occur in the future in the development process of the tumor.
[0043] 2. Transcriptome image library
[0044] The transcriptome image dataset includes the transcriptome of normal tissue and tumor tissue (non-normal tissue). In this embodiment, the transcriptome dataset is divided into 92 categories, 500 samples are randomly taken from each category, and 46000 transcriptome grayscale images are generated according to the method of “transcriptome image generation” described above. These images are the generated model training image library of this embodiment.
[0045] 3. Model training
[0046] In the embodiment of the present application, the generative model adopts a StyleGAN of GAN for model training. In specific implementation, the model training can use a NVIDIA A100 card, 8 pictures per batch, 46000 pictures per round, and can be iteratively trained for 3500 rounds. That is, in an embodiment, the generative transcriptome change process reconstruction model is a model based on a trained generative network, preferably, the generative network can be a generative StyleGAN network, which can improve the accuracy of subsequent generated transcriptome continuity dynamic data.
[0047] As described above, in an embodiment, as shown in FIG. 2, the method of reconstructing a continuously changing transcriptome based on picture processing further includes pre-training the generative transcriptome change process reconstruction model in the following manner:
[0048] Step 201: generating transcriptome picture data;
[0049] Step 202: obtaining a transcriptome picture sample data set; the data set includes positive sample data and negative sample data, the positive sample data is tumor tissue transcriptome picture data, and the negative sample data is normal tissue transcriptome picture data;
[0050] Step 203: training a generative network using the transcriptome picture sample data set to obtain the generative transcriptome change process reconstruction model.
[0051] The above steps 1, transcriptome picture generation, step 2, transcriptome picture library, and step 3, model training, are preparation steps before reconstructing a continuously changing transcriptome, that is, steps of pre-training a generative transcriptome change process reconstruction model.
[0052] In specific implementation, when pre-training the generative transcriptome change process reconstruction model, the embodiment of the present application first obtains transcriptome pictures of each tissue and organ in normal and disease states of a human being; converts the transcriptome pictures into grayscale pictures, each pixel in the picture represents a gene, and the brightness of the pixel represents the expression level of the gene; uses these grayscale pictures to form a training library and train a generative artificial intelligence model (generative transcriptome change process reconstruction model); when subsequently reconstructing transcriptome continuity dynamic data using the model, obtains pictures of cancer-adjacent normal tissue transcriptome and cancer tissue transcriptome of a cancer patient (patient) to be reconstructed, and converts the pictures of the transcriptome to be reconstructed into grayscale pictures and inputs them into the pre-trained generative transcriptome change process reconstruction model to reconstruct the transcriptome change process.
[0053] 4. Reconstruction of tumor occurrence mechanism based on generative artificial intelligence model
[0054] In view of the heterogeneity of tumors, the present embodiment averages the expression of each gene in all samples of the tumor tissue (abnormal tissue) transcriptome of all patients of each tumor to obtain the average transcriptome of the given tumor. The present embodiment also averages the expression of each gene of the cancer-adjacent normal tissue transcriptome (normal tissue transcriptome) of all tumor patients and the transcriptome of the corresponding tissue of all normal persons (non-cancer patients) in the transcriptome data set to obtain the average transcriptome of the normal tissue. Then, the average transcriptome of the given normal tissue and the average transcriptome of the corresponding cancer tissue (abnormal tissue) are input into the trained model, 98-step uniform sampling is performed between each tumor tissue and its normal tissue transcriptome, and then a generative network is used to generate a total of 100 transcriptomes from the normal tissue transcriptome to the corresponding tumor transcriptome (abnormal tissue transcriptome). The 100 transcriptomes constitute the average occurrence path of the given tumor.
[0055] As can be seen from the above, in one embodiment, the method of reconstructing continuously changing transcriptomes based on image processing further includes the following steps of discovering universal laws of tumor occurrence and evolution:
[0056] The average expression of each gene in all samples of the abnormal tissue transcriptome of all individuals of each tumor is taken as the average transcriptome of the abnormal tissue.
[0057] The average expression of each gene of the normal tissue transcriptome of all abnormal individuals and the transcriptome of the corresponding tissue of all normal individuals in the transcriptome data set is taken as the average transcriptome of the normal tissue.
[0058] The average transcriptome of the normal tissue and the average transcriptome of the corresponding abnormal tissue are input into the trained generative transcriptome change process reconstruction model, N-step uniform sampling is performed, and N+2 transcriptome continuity dynamic data from the normal tissue transcriptome to the corresponding abnormal transcriptome are generated; N is a positive integer greater than 1, and the N+2 transcriptome continuity dynamic data constitute the average occurrence path of the given tumor.
[0059] As can be seen from the above, in one embodiment, the value of N can be 98, which can improve the accuracy of generating transcriptome continuity dynamic data, and thus the transcriptome continuous change process can be accurately reconstructed, and the key functional change events, event occurrence order and future transcriptome change trend in the tumor occurrence and development process can be accurately reconstructed.
[0060] 5. Individualized tumor occurrence mechanism deduction based on generative artificial intelligence model
[0061] The embodiment of the application flips 108 lung cancer patient paracancer normal tissue transcriptomes and their corresponding cancer tissue (non-normal tissue) transcriptomes into a trained model (generative transcriptome change process reconstruction model), and performs 98-step uniform sampling between the paracancer tissue transcriptome and the corresponding cancer tissue transcriptome of each patient to generate 100 transcriptomes. The 100 transcriptomes constitute the individualized tumorigenesis path of a given patient. That is, the input of the generative transcriptome change process reconstruction model is: a picture containing a normal tissue transcriptome and its corresponding non-normal tissue transcriptome, and the output of the model is: N+2 generated transcriptomes continuity dynamic data, that is, reconstructed transcriptome continuity dynamic data.
[0062] 6, Principal component analysis of average tumorigenesis path
[0063] The 3000 transcriptomes derived from 30 common tumor genesis pathways are integrated into an anndata dataset in the embodiment of the present application. Then principal component analysis is performed (as shown in FIG. 3A), and the gene list and weight corresponding to the principal component are exported. The gene list and weight are further used for gene function enrichment analysis and expression level regulation of genes in the tumor genesis process (as shown in FIG. 3B). In an embodiment, the individualized transcriptome evolution pathway is subjected to principal component analysis to obtain the key functional change events and event occurrence order in the tumor genesis and development process represented by the principal component. The key functional change events and event occurrence order in the tumor genesis and development process represented by the principal component can include: the individualized transcriptome evolution pathway is subjected to principal component analysis to obtain the gene list and weight corresponding to the principal component; the key functional change events and event occurrence order in the tumor genesis and development process represented by the principal component are indicated by the gene list and weight, that is, the gene list is used for gene function enrichment analysis, such as some tumors originating (occurring) from inflammation, some tumors originating from mutation of tumor suppressor genes, some tumors originating from dedifferentiation, and some tumors having all these events but different in the order of occurrence. In the embodiment of the present application, the key functional change events and event occurrence order in the tumor genesis and development process represented by the principal component are indicated by the gene list and weight obtained by principal component analysis, which can improve the accuracy and efficiency of reconstructing continuously changing transcriptomes. Specifically, trajectory analysis finds that the adrenal neuroblastoma genesis pathway intersects with the brain tumor genesis pathway, suggesting that neuroblastoma originates from glioblastoma like brain tumor. In clinical sample data, the characteristic proteins of neural precursor cells (SOX10, PLP1) that first increase and then decrease in the adrenal neuroblastoma genesis process have no significant regulation in the bulk-seq transcriptome of clinical samples, and NEUROD2 and NEUROD6 also have only a small increase (as shown in FIG. 3C). On the contrary, the neural precursor cells induced by neuroblastoma revealed by the transcriptome generation-based artificial intelligence model in the embodiment of the present application are highly similar to the Schwann cell precursors (SCP) that play an important role in tumor genesis discovered by single-cell sequencing of clinical samples (as shown in FIGS. 3D-3G; Dong et al., Cancer Cell 2020, Single-Cell Characterization of Malignant Phenotypes and Developmental Trajectories of Adrenal Neuroblastoma).Since Schwann cells are glial cells, the tumor occurrence path deduced by the generative artificial intelligence model (generative transcriptome change process reconstruction model) in the embodiments of the present application is close to the tumor occurrence mechanism based on existing single-cell sequencing and trajectory analysis, but the artificial intelligence model in the embodiments of the present application only needs bulk-seq transcriptome data, and does not need single-cell sequencing data, and the analysis efficiency in the embodiments of the present application is high.
[0064] In order to facilitate understanding of how the present application is implemented, 3A to FIG. 3G are described in detail as follows. FIG. 3A to FIG. 3G are schematic diagrams of generative artificial intelligence-based tumor occurrence mechanism deduction. FIG. 3A is the development trajectory of the transcriptome of thirty kinds of clinically common tumors from normal cancer-adjacent tissue to the corresponding cancer tissue in the principal component I (PC1) and the principal component II (PC2) space. Analysis shows that the occurrence path of neuroblastoma of adrenal gland intersects with that of multiple brain tumors (Glioblastoma Multiforme) in the middle, suggesting that the precursor cells of neuroblastoma and the precursor cells of brain tumors are similar to glial cells. FIG. 3B represents the transcriptome level changes of representative early neural development-related genes in the artificial intelligence deduced neuroblastoma occurrence path. NEUROD2, NEUROD6, PLP1, SOX10 and other genes all show expression regulation of first rising and then falling. In FIG. 3C, the bulk-seq clinical samples do not find that the expression levels of PLP1 and SOX10 are significantly different between neuroblastoma and normal adrenal cancer-adjacent tissue. NEUROD2 and NEUROD6 also only have a small increase. In FIG. 3D, the UMAP cell clustering of single-cell sequencing of embryos and neuroblastoma patient adrenal glands. In FIG. 3E, UMAP clustering analysis finds that neuroblastoma is mainly composed of three types of cells, Schwann cell precursor (SCP), sympathoblast, hematopoietic stem cell (HSC), and immune cell. Trajectory analysis further suggests that SCP is the precursor cell source of neuroblastoma. FIG. 3F to FIG. 3G are SOX10 (F) and PLP1 (G) which are mainly expressed in SCP and have very low expression levels in other cell types. Because Schwann cells are also glial cells, the neuroblastoma precursor cells found by the transcriptome-based generative artificial intelligence model in the embodiments of the present application are glial cells, which are highly consistent with the path analysis of existing single-cell sequencing, suggesting that the transcriptome reconstruction method based on the transcriptome generative artificial intelligence model in the embodiments of the present application has the ability to reproduce the tumor occurrence path. Simply comparing the bulk-seq transcriptome static data of tumor and cancer-adjacent tissue cannot find the intermediate steps of tumor occurrence and the precursor cells of tumor source.
[0065] From the above, the method for finding the universal law of tumorigenesis and evolution in the above step 4 based on the tumor mechanism deduction of the generative artificial intelligence model can further include:
[0066] Perform principal component analysis on the Mx100 transcriptomes deduced for M tumor occurrence paths to obtain a gene list and weights corresponding to the principal components; M is a positive integer greater than 1, and the gene list and weights are used for gene function enrichment analysis and expression level regulation of genes in the tumor occurrence process.
[0067] In specific implementation, the embodiment of the present application finds the universal law of tumorigenesis and evolution based on the reconstructed transcriptome series (transcriptome continuity dynamic data), such as the key functional change event and event occurrence order in the development process of a certain tumor and the transcriptome change trend after a preset period.
[0068] 7, Principal component analysis of lung cancer personalized occurrence path
[0069] The 10800 transcriptomes of 108 lung cancer personalized occurrence path deduction of 108 cases are integrated into an anndata dataset. Then principal component analysis is performed (as shown in FIGS. 4A-4B), and the gene list and weight corresponding to the principal component are exported, wherein the weight is used to sort the genes, remove the genes with small weight, retain the genes with large weight, and perform gene function enrichment analysis (for example, some tumors originate from inflammation, some tumors originate from mutation of tumor suppressor genes, some originate from dedifferentiation, and some tumors have all these events but the order of occurrence is different). In addition, the weight has positive and negative, the weight is positive, indicating that the gene is up-regulated in the principal component direction, the weight is negative, indicating that the gene is down-regulated in the principal component direction, and the absolute value of the weight represents the speed of up-regulation or down-regulation. When the gene list, i.e., the weight, is further used for gene function enrichment analysis and the development-related genes are found, especially the homeobox, keratin protein, and transport protein are highly enriched in PC1 (as shown in FIG. 4C). Keratin protein is a gene enriched in principal component I (PC1), and the method for reconstructing a transcriptome based on a transcriptome-generated artificial intelligence model in the embodiment of the present application suggests that it is up-regulated in the lung cancer occurrence process. The existing lung cancer and paracancer normal tissue bulk-seq transcriptome sequencing also confirms that the representative gene KRT20 of keratin protein is indeed significantly up-regulated in lung cancer, especially in lung squamous cell carcinoma samples (as shown in FIG. 4D). TLX2 and ELAVL3 are representative genes related to neural development, which are significantly up-regulated in the lung cancer occurrence process deduced by artificial intelligence (as shown in FIGS. 4E-4F), but only TLX2 regulation is verified in the transcriptome bulk-seq data of clinical samples (as shown in FIGS. 4G-4H), suggesting that the method for reconstructing a transcriptome based on a transcriptome-generated artificial intelligence model in the embodiment of the present application can find gene regulation that cannot be found by existing transcriptome bulk-seq sequencing, i.e., there are genes of lung cancer.
[0070] In order to facilitate understanding of how the present application is implemented, FIGS. 4A-4K are described in detail as follows.
[0071] FIG. 4A-4K are schematic diagrams of the reconstruction of individualized lung cancer pathogenesis in the embodiments of the present application. FIG. 4A is the reconstructed transcriptional evolution path from normal to lung cancer in 108 lung cancer patients based on the generative artificial intelligence model in the embodiments of the present application. The tangent line of the path indicates the direction of the subsequent evolution of the patient's transcriptome, i.e., the trend of the transcriptome changes. FIG. 4B shows the evolution paths of lung adenocarcinoma (LUAD) and lung squamous cell carcinoma (LUSC), respectively. Lung squamous cell carcinoma evolves along the direction of principal component I (PC1), while lung adenocarcinoma evolves along the diagonal direction of principal component I (PC1) and principal component II (PC2). FIG. 4C shows that gene functional enrichment analysis reveals that principal component I is mainly composed of development-related transcription factors, especially homeobox proteins, transport proteins, and keratin proteins. FIG. 4D shows the expression levels of the representative keratin protein gene KRT20 in adenocarcinoma, squamous cell carcinoma, and normal tissues adjacent to the cancer (log2TPM+1, data source: TCGA). FIG. 4E-4F show the induction of representative neurodevelopment-related genes TLX2 (E) and ELAVL3 (F) during lung cancer development based on the generative artificial intelligence model in the embodiments of the present application. FIG. 4G-4H show the expression levels of neurodevelopment-related genes TLX2 (G) and ELAVL3 (H) in lung cancer and normal tissues adjacent to the cancer, suggesting that conventional transcriptome sequencing cannot find the induced expression of ELAVL3 in lung cancer. FIG. 4I-4J show that existing single-cell sequencing finds that TLX2 (I) and ELAVL3 (J) are only expressed in neuroendocrine-like cells (NE) specific to lung cancer in lung cancer clinical samples, confirming the correctness of the reconstruction of the transcriptome based on the artificial intelligence model in the embodiments of the present application. UMAP in FIG. 4I-4J is the abbreviation of uniform manifold approximation and projection. The small molecule inhibitor of ELAVL3 in FIG. 4K can effectively inhibit tumor cell growth in nude mice.
[0072] Step 4 above is the reconstruction of the tumor pathogenesis mechanism based on the generative artificial intelligence model, and step 6 is the principal component analysis of the average tumor pathogenesis path, which is the step of discovering the universal law of tumor pathogenesis and evolution.
[0073] 8. Analysis of lung cancer single-cell sequencing data sets
[0074] The embodiment of the present application adopts a non-small cell lung cancer single cell sequencing dataset. The dataset contains 892,296 cells with detailed annotations. The dataset analysis and display are shown in FIGS. 4I-4J. The analysis shows that the artificial intelligence deductive discovery that the induced neurodevelopment-related genes TLX2 and ELAVL3 in lung cancer are both highly expressed in tumor sample-specific neuroendocrine cell groups (NE). Therefore, the lung cancer occurrence path deduced by the artificial intelligence is again supported by the single cell sequencing results.
[0075] 9. Tumor cell nude mouse tumorigenicity analysis
[0076] The embodiment of the present application constructs a human tumor cell line DU145 stably expressing ELVAL3 protein. A four-week-old mouse is used for tumor load experiment. Each mouse is subcutaneously injected with 5x10 6 cells in the right lower axillary fossa. After three weeks, the experimental group and the control group are injected with pyrvinium pamoate (0.2 mg / kg) or PBS, respectively, four times a week for five consecutive weeks, the mice are killed, the tumors are removed, and the body weight is measured. The analysis results show that the ELAVL3 small molecule inhibitor can suppress tumor growth (as shown in FIG. 4K). Therefore, the tumor occurrence path deduced based on the generative artificial intelligence model in the embodiment of the present application helps to discover anti-tumor targets.
[0077] The above steps 8 and 9 are steps for verifying the effectiveness of the method for reconstructing continuously changing transcriptomes based on picture processing provided by the embodiment of the present application compared with the prior art, that is, compared with the prior art, the method for reconstructing continuously changing transcriptomes based on picture processing provided by the embodiment of the present application accurately reconstructs the transcriptome continuity dynamic data based on transcriptome pictures, accurately reconstructs the key functional change events, event occurrence order, and future transcriptome change trend in the tumor occurrence and development process.
[0078] In summary, compared with the prior art, the method for reconstructing continuously changing transcriptomes based on picture processing provided by the embodiment of the present application has the following beneficial effects:
[0079] 1. The present application can infer the occurrence process of tumors. Existing omics data are static, that is, only static data of patient collected pathological specimens. The present application can obtain continuity data (such as the transcriptome change process in the lung adenocarcinoma and squamous cell carcinoma occurrence and development process deduced by the artificial intelligence model) by sampling the generative artificial intelligence model space.
[0080] 2. The continuity data makes it possible to deduce the mechanism of disease occurrence and subsequent evolution rule (see FIG. 3A-FIG. 3B, the transcriptional evolution path of adrenal neuroblastoma generated by the artificial intelligence model according to the embodiment of the present application (FIG. 3A) and the process of adrenal neuroblastoma occurrence and carcinogenesis (FIG. 3B)). The artificial intelligence model suggests that adrenal neuroblastoma originates from Schwann cell precursor with high expression of NeuroG2 and NeuroD2. This finding is confirmed by the latest single-cell transcriptome.
[0081] 3. Thus, more effective tumor treatment targets can be further found, and the best treatment plan based on the personalized evolution mechanism can be provided for each tumor patient, that is, the transcriptional evolution path can be used to recommend the best treatment plan for individuals.
[0082] The embodiment of the present application also provides a device for reconstructing continuously changing transcriptome based on picture processing, as described in the following embodiment. Since the principle of solving the problem of the device is similar to that of the method for reconstructing continuously changing transcriptome based on picture processing, the implementation of the device can be referred to the implementation of the method for reconstructing continuously changing transcriptome based on picture processing, and the repeated parts will not be described here.
[0083] FIG. 5 is a structural schematic diagram of the device for reconstructing continuously changing transcriptome based on picture processing in the embodiment of the present application, as shown in FIG. 5, the device comprises:
[0084] The acquisition unit 01 is configured to acquire pictures of the normal tissue transcriptome to be reconstructed and the corresponding non-normal tissue transcriptome;
[0085] The reconstruction unit 02 is configured to input the pictures of the normal tissue transcriptome to be reconstructed and the corresponding non-normal tissue transcriptome into a pre-trained generative transcriptome change process reconstruction model, perform N-step uniform sampling between the normal tissue transcriptome and the corresponding non-normal tissue transcriptome, and generate N+2 transcriptome continuity dynamic data; N is a positive integer greater than 1, and the N+2 transcriptome continuity dynamic data constitutes an individualized transcriptome evolution path. The generative transcriptome change process reconstruction model is obtained by pre-training a generative network using a transcriptome picture sample data set;
[0086] The processing unit 03 is configured to perform principal component analysis on the individualized transcriptome evolution path to obtain a gene list and a weight corresponding to a principal component; the gene list and the weight are used to indicate a key functional change event and an event occurrence order in a tumor occurrence and development process represented by the principal component, and a tangent direction of the individualized transcriptome evolution path is used to indicate a transcriptome change trend that will occur after a preset period of the tumor.
[0087] In an embodiment, the apparatus for reconstructing a continuously changing transcriptome based on picture processing can further comprise a training unit configured to pre-train the generative transcriptome change process reconstruction model according to the following method:
[0088] generating transcriptome picture data;
[0089] obtaining a transcriptome picture sample data set; the data set comprises positive sample data and negative sample data, the positive sample data being tumor tissue transcriptome picture data, and the negative sample data being normal tissue transcriptome picture data;
[0090] training a generative network using the transcriptome picture sample data set to obtain the generative transcriptome change process reconstruction model.
[0091] In an embodiment, when generating the transcriptome picture, each gene is fixed at a corresponding position in a 512x512 coordinate system, the expression amount of the gene corresponds to a 512x512 array pixel value, and the conversion mode of the expression amount to the pixel intensity is log(FPKM+1)x14, with a value range of 0-255.
[0092] In an embodiment, the processing unit is specifically configured to:
[0093] perform principal component analysis on the individualized transcriptome evolution path to obtain a gene list and weights corresponding to principal components;
[0094] the gene list and the weights indicate key functional change events and event occurrence sequences in the tumor occurrence and development process represented by the principal components.
[0095] In an embodiment, the apparatus for reconstructing a continuously changing transcriptome based on picture processing can further comprise:
[0096] a mean transcriptome determination unit of non-normal tissue configured to take the average of the expression amount of each gene in all samples of the non-normal tissue transcriptome of all individuals of each type of tumor as the mean transcriptome of the non-normal tissue;
[0097] a mean transcriptome determination unit of normal tissue configured to combine the normal tissue transcriptome of all non-normal individuals and the transcriptome of the corresponding tissue of all normal individuals in the transcriptome data set, and take the average of the expression amount of each gene as the mean transcriptome of the normal tissue;
[0098] An average occurrence path determination unit is configured to input the normal tissue mean transcriptome and the corresponding non-normal tissue mean transcriptome into a trained generative transcriptome change process reconstruction model, perform N-step uniform order sampling, and generate N+2 transcriptome continuity dynamic data from the normal tissue transcriptome to the corresponding non-normal transcriptome; N is a positive integer greater than 1, and the N+2 transcriptome continuity dynamic data constitutes the average occurrence path of a given tumor.
[0099] An analysis unit is configured to perform principal component analysis on the Mx100 transcriptomes deduced by the M tumor occurrence paths, and obtain a gene list and weights corresponding to the principal components; M is a positive integer greater than 1, and the gene list and the weights are used for gene function enrichment analysis and expression level regulation of genes in the tumor occurrence process.
[0100] In one embodiment, the generative transcriptome change process reconstruction model is a model obtained based on a trained generative network, and preferably, the generative network is a generative StyleGAN network, and the value of N is 98.
[0101] The embodiment of the present application also provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the above-mentioned picture processing-based reconstruction of continuously changing transcriptome method when executing the computer program.
[0102] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the above-mentioned picture processing-based reconstruction of continuously changing transcriptome method.
[0103] The embodiment of the present application also provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the above-mentioned picture processing-based reconstruction of continuously changing transcriptome method.
[0104] Compared with the scheme in the prior art that recognizes the static difference between the non-normal tissue transcriptome and the normal tissue transcriptome by collecting pathological specimens of tumor patients, the scheme for reconstructing the continuously changing transcriptome based on picture processing provided in the embodiment of the present application inputs the picture of the normal tissue transcriptome to be recognized and the picture of the corresponding non-normal tissue transcriptome into a pre-trained generative transcriptome change process reconstruction model, performs N-step uniform sampling between the normal tissue transcriptome and the corresponding non-normal tissue transcriptome, generates N+2 transcriptome continuity dynamic data, and equivalently reconstructs the transcriptome continuity dynamic data; N is a positive integer greater than 1, and the N+2 transcriptome continuity dynamic data constitute an individualized transcriptome evolution path; principal component analysis is performed on the individualized transcriptome evolution path, and the key functional change events and the event occurrence sequence in the tumor occurrence and development process represented by the principal components are obtained; the tangent direction of the individualized transcriptome evolution path is used to indicate the transcriptome change trend that will occur after a preset period of the tumor, and the transcriptome continuity dynamic data reconstructed based on the transcriptome picture is used to accurately reconstruct the transcriptome continuous change process, and further accurately reconstruct the key functional change events, the event occurrence sequence in the tumor occurrence and development process, and the transcriptome change trend that will occur in the future.
[0105] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0106] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce the apparatus for implementing the functions specified in one or more flows in the flowcharts and / or one or more blocks in the block diagrams.
[0107] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks and / or flowchart flow or flows and / or block or blocks of the block diagram.
[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks and / or flowchart flow or flows and / or block or blocks of the block diagram.
[0109] The above specific embodiments are described to further explain the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A method of reconstructing a continuously varying transcriptome based on image processing, characterized in that, The method comprises the following steps: Obtaining pictures of normal tissue transcriptome and corresponding non-normal tissue transcriptome to be reconstructed; Inputting the pictures of normal tissue transcriptome and corresponding non-normal tissue transcriptome to be reconstructed into a pre-trained generative transcriptome change process reconstruction model, performing N-step uniform sampling between the normal tissue transcriptome and the corresponding non-normal tissue transcriptome, and generating N+2 transcriptome continuity dynamic data; N is a positive integer greater than 1, and the N+2 transcriptome continuity dynamic data constitutes an individualized transcriptome evolution path, and the generative transcriptome change process reconstruction model is obtained by pre-training a generative network using a transcriptome picture sample data set; Performing principal component analysis on the individualized transcriptome evolution path to obtain key functional change events and event occurrence sequences in the tumor occurrence and development process represented by principal components, and a tangent line direction of the transcriptome evolution path is used to indicate a transcriptome change trend to be occurred after a preset period.
2. The method of claim 1, wherein, The method further comprises pre-training the generative transcriptome change process reconstruction model in the following manner: Generating transcriptome picture data; Obtaining a transcriptome picture sample data set; the data set comprises positive sample data and negative sample data, the positive sample data is non-normal tissue transcriptome picture data, and the negative sample data is normal tissue transcriptome picture data; Training a generative network using the transcriptome picture sample data set to obtain the generative transcriptome change process reconstruction model.
3. The method of claim 2, wherein, When generating the transcriptome picture, each gene is fixed on a corresponding position of a 512*512 coordinate system, the expression amount of the gene corresponds to a 512*512 array pixel value, and the conversion mode from the expression amount to the pixel emphasis is log(FPKM+1)*14, and the value range is between 0 and 255.
4. The method of claim 1, wherein, Performing principal component analysis on the individualized transcriptome evolution path to obtain key functional change events and event occurrence sequences in the tumor occurrence and development process represented by principal components, comprising: Performing principal component analysis on the individualized transcriptome evolution path to obtain a gene list and weights corresponding to the principal components; Indicating the key functional change events and event occurrence sequences in the tumor occurrence and development process represented by the principal components through the gene list and the weights.
5. The method of claim 1, wherein, The method further comprises: Taking the average of the expression amounts of each gene of the non-normal tissue transcriptome of all individuals of each tumor in all samples as a non-normal tissue average transcriptome; Combining the normal tissue transcriptomes of all non-normal individuals and the transcriptomes of the corresponding tissues of all normal individuals in the transcriptome data set, and taking the average of the expression amounts of each gene as a normal tissue average transcriptome; Inputting the normal tissue average transcriptome and the corresponding non-normal tissue average transcriptome into the trained generative transcriptome change process reconstruction model, performing N-step uniform sampling, and generating N+2 transcriptome continuity dynamic data from the normal tissue transcriptome to the corresponding non-normal transcriptome; N is a positive integer greater than 1, and the N+2 transcriptome continuity dynamic data constitutes an average occurrence path of a given tumor. The Mx100 transcriptomes deduced from M tumor occurrence paths are subjected to principal component analysis to obtain a gene list and weights corresponding to the principal components; M is a positive integer greater than 1, and the gene list and weights are used for gene function enrichment analysis and expression level regulation of genes in the tumor occurrence process.
6. The method of claim 1, wherein, The generative transcriptome change process reconstruction model is a model obtained based on a trained generative network, and N is 98.
7. An apparatus for reconstructing continuously varying transcriptomes based on image processing, characterized in that, The method comprises the following steps: An acquisition unit is configured to acquire a normal tissue transcriptome to be reconstructed and a picture of a corresponding non-normal tissue transcriptome; A reconstruction unit is configured to input the normal tissue transcriptome to be reconstructed and the picture of the corresponding non-normal tissue transcriptome into a pre-trained generative transcriptome change process reconstruction model, perform N-step uniform sampling between the normal tissue transcriptome and the corresponding non-normal tissue transcriptome, and generate N+2 transcriptome continuity dynamic data; N is a positive integer greater than 1, and the N+2 transcriptome continuity dynamic data constitute an individualized transcriptome evolution path; the generative transcriptome change process reconstruction model is obtained by pre-training a generative network using a transcriptome picture sample data set; A processing unit is configured to perform principal component analysis on the individualized transcriptome evolution path to obtain a gene list and weights corresponding to the principal components; the gene list and weights are used to indicate key functional change events and event occurrence sequences in a tumor occurrence and development process represented by the principal components, and a tangent line of the transcriptome evolution path is used to indicate a transcriptome change trend that will occur after a preset time period.
8. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method of any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1 to 6.
10. A computer program product, characterised in that, The computer program product comprises a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Cancer transcriptome data processing method based on gene co-expression network analysis
CN114360642A
Transcriptome image generation device, method and application
CN114882955A
Method and system for identifying spatial transcriptome cell expression pattern
CN115732034A
Tumor virtual three-dimensional transcriptome construction method and device, equipment and storage medium
CN116129999A
Method, apparatus and computer program product for providing a multi-omics framework for estimating temporal disease trajectories
US20220399120A1