Gene module construction method and device, computer equipment and storage medium

By deconvolution and similarity clustering of single-cell and spatiotemporal transcriptome data, combined with the HotSpot model to construct gene supermodules, the problems of poor compatibility of multimodal data and batch effects were solved, and the accuracy and consistency of gene module construction were improved.

CN120613009APending Publication Date: 2025-09-09SHENZHEN HUADA SANJIAN QIFA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410259476.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies have poor compatibility with multimodal data and are unable to effectively address the impact of batch effects, resulting in low accuracy in gene module construction.

Method used

By sequencing single-cell and spatiotemporal transcriptome data, constructing a transcriptome gene expression matrix, combining it with a multi-layer perceptron model for deconvolution processing, and using the HotSpot model to construct gene modules, similarity clustering and integration were performed to obtain gene supermodules.

Benefits of technology

It improves the compatibility of multimodal data, reduces batch effects, and improves the accuracy and consistency of gene module construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120613009A_ABST
    Figure CN120613009A_ABST
Patent Text Reader

Abstract

The invention relates to a gene module construction method and device, computer equipment and a storage medium. The method comprises the steps that gene sequencing is conducted according to pathological sample queue data, a transcription data set is obtained, the transcription data set comprises single cell data and space-time transcriptome data, transcriptome preprocessing is conducted on the space-time transcriptome data, and a transcriptome gene expression matrix is obtained; carrying out deconvolution processing according to the type of the single cell data and the transcriptome gene expression matrix to obtain embedded matrix data, carrying out gene module construction according to the embedded matrix data and a preset target gene construction model to obtain a gene module of each sample slice, carrying out similarity clustering on the gene modules to obtain a gene weight list, and storing the gene weight list in a database; and integrating the gene modules according to the gene weight list to obtain a gene hypermodule. Therefore, by adopting the method, the compatibility of the multi-modal data can be improved, the batch effect can be reduced, and the gene module construction accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data analysis technology, and in particular to a gene module construction method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the development of spatiotemporal transcriptomics technology, large-scale analysis of transcriptomes with single-cell resolution can comprehensively classify cell types and states in different tissues. The emergence of single-cell resolution spatial omics provides precise spatial location information of cells, which is suitable for the analysis of spatiotemporal transcriptome gene data and the mining of molecular features with spatiotemporal specificity from the data, thereby improving the accuracy of gene module construction.

[0003] However, there is currently no comprehensive solution for constructing gene modules that integrate single-cell and spatial groups. The more mainstream solution is to calculate the Jaccard similarity between non-negative matrix decompositions (NMF) based on the results of spatial group or single-cell expression matrices to merge samples or discover gene programs across cell types. However, this merging method has poor compatibility with multimodal data, and when discussing the convergent characteristics of multiple samples, it is difficult to use a low-dimensional representation to describe the similarities beyond batch effects. In other words, it is difficult to effectively address the impact of batch effects, resulting in low accuracy in gene module construction. Therefore, how to achieve a gene module construction method that improves the compatibility of multimodal data and reduces batch effects has become an urgent problem that needs to be solved. Summary of the Invention

[0004] Based on this, it is necessary to provide a gene module construction method, device, computer equipment, and storage medium that can improve the compatibility of multimodal data and reduce batch effects in response to the above technical problems.

[0005] In a first aspect, the present application provides a method for constructing a gene module, comprising:

[0006] Gene sequencing is performed based on the pathological sample cohort data to obtain a transcriptional dataset; wherein the transcriptional dataset includes: single-cell data and spatiotemporal transcriptome data;

[0007] Performing transcriptome preprocessing on the spatiotemporal transcriptome data to obtain a transcriptome gene expression matrix;

[0008] Performing deconvolution processing according to the type of the single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data;

[0009] Constructing a gene module based on the embedded matrix data and a preset target gene construction model to obtain a gene module for each sample slice;

[0010] Performing similarity clustering on the gene modules to obtain a gene weight list;

[0011] The gene modules are integrated according to the gene weight list to obtain a gene supermodule.

[0012] In one embodiment, the type includes single cell type data, and the method further comprises:

[0013] Performing single-cell preprocessing clustering on the single-cell data to obtain a clustering result;

[0014] The single cell data is annotated according to the clustering result to obtain the single cell type data.

[0015] In one embodiment, performing single-cell pre-processing clustering on the single-cell data to obtain a clustering result includes:

[0016] Performing gene sequence comparison on the single cell data and preset genome data to obtain comparison data;

[0017] Performing data preprocessing on the comparison data according to a preset preprocessing strategy to obtain a gene count matrix;

[0018] Cluster analysis is performed on the gene count matrix according to a preset cluster analysis strategy to obtain the clustering result.

[0019] In one embodiment, the type of the single-cell data includes single-cell type data; and performing deconvolution processing based on the type of the single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data includes:

[0020] Deconvolving the single cell type data to obtain cell subclass label data;

[0021] The cell subclass label data is embedded into the transcriptome gene expression matrix according to a preset multi-layer perceptron model to obtain the embedded matrix data.

[0022] In one embodiment, integrating the gene modules according to the gene weight list to obtain a gene supermodule includes:

[0023] Calculating the Jaccard similarity between the gene weight preset rankings of each gene module according to the gene weight list and the gene module to obtain a Jaccard similarity matrix;

[0024] Performing hierarchical clustering on the Jaccard similarity matrix to obtain a hierarchical clustering result;

[0025] Performing super-module gene screening based on the hierarchical clustering results and the gene module to obtain super-module genes;

[0026] The supermodule genes are integrated to obtain the gene supermodule.

[0027] In one embodiment, the super-module gene screening is performed based on the hierarchical clustering results and the gene module to obtain the super-module gene, including:

[0028] Performing module similarity judgment on the gene modules according to the hierarchical clustering results to obtain inter-module similarity data;

[0029] Performing multi-category clustering on the gene modules according to the inter-module similarity data to obtain multi-category clustering data;

[0030] Counting the gene frequencies in each category according to the multi-category clustering data to obtain the gene frequencies of each category;

[0031] The gene modules are screened according to the frequencies of the genes in each category to obtain the supermodule genes.

[0032] In one embodiment, after performing hierarchical clustering on the Jaccard similarity matrix to obtain a hierarchical clustering result, the method further includes:

[0033] Perform heat map drawing according to the hierarchical clustering result and a preset heat map drawing function to obtain a correlation heat map;

[0034] Calculating the cellular level of each gene module according to a preset scoring model and the correlation heat map to obtain cellular level data;

[0035] Drawing is performed according to the cell-level data and the heat map drawing function to obtain a spatial distribution pseudo-color map, and the spatial distribution pseudo-color map is displayed.

[0036] In a second aspect, the present application also provides a gene module construction device, comprising:

[0037] A sequencing module is used to perform gene sequencing based on the pathological sample cohort data to obtain a transcriptional dataset; wherein the transcriptional dataset includes: single-cell data and spatiotemporal transcriptome data;

[0038] A preprocessing module, configured to perform transcriptome preprocessing on the spatiotemporal transcriptome data to obtain a transcriptome gene expression matrix;

[0039] a deconvolution module, configured to perform deconvolution processing according to the type of the single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data;

[0040] A construction module is used to construct a gene module according to the embedded matrix data and a preset target gene construction model to obtain a gene module for each sample slice;

[0041] A similarity clustering module is used to perform similarity clustering on the gene modules to obtain a gene weight list;

[0042] An integration module is used to integrate the gene modules according to the gene weight list to obtain a gene supermodule.

[0043] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0044] Gene sequencing is performed based on the pathological sample cohort data to obtain a transcriptional dataset; wherein the transcriptional dataset includes: single-cell data and spatiotemporal transcriptome data;

[0045] Performing transcriptome preprocessing on the spatiotemporal transcriptome data to obtain a transcriptome gene expression matrix;

[0046] Performing deconvolution processing according to the type of the single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data;

[0047] Constructing a gene module based on the embedded matrix data and a preset target gene construction model to obtain a gene module for each sample slice;

[0048] Performing similarity clustering on the gene modules to obtain a gene weight list;

[0049] The gene modules are integrated according to the gene weight list to obtain a gene supermodule.

[0050] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0051] Gene sequencing is performed based on the pathological sample cohort data to obtain a transcriptional dataset; wherein the transcriptional dataset includes: single-cell data and spatiotemporal transcriptome data;

[0052] Performing transcriptome preprocessing on the spatiotemporal transcriptome data to obtain a transcriptome gene expression matrix;

[0053] Performing deconvolution processing according to the type of the single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data;

[0054] Constructing a gene module based on the embedded matrix data and a preset target gene construction model to obtain a gene module for each sample slice;

[0055] Performing similarity clustering on the gene modules to obtain a gene weight list;

[0056] The gene modules are integrated according to the gene weight list to obtain a gene supermodule.

[0057] The above-mentioned gene module construction method, device, computer equipment, and storage medium first perform gene sequencing based on the pathological sample cohort data to obtain a transcription data set, wherein the transcription data set includes: single-cell data and spatiotemporal transcriptome data, then perform transcriptome preprocessing on the spatiotemporal transcriptome data to obtain a transcriptome gene expression matrix, and then perform deconvolution processing based on the type of single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data that combines the characteristics of the two data. Then, gene modules are constructed based on the embedded matrix data and the preset target gene construction model to obtain the gene module of each sample slice, and the gene modules are similarly clustered to obtain a gene weight list. Finally, the gene modules are integrated according to the gene weight list to obtain a gene super module. Therefore, the compatibility of multimodal data is improved by combining the characteristics of multimodal data, and the batch effect caused by a large number of sample slices is reduced by using similarity clustering and weight integration for the gene modules corresponding to a large number of sample slices. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0059] Figure 1 A diagram of an application environment of a gene module construction method in one embodiment;

[0060] Figure 2 Schematic diagram of a process for constructing a gene module in one embodiment;

[0061] Figure 3 A diagram of pathological sample cohort data in one embodiment;

[0062] Figure 4 Schematic diagram of the process of annotating single cell type data in one embodiment;

[0063] Figure 5 A data diagram of a marker gene list of a supermodule in one embodiment;

[0064] Figure 6 A flowchart of a visual image display step in one embodiment;

[0065] Figure 7 A grayscale image of a pseudo-color image of the spatial distribution of a gene module on a slice in one embodiment;

[0066] Figure 8 This is a prognostic survival analysis diagram of the TCGA-PAAD dataset of MGM6 in another embodiment;

[0067] Figure 9 This is a prognostic survival analysis diagram of the TCGA-PAAD dataset of MGM11 in another embodiment;

[0068] Figure 10 A structural block diagram of a gene module construction device in one embodiment;

[0069] Figure 11 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0070] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0071] The gene module construction method provided in the embodiments of the present application can be applied to Figure 1 In the application environment shown. The data storage system obtains pathological sample queue data for gene sequencing to obtain a transcription data set, wherein the transcription data set includes: single cell data and spatiotemporal transcriptome data, and the following steps are implemented by server 104: deconvolution processing is performed according to the type of single cell data and the transcriptome gene expression matrix to obtain embedded matrix data, gene module construction is performed according to the embedded matrix data and the preset target gene construction model, the gene module of each sample slice is obtained, the gene module is clustered similarly to obtain a gene weight list, the gene module is integrated according to the gene weight list to obtain a gene super module, and finally the obtained gene super module is transmitted to the terminal 102 via the network. Among them, the data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, etc. The server 104 can be implemented with an independent server or a server cluster consisting of multiple servers.

[0072] In an exemplary embodiment, Figure 2 As shown, a gene module construction method is provided, and the method is applied to Figure 1The server 104 in the example is used to illustrate the process, including the following steps 202 to 212.

[0073] Step 202 : Perform gene sequencing based on the pathology sample cohort data to obtain a transcriptional dataset; wherein the transcriptional dataset includes: single-cell data and spatiotemporal transcriptome data.

[0074] The pathological sample cohort data may include fresh tumor tissue and adjacent tissue samples, as well as other data, which can be referenced. Figure 3 .

[0075] In some embodiments, pathology sample cohort data includes samples of primary pancreatic cancer patients and their matched fresh tumor tissue and adjacent tumor tissue. For example, matched frozen tumor tissue and adjacent tumor tissue are collected from pancreatic cancer patients and cohorted to obtain pathology sample cohort data. The frozen tumor tissue can be at least 2 cm from the tumor boundary, and the adjacent tumor tissue can be at least 2 cm from the tumor boundary.

[0076] In some embodiments, single-cell transcriptome sequencing is performed by surgically removing pancreatic marginal zone tissue from liver cancer patients, dissociating it, and preparing a single-cell suspension. scRNA-seq libraries are prepared using the DNBelab C4 system. The sequencing libraries are sequenced using the DIPSEQ T10 sequencer to generate single-cell data. The marginal zone tissue can be a 1 cm wide region centered on the tumor boundary.

[0077] In some embodiments, spatiotemporal transcriptome sequencing is performed by collecting tumor tissue, attaching the edge region to the surface of a Stereo-seq chip, and staining adjacent tissue sections attached to a slide using H&E staining. Stereo-seq is then used for spatiotemporal transcriptome sequencing. Briefly, to generate a DNB array for in situ RNA capture, oligonucleotides containing random 25-nucleotide coordinate identities are first synthesized and circularized using T4 DNA ligase and a splint. DNBs are then generated by rolling circle amplification and loaded onto the patterned chip. Next, single-end sequencing is performed on a DNBSEQ-T10 sequencer using the SE25 sequencing strategy to generate spatiotemporal transcriptome data. The tumor tissue can be a 1×1 cm tissue region at least 2 cm from the tumor boundary; the edge region can be a 1×1 cm tissue region centered on the tumor boundary; and the patterned chip can be 10 mm × 10 mm in size. For simplicity, in the following examples, nucleotide coordinate identities are referred to as CIDs, and the spatiotemporal omics technology is referred to as Stereo-seq.

[0078] Step 204 : performing transcriptome preprocessing on the spatiotemporal transcriptome data to obtain a transcriptome gene expression matrix.

[0079] In some embodiments, transcriptome preprocessing is performed on spatiotemporal transcriptome data, and the Stereo-seq preprocessing process SAW can be used to obtain a transcriptome gene expression matrix with cellular level resolution, wherein the transcriptome gene expression matrix contains spatial location information.

[0080] Specifically, FASTQ files were aligned to the reference genome (GRCh38) using STAR. Mapped reads with a MAPQ > 10 were counted and annotated to their corresponding genes. UMIs with the same CID and identical locus were collapsed, allowing for one mismatch to correct for sequencing and PCR errors. This information was used to generate an expression matrix containing CIDs. Small fields of view of nucleic acid staining microscopy images were then stitched together based on image overlap and the microarray grid lines. The microarray expression matrix was then plotted as a grayscale image based on the UMI number at each point. The microarray grid lines in the grayscale image were overlaid with the staining image and manually aligned. The aligned staining image was then exported. Cells were then segmented using a U-Net-based deep neural network, and the microarray expression matrix was cell-corrected using the GMM method. A transcriptome gene expression matrix with cell labels and spatial location information was generated by cell binning. A cell bin can be the region of a single cell after cell segmentation. For simplicity, the segmented single cell region will be referred to as a cell bin in the subsequent examples.

[0081] Step 206 : Perform deconvolution processing based on the type of single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data.

[0082] In some embodiments, the type of single-cell data includes single-cell type data, and the single-cell data deconvolution method can be used to annotate the transcriptome gene expression matrix of each cellbin with spatial group CAF cells to obtain embedded matrix data of different cell types in each cellbin.

[0083] Step 208 : constructing a gene module based on the embedded matrix data and a preset target gene construction model to obtain a gene module for each sample slice.

[0084] In some embodiments, the preset target gene construction model can be HotSpot, which constructs the gene module of each sample slice by taking the embedding matrix data as the input of HotSpot and adopting the embedding-based similarity graph method proposed by HotSpot.

[0085] Step 210 , perform similarity clustering on the gene modules to obtain a gene weight list.

[0086] In some embodiments, HotSpot is used to perform spatial cellbin similarity calculation on the gene module to obtain a similarity calculation result, and then the create_knn_graph function is used to construct a knn graph based on the similarity calculation result, and then the compute_autocorrelations() function is used to calculate the weight of the gene in the knn graph to obtain a gene weight list.

[0087] Step 212: Integrate the gene modules according to the gene weight list to obtain a gene supermodule.

[0088] In some embodiments, a gene weight list can be used to determine a preset number of genes with the highest frequency of occurrence in a gene module, and the preset number of genes can be used as a gene supermodule, thereby improving the accuracy of gene module construction.

[0089] In the above-mentioned gene module construction method, gene sequencing is first performed based on the pathological sample cohort data to obtain a transcriptional dataset, wherein the transcriptional dataset includes: single-cell data and spatiotemporal transcriptome data, and then the spatiotemporal transcriptome data is subjected to transcriptome preprocessing to obtain a transcriptome gene expression matrix, and then deconvolution processing is performed based on the type of single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data that combines the characteristics of the two data, and then gene modules are constructed based on the embedded matrix data and the preset target gene construction model to obtain the gene module of each sample slice, and the gene modules are clustered by similarity to obtain a gene weight list, and finally the gene modules are integrated according to the gene weight list to obtain a gene super module. Therefore, the compatibility of multimodal data is improved by combining the characteristics of multimodal data, and the batch effect caused by a large number of sample slices is reduced by using similarity clustering and weight integration for the gene modules corresponding to a large number of sample slices.

[0090] In some embodiments, the type of single-cell data includes single-cell type data. Before step 203, the gene module construction method further includes: a single-cell type data annotation step.

[0091] In an exemplary embodiment, Figure 4 As shown, the single cell type data annotation step may include but is not limited to the following steps 402 to 404. Among them:

[0092] Step 402: Perform single-cell pre-processing clustering on the single-cell data to obtain a clustering result.

[0093] Step 404 : annotate the single cell data according to the clustering result to obtain single cell type data.

[0094] In some embodiments, single-cell preprocessing and clustering are performed on single-cell data to obtain clustering results, including: performing gene sequence comparison between the single-cell data and preset genome data to obtain comparison data; performing data preprocessing on the comparison data according to a preset preprocessing strategy to obtain a gene count matrix; and performing cluster analysis on the gene count matrix according to a preset cluster analysis strategy to obtain clustering results.

[0095] In some embodiments, the preset genomic data uses the GRCh38 human reference genome. The preset preprocessing strategies include: filtering alignment, gene annotation, UMI correction, and UMI count matrix generation. The preset clustering analysis strategies include: quality filtering, sample integration, normalization, scaling and dimensionality reduction visualization, and clustering. The GRCh38 human reference genome can be Ensembl 98.

[0096] Specifically, single-cell FASTQ files are first aligned to Ensembl 98 to generate alignment data. After adjusting the alignment quality score, filtering and alignment, gene annotation, UMI correction, and UMI count matrix generation are performed to generate a gene count matrix. The gene count matrices of all samples are then merged, followed by quality filtering, sample integration, normalization, scaling, dimensionality reduction, visualization, and clustering to obtain accurate clustering results. Clustering results can be more intuitively viewed using visual graphs. MAPQ can be used to assess alignment quality.

[0097] In the above embodiment, the single-cell data is first compared with the preset genome data for gene sequence to obtain comparison data, and then the comparison data is preprocessed according to the preset preprocessing strategy to obtain a gene count matrix. Finally, the gene count matrix is ​​clustered according to the preset clustering analysis strategy to obtain accurate clustering results, which can effectively improve the accuracy of the clustering results, thereby improving the accuracy of the type annotation of the single-cell data in subsequent steps.

[0098] In some embodiments, deconvolution processing is performed based on the type of single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data, including: deconvolution of single-cell type data to obtain cell subclass label data; and embedding the cell subclass label data into the transcriptome gene expression matrix according to a preset multi-layer perceptron model to obtain embedded matrix data.

[0099] Specifically, the pre-set multilayer perceptron model can use SPACEL's Spoint module, which is an embedded multilayer perceptron model with a probabilistic model. Single-cell type data is deconvolved to obtain cell subtype label data. This cell subtype label data is then convolved with the transcriptome gene expression matrix using the Spoint module to obtain an embedding matrix for each cell bin's different cell types. The multilayer perceptron model can be abbreviated as an MLP model.

[0100] In the above embodiment, by deconvolving the single cell type data, the cell subclass label data is obtained, and the cell subclass label data is embedded in the transcriptome gene expression matrix according to the preset multi-layer perceptron model to obtain embedded matrix data, which can improve the compatibility of multimodal data, so that the constructed embedded matrix data can effectively reflect the gene characteristics of single cell data and transcriptome data, thereby improving the accuracy of subsequent gene module construction.

[0101] In some embodiments, gene modules are integrated according to a gene weight list to obtain a gene supermodule, including: calculating the Jaccard similarity between the preset rankings of gene weights of each gene module according to the gene weight list and the gene module to obtain a Jaccard similarity matrix; performing hierarchical clustering on the Jaccard similarity matrix to obtain a hierarchical clustering result; performing supermodule gene screening according to the hierarchical clustering result and the gene module to obtain a supermodule gene; and integrating the supermodule genes to obtain a gene supermodule.

[0102] Specifically, we first calculated the Jaccard similarity between the top 50 genes in each gene module to obtain a Jaccard similarity matrix. We then performed Ward hierarchical clustering on the Jaccard similarity matrix to obtain the hierarchical clustering results. Based on the hierarchical clustering results and the gene modules, we screened for supermodule genes to obtain supermodule genes. These supermodule genes were then integrated to obtain gene supermodules.

[0103] The gene supermodule data may include multiple supermodules, and the supermodule may include multiple genes with a preset number of weighted proportions. In addition, it may also include other data. For details, please refer to Figure 5 .

[0104] In the above embodiment, the Jaccard similarity between the preset rankings of the gene weights of each gene module is first calculated based on the gene weight list and the gene module to obtain a Jaccard similarity matrix. Then, the Jaccard similarity matrix is ​​subjected to hierarchical clustering to obtain a hierarchical clustering result. Then, super-module genes are screened based on the hierarchical clustering result and the gene module to obtain super-module genes. Finally, the super-module genes are integrated to obtain a gene super-module. Constructing super-module genes through hierarchical clustering and gene screening can effectively improve the accuracy of gene module construction.

[0105] In some embodiments, super-module genes are screened based on hierarchical clustering results and gene modules to obtain super-module genes, including: judging module similarity of gene modules based on the hierarchical clustering results to obtain inter-module similarity data, performing multi-category clustering on gene modules based on the inter-module similarity data to obtain multi-category clustering data, counting the gene frequencies in each category based on the multi-category clustering data to obtain the gene frequencies of each category, and screening gene modules based on the gene frequencies of each category to obtain super-module genes.

[0106] In some embodiments, a cluster heat map is drawn based on the hierarchical clustering results and the gene modules, and the degree of similarity between the modules is determined to obtain a module similarity determination result. The gene modules are then clustered into several categories based on the similarity determination results. The frequency of genes in each category is counted to obtain the 50 most frequently occurring supermodule genes, and the supermodule genes are integrated into a gene supermodule.

[0107] In the above embodiment, module similarity judgment is performed on the gene module according to the hierarchical clustering result to obtain inter-module similarity data, and the gene module is subjected to multi-category clustering according to the inter-module similarity data to obtain multi-category clustering data. According to the multi-category clustering data, the gene frequency in each category is counted to obtain the gene frequency of each category, and the gene module is screened according to the gene frequency of each category to obtain super-module genes. By means of multi-category clustering, the gene frequency of each category is obtained by statistics, and the gene module is screened according to the gene frequency of each category, which can effectively improve the accuracy of gene screening, thereby improving the accuracy of gene module construction.

[0108] In some embodiments, after performing hierarchical clustering on the Jaccard similarity matrix to obtain the hierarchical clustering results, the method further includes: a visualization image display step.

[0109] In an exemplary embodiment, Figure 6 As shown, the visual image display step may include but is not limited to the following steps 602 to 606. In which:

[0110] Step 602 : Perform heat map drawing based on the hierarchical clustering result and a preset heat map drawing function to obtain a correlation heat map.

[0111] Step 604 : Calculate the cellular level of each gene module according to the preset scoring model and correlation heat map to obtain cellular level data.

[0112] Step 606 , performing drawing based on the cell-level data and the heat map drawing function to obtain a spatial distribution pseudo-color map, and displaying the spatial distribution pseudo-color map.

[0113] In some embodiments, the preset heat map drawing function specifically uses the plot_local_correlations() function. First, the correlation heat map is drawn using the plot_local_correlations() function, and then the score of each module in the cell bin is calculated using the AUCell software, and a pseudo-color map of the spatial distribution of the score is drawn and displayed.

[0114] Among them, there are spatial distribution heterogeneity analyses of multiple genes on the spatial distribution pseudo-color map. For details, please refer to Figure 7 .

[0115] In the above embodiment, a heat map is drawn based on the hierarchical clustering results and a preset heat map drawing function to obtain a correlation heat map, and the cellular level of each gene module is calculated according to the preset scoring model and the correlation heat map to obtain cellular level data. The spatial distribution pseudo-color map is drawn based on the cellular level data and the heat map drawing function, and the spatial distribution pseudo-color map is displayed. This allows users to analyze gene modules more intuitively.

[0116] In another embodiment, the gene set of each gene module is subjected to COX regression survival analysis in the TCGA-PAAD data set. Specifically, the gene supermodule constructed according to gene expression is used to perform TCGA PAAD cohort overall survival (OS) analysis using the gepia2 website, and the log-rank test is used to perform hypothesis testing and draw the Kaplan-Meier survival curve of the COX proportional hazard model. Therefore, the public data set survival analysis verification of the gene module is achieved, and the authenticity of the data analysis is improved.

[0117] Among them, the gene set of each gene module is subjected to COX regression survival analysis in the TCGA-PAAD dataset to generate a prognostic survival analysis graph. For example, the prognostic survival analysis of the TCGA-PAAD dataset of MGM6 is as follows: Figure 8 As shown in the TCGA-PAAD dataset, the prognostic survival analysis of MGM11 is as follows Figure 9 shown.

[0118] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0119] Based on the same inventive concept, the present application also provides a gene module construction device for implementing the aforementioned gene module construction method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more gene module construction device embodiments provided below can be found in the above-mentioned limitations of the gene module construction method and will not be repeated here.

[0120] In an exemplary embodiment, Figure 10 As shown, a gene module construction device is provided, including: a sequencing module 1001, a preprocessing module 1002, a deconvolution module 1003, a construction module 1004, a similarity clustering module 1005 and an integration module 1006, wherein:

[0121] The sequencing module 1001 is used to perform gene sequencing based on the pathological sample cohort data to obtain a transcriptional dataset, wherein the transcriptional dataset includes: single-cell data and spatiotemporal transcriptome data.

[0122] The preprocessing module 1002 is used to perform transcriptome preprocessing on the spatiotemporal transcriptome data to obtain a transcriptome gene expression matrix.

[0123] The deconvolution module 1003 is used to perform deconvolution processing according to the type of single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data.

[0124] The construction module 1004 is used to construct a gene module according to the embedded matrix data and the preset target gene construction model to obtain the gene module of each sample slice.

[0125] The similarity clustering module 1005 is used to perform similarity clustering on the gene modules to obtain a gene weight list.

[0126] The integration module 1006 is used to integrate the gene modules according to the gene weight list to obtain a gene supermodule.

[0127] In some embodiments, the preprocessing module 1002 is further configured to perform single-cell preprocessing clustering on the single-cell data to obtain clustering results; and perform type annotation on the single-cell data according to the clustering results to obtain single-cell type data.

[0128] In some embodiments, the preprocessing module 1002 is further used to perform gene sequence comparison between the single-cell data and the preset genome data to obtain comparison data; perform data preprocessing on the comparison data according to a preset preprocessing strategy to obtain a gene count matrix; and perform cluster analysis on the gene count matrix according to a preset cluster analysis strategy to obtain a clustering result.

[0129] In some embodiments, the deconvolution module 1003 is further used to deconvolute single cell type data to obtain cell subclass label data; and embed the cell subclass label data into the transcriptome gene expression matrix according to a preset multi-layer perceptron model to obtain embedded matrix data.

[0130] In some embodiments, the integration module 1006 is further used to calculate the Jaccard similarity between the preset rankings of gene weights of each gene module based on the gene weight list and the gene module to obtain a Jaccard similarity matrix; perform hierarchical clustering on the Jaccard similarity matrix to obtain a hierarchical clustering result; perform super-module gene screening based on the hierarchical clustering result and the gene module to obtain a super-module gene; and integrate the super-module genes to obtain a gene super-module.

[0131] In some embodiments, the integration module 1006 is further used to perform module similarity judgment on the gene modules based on the hierarchical clustering results to obtain inter-module similarity data; perform multi-category clustering on the gene modules based on the inter-module similarity data to obtain multi-category clustering data; perform statistics on the gene frequencies in each category based on the multi-category clustering data to obtain the gene frequencies of each category; and screen the gene modules based on the gene frequencies of each category to obtain super-module genes.

[0132] In some embodiments, the integration module 1006 is also used to display a visual image, including: drawing a heat map based on the hierarchical clustering results and a preset heat map drawing function to obtain a correlation heat map; calculating the cellular level of each gene module based on a preset scoring model and the correlation heat map to obtain cellular level data; drawing based on the cellular level data and the heat map drawing function to obtain a spatial distribution pseudo-color map, and displaying the spatial distribution pseudo-color map.

[0133] In the above-mentioned gene module construction device, gene sequencing is first performed according to the pathological sample queue data to obtain a transcription data set, which includes: single-cell data and spatiotemporal transcriptome data. The spatiotemporal transcriptome data is subjected to transcriptome preprocessing to obtain a transcriptome gene expression matrix. Deconvolution processing is performed according to the type of single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data combining the characteristics of the two data. Gene module construction is performed based on the embedded matrix data and the preset target gene construction model to obtain the gene module of each sample slice, and the gene module is clustered by similarity to obtain a gene weight list. Finally, the gene module is integrated according to the gene weight list to obtain a gene super module. Therefore, the compatibility of multimodal data is improved by combining the characteristics of multimodal data, and the batch effect of a large number of sample slices is reduced by adopting similarity clustering and weight integration for the gene modules corresponding to a large number of sample slices.

[0134] Each module in the aforementioned gene module construction device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0135] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 11 As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface and display unit are connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be achieved via Wi-Fi, mobile cellular networks, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a gene module construction method. The display unit of the computer device is used to produce a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display or an electronic ink display.

[0136] Those skilled in the art will understand that Figure 11The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0137] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0138] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0139] It should be noted that the user information (including but not limited to sample queue information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0140] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0141] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0142] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for constructing a gene module, characterized in that: The method comprises: Gene sequencing is performed based on the pathological sample cohort data to obtain a transcriptional dataset; wherein the transcriptional dataset includes: single-cell data and spatiotemporal transcriptome data; Performing transcriptome preprocessing on the spatiotemporal transcriptome data to obtain a transcriptome gene expression matrix; Performing deconvolution processing according to the type of the single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data; Constructing a gene module based on the embedded matrix data and a preset target gene construction model to obtain a gene module for each sample slice; Performing similarity clustering on the gene modules to obtain a gene weight list; The gene modules are integrated according to the gene weight list to obtain a gene supermodule.

2. The method according to claim 1, characterized in that The type includes single cell type data, and the method further includes: Performing single-cell preprocessing clustering on the single-cell data to obtain a clustering result; The single cell data is annotated according to the clustering result to obtain the single cell type data.

3. The method according to claim 2, characterized in that The performing single cell preprocessing clustering on the single cell data to obtain a clustering result includes: Performing gene sequence comparison on the single cell data and preset genome data to obtain comparison data; Performing data preprocessing on the comparison data according to a preset preprocessing strategy to obtain a gene count matrix; Cluster analysis is performed on the gene count matrix according to a preset cluster analysis strategy to obtain the clustering result.

4. The method according to claim 1, wherein The type of the single-cell data includes single-cell type data; the deconvolution processing is performed according to the type of the single-cell data and the transcriptome gene expression matrix to obtain the embedded matrix data, including: Deconvolving the single cell type data to obtain cell subclass label data; The cell subclass label data is embedded into the transcriptome gene expression matrix according to a preset multi-layer perceptron model to obtain the embedded matrix data.

5. The method according to any one of claims 1 to 4, characterized in that The step of integrating the gene modules according to the gene weight list to obtain a gene supermodule includes: Calculating the Jaccard similarity between the gene weight preset rankings of each gene module according to the gene weight list and the gene module to obtain a Jaccard similarity matrix; Performing hierarchical clustering on the Jaccard similarity matrix to obtain a hierarchical clustering result; Performing super-module gene screening based on the hierarchical clustering results and the gene module to obtain super-module genes; The supermodule genes are integrated to obtain the gene supermodule.

6. The method according to claim 5, characterized in that The super-module gene screening is performed according to the hierarchical clustering result and the gene module to obtain the super-module gene, comprising: Performing module similarity judgment on the gene modules according to the hierarchical clustering results to obtain inter-module similarity data; Performing multi-category clustering on the gene modules according to the inter-module similarity data to obtain multi-category clustering data; Counting the gene frequencies in each category according to the multi-category clustering data to obtain the gene frequencies of each category; The gene modules are screened according to the frequencies of the genes in each category to obtain the supermodule genes.

7. The method according to claim 5, characterized in that After performing hierarchical clustering on the Jaccard similarity matrix to obtain a hierarchical clustering result, the method further includes: Perform heat map drawing according to the hierarchical clustering result and a preset heat map drawing function to obtain a correlation heat map; Calculating the cellular level of each gene module according to a preset scoring model and the correlation heat map to obtain cellular level data; Drawing is performed according to the cell-level data and the heat map drawing function to obtain a spatial distribution pseudo-color map, and the spatial distribution pseudo-color map is displayed.

8. A gene module construction device, characterized in that: The device comprises: A sequencing module is used to perform gene sequencing based on the pathological sample cohort data to obtain a transcriptional dataset; wherein the transcriptional dataset includes: single-cell data and spatiotemporal transcriptome data; A preprocessing module, configured to perform transcriptome preprocessing on the spatiotemporal transcriptome data to obtain a transcriptome gene expression matrix; a deconvolution module, configured to perform deconvolution processing according to the type of the single-cell data and the transcriptome gene expression matrix to obtain embedded matrix data; A construction module is used to construct a gene module according to the embedded matrix data and a preset target gene construction model to obtain a gene module for each sample slice; A similarity clustering module is used to perform similarity clustering on the gene modules to obtain a gene weight list; An integration module is used to integrate the gene modules according to the gene weight list to obtain a gene supermodule.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.