Cell type and cell abundance identification method and system based on cross-modal training

Through the cross-modal joint representation learning framework, the problem of difficult to achieve fine-grained cell type recognition in the prior art is solved, and efficient prediction of cell abundance and cell type from histopathological images is achieved.

CN120148030AActive Publication Date: 2025-06-13NANKAI UNIV

Patent Information

Application Number
CN202510327527.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-13
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The prior art is difficult to achieve fine-grained recognition in cell type recognition, and relying on manual annotation and gene expression data, it fails to effectively utilize the morphological patterns in histopathological images.

Method used

A cross-modal joint representation learning framework is adopted to integrate image morphology and molecular gene expression patterns, and fine-grained cell types and cell abundance are identified from histopathological images through cross-modal training models.

Benefits of technology

The ability to predict cell abundance from histological images is improved, the spatial distribution of fine-grained cell types is revealed, and abundance information for predicting fine-grained cell types and corresponding cell types is achieved only from histopathological images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148030A_ABST
    Figure CN120148030A_ABST
Patent Text Reader

Abstract

The invention provides a cell type and cell abundance identification method and system based on cross-modal training, relates to the field of image processing, and aims to solve the problems that the existing identification technology mostly depends on non-standard factors such as manual labeling position annotation information and the like, abundant morphological modes in a tissue pathological image are not fully utilized, and the identification accuracy is poor. And the identification reliability and accuracy are influenced. The method comprises the following steps: acquiring spatial transcriptomics data matched with pathological image-gene expression, and preprocessing the spatial transcriptomics data to obtain high-expression gene expression data, local image blocks, cell types and abundance tags; constructing a cross-modal joint representation learning model, and inputting high-expression gene expression data and local image blocks into the model for training; and predicting a to-be-predicted histological image based on the trained model. According to the method, the problems in the prior art are solved, the capability of predicting the cell abundance from the histological image is improved, and the spatial distribution of fine-grained cell types is fully revealed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and particularly relates to a method and system for identifying cell types and cell abundances based on cross-modal training. Background Art

[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Cells are the basic units of life and constitute the structural and functional basis of the tissues and organs of organisms. The cellular structure of tissues, referring to the spatial arrangement and morphological characteristics of cells, provides key insights into how cell interactions contribute to biological behaviors, including tissue development, disease progression, and treatment response. In cancer research, the cellular spatial structure can elucidate the key features of the tumor microenvironment, heterogeneity, tumor lymphocyte infiltration, and their impact on patient prognosis and personalized treatment strategies.

[0004] Currently, the methods for cell type identification can be roughly divided into two groups.

[0005] (1) Methods based on computational pathology, which achieve nuclear instance segmentation and classification through deep learning. These methods can identify cell types in histopathological images, but they are limited to identifying coarse-grained cell categories, usually no more than five main cell types. This limits the exploration of more refined cell type subtypes. Moreover, these cell type identification methods all rely on manually annotated location annotation information, which introduces non-standardized factors and affects the reliability of the final inference and identification results.

[0006] (2) Methods based on spatial transcriptomics (ST), such as Seurat, RCTD, and Cell2location, usually adopt cell type deconvolution algorithms to estimate the cell types and proportions within each grid of the spatial transcriptome using gene expression data. These methods integrate and utilize single-cell RNA sequencing (scRNA-seq) reference transcriptome data to achieve fine-grained cell type resolution based on spatial transcriptome profiles. However, these methods mainly rely on gene expression and fail to utilize the rich morphological patterns present in histopathological images. Summary of the Invention

[0007] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a method and system for identifying cell types and cell abundances based on cross-modal training. Through a cross-modal joint representation learning framework, it integrates image morphology and molecular gene expression patterns, enhances the interaction between different modalities, and realizes the identification of fine-grained cell types and the abundance information of corresponding cell types only from histopathological images.

[0008] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0009] The first aspect of the present invention provides a method for identifying cell types and cell abundances based on cross-modal training, including:

[0010] Obtain spatially transcriptomic data with matched pathological images and gene expressions;

[0011] Preprocess the spatially transcriptomic data with matched pathological images and gene expressions to obtain a preprocessed comprehensive dataset, which includes highly expressed gene expression data, local image patches, cell type, and cell abundance labels;

[0012] Construct a cross-modal joint representation learning model, input the highly expressed gene expression data and local image patches into the cross-modal joint representation learning model for training, and obtain a trained cross-modal joint representation learning model;

[0013] Based on the trained cross-modal joint representation learning model, predict the cell type and cell abundance of the histological image to be identified.

[0014] As an implementation, the specific process of preprocessing the spatially transcriptomic data with matched pathological images and gene expressions is as follows:

[0015] For the spatially transcriptomic data with matched pathological images and gene expressions, use the cell type deconvolution method to generate fine-grained cell types and corresponding cell abundance labels;

[0016] Extract local image patches and highly expressed gene expression data from the spatially transcriptomic data with matched pathological images and gene expressions;

[0017] Construct a comprehensive dataset with fine-grained cell types, corresponding cell abundance labels, local image patches, and highly expressed gene expression data.

[0018] As an implementation, construct a cross-modal joint representation learning model, where the cross-modal joint representation learning model includes a morphological modality representation module, a molecular modality representation module, and a cross-modal embedding alignment module.

[0019] As an implementation, the specific process of inputting the highly expressed gene expression data and local image patches into the cross-modal joint representation learning model for training is as follows:

[0020] Input the local image patches into the morphological modality representation module to obtain image feature embeddings;

[0021] Input the highly expressed gene expression data into the molecular modality representation module to obtain molecular expression embeddings;

[0022] Input the image feature embedding and the molecular expression embedding into the cross-modal embedding alignment module to obtain the predicted abundance value of cell types and the original gene expression pattern respectively.

[0023] As an implementation, input the local image patch into the morphological modality representation module, where the morphological modality representation module includes a feature backbone module and a transformation layer. The specific process is as follows:

[0024] Extract the morphological features of the local image patch through the feature backbone module;

[0025] Input the morphological features of the local image patch into the transformation layer for non-linear transformation to obtain the image patch feature embedding.

[0026] As an implementation, input the high-expression gene expression data into the molecular modality representation module, where the molecular modality representation module includes a self-normalizing network module and a transformation layer. The specific process is as follows:

[0027] Through the self-normalizing network module, enhance the features of the high-expression gene expression data to obtain the enhanced high-expression gene expression data;

[0028] Through the transformation layer, map the enhanced high-expression gene expression data to the molecular expression embedding.

[0029] As an implementation, input the image feature embedding and the molecular expression embedding into the cross-modal embedding alignment module, which includes a multi-layer perceptron and a decoder. The specific process is as follows:

[0030] Align the image patch feature embedding and the molecular expression embedding;

[0031] Input the aligned image patch feature embedding into the multi-layer perceptron to obtain the predicted abundance value of cell types;

[0032] Input the aligned molecular expression embedding into the decoder to reconstruct the original gene expression pattern.

[0033] As an implementation, optimize the model through the overall loss. The overall loss formula is:

[0034]

[0035] Where and represent the cross-modal consistency comparison loss cell abundance prediction loss and gene expression reconstruction loss respectively; the balance weights of θ morph 、θ molec 、 and are and For the parameters in , the argmin(·) term aims to find the optimization parameters that minimize the overall loss function

[0036] The second aspect of the present invention provides a cell type and cell abundance recognition system based on cross-modal training, including:

[0037] A data acquisition module for acquiring spatially transcriptomic data with matched pathological images and gene expressions;

[0038] A data processing module for preprocessing the spatially transcriptomic data with matched pathological images and gene expressions to obtain a preprocessed comprehensive data set, including highly expressed gene expression data, local image patches, cell type, and cell abundance labels;

[0039] A model construction and training module for constructing a cross-modal joint representation learning model, inputting the highly expressed gene expression data and local image patches into the cross-modal joint representation learning model for training, and obtaining a trained cross-modal joint representation learning model;

[0040] A model recognition module for predicting cell type and cell abundance of a histological image to be recognized based on the trained cross-modal joint representation learning model.

[0041] The third aspect of the present invention provides a computer device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the method described in the first aspect of the present invention.

[0042] The above one or more technical solutions have the following beneficial effects:

[0043] In this embodiment, by constructing a cross-modal joint representation learning model, cross-modal training for identifying fine-grained cell types and cell abundances from pathological images is achieved. This model not only improves the ability to predict cell abundances from histological images but also reveals the spatial distribution of fine-grained cell types, enabling the prediction of fine-grained cell types and the abundance information of corresponding cell types solely from tissue pathological images.

[0044] In this embodiment, a morphological modal representation module is used to learn the morphological patterns presented in local image patches. Meanwhile, a molecular modal representation module is introduced to extract the inherent key molecular features in gene expression data through a gene expression reconstruction task, making full use of the rich morphological patterns existing in tissue pathological images.

[0045] ​In this embodiment, by using cross-modal embedding during the training process to align the embedding spaces of the morphological and molecular modalities, the image morphology and molecular gene expression patterns can be integrated, presenting a better fine-grained cell type recognition effect, which reveals that the molecular gene expression information enhances the morphological pattern features in the image.

[0046] In this embodiment, the constructed cross-modal joint representation learning model exhibits more excellent capabilities during the inference application stage. It can analyze the fine-grained cell spatial distribution, reveal the co-localization interaction patterns between cell types, and provide valuable insights into the intercellular spatial representation within the tumor ecosystem.

[0047] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0049] Figure 1 It is a flowchart of a method for identifying cell types and cell abundances based on cross-modal training in the first embodiment.

[0050] Figure 2 It is a flowchart for preprocessing spatial transcriptome data in the first embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0052] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0053] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0054] Embodiment 1

[0055] This embodiment discloses a method for identifying cell types and cell abundances based on cross-modal training.

[0056] To more clearly illustrate this embodiment, the implementation process of a method for identifying cell types and cell abundances based on cross-modal training can be specifically described as follows:

[0057] As Figure 1As shown, a method for identifying cell types and cell abundances based on cross-modal training includes:

[0058] S1. Obtain spatially transcriptomic data with matched pathological images and gene expressions;

[0059] S2. Preprocess the spatially transcriptomic data with matched pathological images and gene expressions to obtain a preprocessed comprehensive dataset, which includes highly expressed gene expression data, local image patches, cell type, and cell abundance labels;

[0060] S3. Construct a cross-modal joint representation learning model, input the highly expressed gene expression data and local image patches into the cross-modal joint representation learning model for training, and obtain a trained cross-modal joint representation learning model;

[0061] S4. Based on the trained cross-modal joint representation learning model, predict the cell type and cell abundance of the histological image to be identified.

[0062] In step S1, spatially transcriptomic data with matched pathological images and gene expressions is obtained.

[0063] In this embodiment, the data used is complete and spatially transcriptomic data with matched morphological pathological images and gene expressions.

[0064] This data is spatially transcriptomic data, which includes basic information, gene expression information at different sequenced positions, and the corresponding pathological section image regions. The spatially transcriptomic data is directly downloaded from the existing research sequencing data.

[0065] In this embodiment, a single-cell transcriptomic atlas is simultaneously downloaded and obtained from the existing research sequencing data.

[0066] After the above steps, the original spatially transcriptomic data is downloaded, providing a comprehensive data basis for subsequent operations and model training, obtaining the required dataset, ensuring the integrity and accuracy of the data, and providing support for subsequent analysis.

[0067] As Figure 2 shown, in step S2, the spatially transcriptomic data with matched pathological images and gene expressions is preprocessed to obtain a preprocessed comprehensive dataset, which includes highly expressed gene expression data, local image patches, cell type, and cell abundance labels.

[0068] In this embodiment, the preprocessing of the original spatially transcriptomic data is to generate fine-grained cell types and corresponding cell abundances using a cell type deconvolution method

[0069] Preprocess the spatial transcriptomics data of pathological image-gene expression matching. The specific process is as follows:

[0070] (1) For the spatial transcriptomics data of pathological image-gene expression matching, use the cell type deconvolution method to generate fine-grained cell types and corresponding cell abundance labels.

[0071] Based on this single-cell transcriptomics atlas as a reference, use the cell deconvolution method D(·) to estimate the cell abundance of the spatial transcriptomics data.

[0072] Generate cell abundance estimates, that is, cell abundance labels, according to the single-cell transcriptomics atlas and the gene expression data of the spatial transcriptomics data grid.

[0073] In this embodiment, use the reference single-cell transcriptomics atlas and the gene expression data of the spatial transcriptomics data grid to generate cell abundance estimates as the true labels for subsequent model training.

[0074] (2) Extract local image patches and highly expressed gene expression data from the spatial transcriptomics data of pathological image-gene expression matching.

[0075] In this embodiment, the specific process is as follows: 1) Extract the highly expressed genes in the previous step.

[0076] Specifically, first, standardize the original data using TPM (Transcripts Per Million), or alternatively, the FPKM (Fragments Per Kilobase Million) normalization method can be used. The purpose is to eliminate the influence of sequencing depth and gene length.

[0077] Then, filter out low-expressed genes to remove genes with expression levels below a certain threshold in the cohort samples.

[0078] Furthermore, perform a log2 logarithmic transformation on the gene matrix to reduce the influence of data skewed distribution.

[0079] Finally, sort all genes in descending order according to the selected expression level index (average expression) and extract the gene expression results corresponding to the top 250 genes.

[0080] 2) Extract local image patches.

[0081] Extract local image patches of 224×224 pixels in size from the histological pathological image at a magnification of 20×.

[0082] Specifically, first, the Otsu threshold method is used to distinguish the tissue area from the background and extract the effective tissue area; then, for the extracted tissue area, a non-overlapping sliding window image sampling method with a fixed step size (such as 224 pixels here) is used to extract image patches of 224×224 pixels.

[0083] (3) Construct a comprehensive dataset by combining fine-grained cell types, corresponding cell abundance labels, local image patches, and high-expressed gene expression data.

[0084] In this embodiment, a comprehensive dataset X is finally generated i , including high-expressed gene expression data local image patches and labels of cell types and abundances

[0085] The comprehensive dataset is represented as:

[0086]

[0087] After the above steps, the original spatial transcriptomics data is preprocessed to obtain label information of cell types and cell abundances for subsequent model training.

[0088] As Figure 1 shown, in step S3, a cross-modal joint representation learning model is constructed, and the high-expressed gene expression data and local image patches are input into the cross-modal joint representation learning model for training to obtain a trained cross-modal joint representation learning model.

[0089] S3-1. Construct a cross-modal joint representation learning model.

[0090] In this embodiment, the cross-modal joint representation learning model includes three key components: a morphological modal representation module for local image patches in the histological pathology images of spatial transcriptomics, a molecular modal representation module based on high-expressed gene expression data, and a cross-modal embedding alignment module for integrating tissue morphology and molecular gene patterns in the training stage.

[0091] (1) Morphological modal representation module.

[0092] In this embodiment, the morphological modal representation module includes a feature backbone module and a transformation layer, and the morphological modal representation module is used to learn the morphological patterns presented in local image patches.

[0093] (2) Molecular modal representation module.

[0094] In this embodiment, the molecular modal representation module includes a self-normalizing network module and a transformation layer. The molecular modal representation module is introduced to extract the inherent key molecular features in gene expression data through a gene expression reconstruction task.

[0095] (3) Cross-modal Embedding Alignment Module.

[0096] In this embodiment, the cross-modal embedding alignment module includes a multi-layer perceptron and a decoder.

[0097] The cross-modal embedding alignment module aligns the embedding spaces of the morphology and molecular modalities during the training process, integrates the image morphology and molecular gene expression patterns, and enhances the interaction between different modalities.

[0098] S3-2. Input the high-expression gene expression data and local image patches into the cross-modal joint representation learning model for training to obtain a trained cross-modal joint representation learning model.

[0099] Inputting the high-expression gene expression data and local image patches into the cross-modal joint representation learning model for training, the specific process is as follows:

[0100] S3-2-1. Input the local image patches into the morphological modality representation module to obtain image feature embeddings.

[0101] In this embodiment, the morphological modality representation module is denoted as which aims to directly extract morphological features from local image patches of spatial transcriptomics data.

[0102] Inputting the local image patches into the morphological modality representation module, the specific process is as follows:

[0103] (1) Extract the morphological features of the local image patches through the feature backbone module.

[0104] In this embodiment, an existing pathological foundation model (PFM) is used as the feature backbone structure and it is fine-tuned using the local image patch data of spatial transcriptomics data to capture cell and histological morphological patterns.

[0105] Among them, due to the strong image representation ability of the PFM pathological foundation model, the PFM model is selected. Considering the large number of parameters of the PFM foundation model, the LoRA fine-tuning method is used. LoRA is an efficient fine-tuning method that significantly reduces the number of parameters to be trained by introducing low-rank decomposition (Low-Rank Decomposition) into the weight matrix of the pre-trained model while maintaining the model performance. Its core idea is to approximate the weight update with the product of two low-rank matrices.

[0106] Specifically, first load the pre-trained pathological foundation model (PFM), freeze most of the parameters of the pre-trained model, and only train the low-rank matrices A and B introduced by LoRA while keeping the original weight W unchanged.

[0107] During the inference phase, the updates of the LoRA module are merged into the original weights:

[0108] W′ = W + A·B (2)

[0109] Therefore, the core formula of LoRA is the weight update of low-rank factorization: By adopting this fine-tuning strategy, the number of parameters to be trained can be significantly reduced while maintaining the model performance.

[0110] Specifically, local image patch data of 224×224 pixels is used as the input of this branch, and then the fine-tuned PFM model is used to extract relevant morphological features.

[0111] (2) The morphological features of the local image patches are input into the transformation layer for non-linear transformation to obtain image patch feature embeddings.

[0112] In this embodiment, the transformation layer performs non-linear transformation on the features, which can ensure that the morphological feature embeddings applicable to this task are learned.

[0113] The obtained morphological features are non-linearly transformed through the processing of the projection head in the transformation layer to generate morphological-modal embeddings a with a dimension of d i.e., image patch feature embeddings, and the formula is:

[0114]

[0115] After the above steps, through the morphological-modal representation module, the morphological patterns presented in the local image patches of the H&E stained pathological images in the spatial transcriptomics data can be learned, and at the same time, it provides one of the key data for subsequent integration of image morphology and molecular gene expression patterns.

[0116] S3-2-2. Input the high-expression gene expression data into the molecular-modal representation module to obtain molecular expression embeddings.

[0117] In this embodiment, the molecular-modal representation module is denoted as and is used to learn the molecular patterns at the gene level from the gene expression data corresponding to the local image patches of the spatial transcriptomics data.

[0118] The process of inputting the high-expression gene expression data into the molecular-modal representation module is as follows:

[0119] (1) Through the self-normalizing network module, the high-expression gene expression data is enhanced in features to obtain enhanced high-expression gene expression data.

[0120] In this example, the highly expressed gene data is processed through an optimized self-normalizing network module SNN to obtain enhanced highly expressed gene expression data, denoted as

[0121] By using a self-normalizing network for molecular representation learning to enhance the expression ability of molecular feature representations.

[0122] Among them, SNN is a prior art, and a single-cell based model can also be improved and utilized to further enhance the feature representation ability of gene expression data.

[0123] (2) Through a transformation layer, the enhanced highly expressed gene expression data is mapped into molecular expression embeddings.

[0124] In this embodiment, in the transformation layer, the projection head maps the features into embeddings of the molecular modality. These embeddings are called gene expression embeddings That is, molecular expression embeddings, and their dimension is also d a

[0125]

[0126] After the above steps, key molecular features inherent in the gene expression data are learned and extracted through a gene expression reconstruction task, and are used for feature alignment with morphological pattern features.

[0127] S3-2-3. Input the image feature embeddings and molecular expression embeddings into a cross-modal representation learning module to obtain the predicted abundance values of cell types and the original gene expression patterns respectively.

[0128] In this embodiment, based on the morphological modality representation and the molecular modality feature representation a cross-modal embedding alignment design is proposed to combine the gene expression patterns with histological features.

[0129] Inputting the image feature embeddings and molecular expression embeddings into the cross-modal representation learning module, the specific process is as follows:

[0130] (1) Align the image patch feature embeddings and molecular expression embeddings.

[0131] In this embodiment, in order to unify the features of morphological histology and molecular expression, a cross-modal consistency comparison loss is constructed to align the embedding spaces of the two modalities.

[0132] Specifically, for a sample set with a batch size of B Morphological modality image embedding (patch feature embedding) and molecular modality gene expression embedding (molecular expression embedding) in pairs The learning objective is defined as:

[0133]

[0134] where ||·|| 2 represents the L2 norm, aiming to minimize the distance between the paired morphological and molecular modality embeddings from the same sample.

[0135] (2) Input the aligned patch feature embeddings into a multi-layer perceptron to obtain the predicted abundance values of cell types.

[0136] In this embodiment, in order to establish an end-to-end training framework for cross-modal embedding alignment and cell type prediction, the cell abundance estimation label y m is used as a supervision signal to learn the cell type pattern from the histological image p m .

[0137] Specifically, the aligned embedding of the morphological modality is fed into a multi-layer perceptron to predict the cell abundance spectrum. This process linearly transforms the aligned morphological modality image embedding (patch feature embedding) into the predicted abundance values of cell types The loss function is defined as:

[0138]

[0139] Given that the cell abundance label y m is continuous, a regression strategy is adopted, and the root mean square error loss is used as the objective function to optimize the model for predicting the cell abundance spectrum.

[0140] (3) Input the aligned molecular expression embeddings into a decoder to reconstruct the original gene expression pattern.

[0141] In this embodiment, in order to enhance the model's ability to extract the inherent molecular information s m in gene expression data, a gene expression reconstruction task is proposed. The aligned molecular modality gene expression embedding (molecular expression embedding) is input into the decoder block to reconstruct the original gene expression pattern using the learned cross-modal molecular representation , and the formula is:

[0142]

[0143] where Represents the reconstructed gene expression value, Represents the gene expression reconstruction loss function.

[0144] This method effectively captures molecular information to enrich histological representations, thus contributing to the identification of fine-grained cell types from histological images.

[0145] After the above steps, by using cross-modal embeddings to align the embedding spaces of the morphological and molecular modalities during training, it is possible to integrate the image morphology and molecular gene expression patterns, presenting a better fine-grained cell type recognition effect, which reveals that the molecular gene expression information enhances the morphological pattern features in the image.

[0146] S3-3. Construct the overall loss optimization.

[0147] In this embodiment, the model is optimized through the overall loss, and the overall loss objective Is expressed as the weighted sum of three loss objective terms, and the formula is:

[0148]

[0149] Among them, And Respectively represent the cross-modal consistency alignment loss Cell abundance prediction loss And gene expression reconstruction loss Of the balance weights; θ morph 、θ molec 、 And Is And The parameters in, and the argmin(·) term aims to find the optimization parameters that minimize the overall loss function Of

[0150] Such as Figure 1 Shown, in step S4, based on the trained cross-modal joint representation learning model, the histological image to be recognized is used for cell type and cell abundance prediction.

[0151] In this embodiment, the histological image is directly fed into the fine-tuned PFMs encoder and projection head for processing to extract cross-modal embedding features, so as to capture the morphological patterns enhanced by molecular gene expression information. Then the MLP regression head converts the obtained embeddings into predicted fine-grained cell types and corresponding cell abundance values.

[0152] In this embodiment, by integrating and utilizing the rich morphological patterns existing in histopathological images, a cross-modal feature alignment training method for pathological morphology-molecular genes is proposed, and finally, the automatic prediction of fine-grained cell types and the abundance information of corresponding cell types is realized in the inference stage. Compared with the method based on computational pathology, the method of this embodiment can identify more and finer-grained cell types and the corresponding cell richness; compared with the method based on spatial transcriptomics (ST), the method of this embodiment jointly utilizes the rich morphological patterns existing in histopathological images, and can realize the prediction of fine-grained cell types and the abundance information of corresponding cell types only from histopathological images, further reflecting the spatial distribution of their respective cell types, and providing the interaction patterns between different cell types, which has the potential to reveal key insights into how cell communication promotes biological behavior.

[0153] Embodiment 2

[0154] The purpose of this embodiment is to provide a cell type and cell abundance recognition system based on cross-modal training, including:

[0155] A data acquisition module, configured to acquire spatial transcriptomics data with matched pathological images and gene expressions;

[0156] A data processing module, configured to preprocess the spatial transcriptomics data with matched pathological images and gene expressions to obtain a preprocessed comprehensive data set, including high-expression gene expression data, local image patches, cell types, and cell abundance labels;

[0157] A model construction and training module, configured to construct a cross-modal joint representation learning model, input the high-expression gene expression data and local image patches into the cross-modal joint representation learning model for training, and obtain a trained cross-modal joint representation learning model;

[0158] A model recognition module, configured to predict cell types and cell abundances for the histological image to be recognized based on the trained cross-modal joint representation learning model.

[0159] Based on providing a cell type and cell abundance recognition system based on cross-modal training, the method steps in Embodiment 1 are implemented.

[0160] Embodiment 3

[0161] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the above method are implemented.

[0162] Embodiment 4

[0163] The purpose of this embodiment is to provide a computer-readable storage medium.

[0164] A computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the above method are executed.

[0165] Example Five

[0166] The purpose of this embodiment is to provide a computer program product containing instructions, which, when running on a computer, enables the computer to execute the methods and functions involved in any one of the above embodiments.

[0167] The steps involved in the devices of the above embodiments correspond to those of Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0168] Those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device for execution by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0169] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.

Claims

1. A method for identifying cell types and cell abundance based on cross-modal training, characterized in that: include: Acquire spatial transcriptomics data matching pathological images and gene expression; Preprocessing the spatial transcriptomics data of the pathological image-gene expression matching to obtain a preprocessed comprehensive data set, which includes highly expressed gene expression data, local image blocks, cell types and cell abundance labels; Construct a cross-modal joint representation learning model, input the highly expressed gene expression data and the local image blocks into the cross-modal joint representation learning model for training, and obtain a trained cross-modal joint representation learning model; Based on the trained cross-modal joint representation learning model, cell type and cell abundance prediction is performed on the histological images to be predicted.

2. A method for identifying cell types and cell abundance based on cross-modal training as claimed in claim 1, characterized in that: The spatial transcriptomics data of the pathological image-gene expression matching is preprocessed, and the specific process is as follows: For the spatial transcriptomics data of pathological image-gene expression matching, a cell type deconvolution method is used to generate fine-grained cell types and corresponding cell abundance labels; Extracting local image blocks and highly expressed gene expression data from the spatial transcriptomics data of the pathological image-gene expression matching; A comprehensive dataset is constructed by combining fine-grained cell types, corresponding cell abundance labels, local image patches, and highly expressed gene expression data.

3. A method for identifying cell types and cell abundance based on cross-modal training as claimed in claim 1, characterized in that: A cross-modal joint representation learning model is constructed, wherein the cross-modal joint representation learning model includes a morphological modality representation module, a molecular modality representation module and a cross-modal embedding alignment module.

4. A method for identifying cell types and cell abundance based on cross-modal training as claimed in claim 1, characterized in that: The highly expressed gene expression data and local image patches are input into the cross-modal joint representation learning model for training. The specific process is as follows: Input the local image block into the morphological modality representation module to obtain image feature embedding; The expression data of highly expressed genes are input into the molecular modality representation module to obtain molecular expression embedding; The image feature embedding and molecular expression embedding are input into the cross-modal embedding alignment module to obtain the predicted abundance value and original gene expression pattern of the cell type, respectively.

5. A method for identifying cell types and cell abundance based on cross-modal training as claimed in claim 4, characterized in that: The local image block is input into the morphological modality representation module, where the morphological modality representation module includes a feature backbone module and a transformation layer. The specific process is as follows: The morphological features of the local image blocks are extracted through the feature backbone module; The morphological features of the local image block are input into the transformation layer for nonlinear transformation to obtain the image block feature embedding.

6. A method for identifying cell types and cell abundance based on cross-modal training as claimed in claim 4, characterized in that: The expression data of highly expressed genes are input into the molecular modality representation module, wherein the molecular modality representation module includes a self-normalization network module and a transformation layer. The specific process is as follows: The feature enhancement of the highly expressed gene expression data is performed through the self-normalization network module to obtain the enhanced highly expressed gene expression data; Through the transformation layer, the enhanced highly expressed gene expression data are mapped to molecular expression embeddings.

7. A method for identifying cell types and cell abundance based on cross-modal training as claimed in claim 4, characterized in that: The image feature embedding and the molecular expression embedding are input into the cross-modal embedding alignment module, which includes a multi-layer perceptron and a decoder. The specific process is as follows: Align image patch feature embeddings with molecule expression embeddings; The aligned image patch features are embedded and input into a multi-layer perceptron to obtain the predicted abundance value of the cell type; The aligned molecular expression embeddings are input to the decoder to reconstruct the original gene expression patterns.

8. A method for identifying cell types and cell abundance based on cross-modal training as claimed in claim 1, characterized in that: The model is optimized by the overall loss, and the overall loss formula is: in and Represent the cross-modal consistency comparison loss Cell abundance predicts loss and gene expression reconstruction loss The balance weight of θ morph ,θ molec , and yes and The argmin(·) term aims to find the parameter that minimizes the overall loss function Optimization parameters 9. A cell type and cell abundance recognition system based on cross-modal training, characterized in that: include: Data acquisition module, used to obtain spatial transcriptomics data matching pathological images and gene expression; A data processing module, used for preprocessing the spatial transcriptomics data of the pathological image-gene expression matching to obtain a preprocessed comprehensive data set, which includes highly expressed gene expression data, local image blocks, cell types and cell abundance labels; The model building and training module is used to build a cross-modal joint representation learning model, input the highly expressed gene expression data and the local image blocks into the cross-modal joint representation learning model for training, and obtain a trained cross-modal joint representation learning model; The model recognition module is used to predict the cell type and cell abundance of the histological images to be identified based on the trained cross-modal joint representation learning model.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Method and system for identifying spatial transcriptome cell expression pattern

    CN115732034A

  • Systems and Methods for Deconvolving Cell Types in Histology Slide Images, Using Super-Resolution Spatial Transcriptomics Data

    US20240161519A1

Cited By

  • Single cell transcriptome data and text description conjoint analysis method based on multi-modal language model

    CN120452543A

  • Text labeling method and device for cell image, electronic equipment and program product

    CN120510612A

  • Spatial omics-based intestinal cancer metastasis prediction method and device, medium and equipment

    CN120913863A

  • Metastasis prediction method, device, medium and equipment for colorectal cancer based on spatial omics

    CN120913863B

  • Breast cancer detection method and system based on morphological image and space transcriptome cross-graph collaborative learning

    CN121582228A