Multi-modal integration method for unicellular omics based on linear selective state space

Through a multimodal integration method based on linear selective state space, transcriptome data and proteomic data are integrated, and the problem of insufficient data integration in the prior art is solved, achieving more accurate cell type recognition.

CN120220823APending Publication Date: 2025-06-27GREATER BAY AREA UNIV (IN PREPARATION)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510140936.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate transcriptomics and proteomics data, resulting in insufficient accuracy and credibility of cell type recognition.

Method used

A multimodal integration method based on linear selective state space is adopted to enhance the fusion network by constructing a selective state space, integrating transcriptome data and proteomic data, extracting common features and unique details, and improving the accuracy of cell type recognition.

Benefits of technology

The effective integration of transcriptomic data and proteomic data is achieved, the accuracy and credibility of cell type recognition are improved, and the overall fusion performance is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220823A_ABST
    Figure CN120220823A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biological information, and particularly discloses a multi-modal integration method for unicellular omics based on a linear selective state space. Comprising the following steps: (1) performing quality control on original multi-omics data, and removing low-quality samples; (2) carrying out normalization and hypervariant gene selection on the processed multi-omics data; (3) constructing an omics feature fusion module, integrating common features of multiple omics, constructing an omics feature enhancement module, and retaining and enhancing unique details of each omics; (4) constructing a selective state space fusion module, and integrating the unique information and common features of each group of science; (5) constructing a selective state space enhanced fusion network, and connecting modules in series to form an integral network structure; and (6) clustering the cells and comparing the clustered cells with real labels to obtain clustering performance. According to the method, the transcriptome data and the proteome data can be well integrated, and unique information is reserved while common features of the multi-modal data are extracted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics technology, and particularly relates to a multimodal integration method for single-cell omics based on a linear selective state space. Background Art

[0002] Sequencing-based transcriptome and cellular indexing of transcript epitopes (CITE-seq) is a multi-omics technology. This technology can combine gene expression information and cell surface protein information at the single-cell level through sequencing, so as to achieve the purpose of simultaneously detecting protein information on the cell surface and intracellular transcriptome information, including transcriptome data and proteome data. CITE-seq can more deeply distinguish cell heterogeneity, more accurately mine specific cell types, and explore the mechanisms behind biological phenomena such as treatment resistance.

[0003] Transcriptome data refers to the summary of all RNA molecules in an organism and the measurement results of their expression levels. These data reflect the gene expression status of different tissues and cell types, and have important value in fields such as gene function research, exploration of disease mechanisms, and drug development.

[0004] Proteomics data refers to the measurement results of the summary of all protein molecules in an organism. It mainly includes the expression level of proteins, protein activity, the modified status, the interaction with other proteins or molecules, subcellular localization, and three-dimensional structure. These data usually exist in the form of experimental data sets and are used for bioinformatics and statistical analysis to study the protein functions and regulatory mechanisms in the process of life activities.

[0005] Traditional data integration methods often fail to fully utilize the unique information between different omics data, resulting in limitations in classification accuracy and biological feature analysis ability. Integrating multimodal data, especially transcriptomics data and proteomics data, is very difficult, and how to further improve the accuracy and reliability of cell type recognition using the integrated data is also one of the problems. Therefore, it is necessary to provide a multimodal integration method for single-cell omics based on a linear selective state space to effectively integrate transcriptomics and proteomics data, so as to improve the accuracy and reliability of cell type recognition. Summary of the Invention

[0006] The present invention aims to solve at least one of the technical problems existing in the above-mentioned prior art. For this purpose, the present invention proposes a multimodal integration method for single-cell omics based on a linear selective state space to effectively integrate transcriptomics and proteomics data, so as to improve the accuracy and reliability of cell type recognition.

[0007] The first aspect of the present invention provides a multimodal integration method for single-cell omics based on a linear selective state space.

[0008] Specifically, it includes the following steps:

[0009] (1) Perform quality control on the original multi-omics data to remove low-quality samples;

[0010] (2) Normalize and select highly variable genes for the processed multi-omics data;

[0011] (3) Construct an omics feature fusion module to integrate the common features of multi-omics, and construct an omics feature enhancement module to retain and enhance the unique details of each omics;

[0012] (4) Construct a selective state space fusion module to integrate the unique information and common features of each omics;

[0013] (5) Construct a selective state space enhanced fusion network, and connect each module in series to form an overall network structure;

[0014] (6) Cluster the cells and compare them with the true labels to obtain the clustering performance.

[0015] Preferably, in step (1), the low-quality samples include at least one of cells with less than 200 gene expressions, genes expressed in less than 3 cells, and cells with a mitochondrial gene proportion exceeding 20%.

[0016] Preferably, in step (2), the normalization is to perform RPKM normalization on the transcriptome data and CLR normalization on the proteome data.

[0017] Preferably, the RPKM normalization is to divide the number of reads mapped to a gene by the number of reads mapped to all genomes.

[0018] Preferably, the CLR normalization is the centered log-ratio transformation to make the proteome data conform to the normal distribution rule.

[0019] Preferably, in step (2), the method for selecting highly variable genes includes selecting 3500 - 4500 genes with high variability and all proteins to form a normalized data matrix for each modality as the model input. The preprocessing process is all executed using the integrated functions of Scanpy.

[0020] Preferably, in step (3), the common features and unique details are input into the selective state space fusion module for integration together, and the integration formula is:

[0021] F = RNA + ADT

[0022] O R= RNA + LDC(RNA) + F * sigmoid(GAP(RNA - ADT))

[0023] O A = ADT + LDC(ADT) + F * sigmoid(GAP(ADT - RNA));

[0024] Among them, F is the feature of the rough fusion of two modalities, and O R is the common feature and unique details of the transcriptomic RNA, and O A is the common feature and unique details of the protein ADT. GAP is the global pooling operation, sigmoid is the activation function operation, LDC is the learnable descriptive convolution operation, RNA is the transcriptomic information, and ADT is the proteomic information.

[0025] Further preferably, in step (3), the omics feature fusion module first receives the output of each modality from the adaptive max pooling layer. Secondly, after fusing the two modalities, it inputs them together with each modality into the omics feature enhancement module. During the enhancement process, the difference features are obtained by subtracting elements from different modality features, the mapping of the difference features is enhanced, and then the difference features are merged with the original features, and the additional modality supplementary information is used to enrich the difference features. This process effectively extracts and amplifies the inherent common features and unique details in the omics, thereby improving the overall fusion performance. Finally, the enhanced common features and unique details of each modality are input into the selective state space fusion module for integration.

[0026] Preferably, in step (4), the selective state space fusion module receives the intermediate states of each modality output by the omics feature fusion module. The enhanced modalities are fused twice to obtain the mixed features, and the formula is:

[0027] F’ n = Li(LN(O R )) * Li(LN(O A ))

[0028] F n = F’ n + Li(LN(O R )) + Li(LN(O A ));

[0029] Among them, O R is the enhanced transcriptomic feature, O A is the enhanced proteomic feature, Li is the linear layer, LN is the normalization layer, and F n is the mixed feature.

[0030] Preferably, the hybrid features are integrated with the original modality and the channel attention mechanism is used to capture the global dependencies between different channels and reduce channel redundancy to obtain the integrated result. The formula is as follows:

[0031] O’ R = ECA(Li(Li(LN(O R )) * Mamba(F n )))

[0032] O’ A = ECA(Li(Li(LN(O A )) * Mamba(F n )));

[0033] where ECA is the channel attention mechanism, and O’ R is the integrated transcriptomics result, and O’ A is the integrated proteomics.

[0034] Preferably, in step (5), the overall network structure includes the training of the pre - self - masking encoder and the omics data operation.

[0035] More preferably, the training of the pre - self - masking encoder includes: first receiving the normalized data matrices of each modality obtained from step (2), respectively using the RNA self - masker and protein self - masker with the embedded linear selective state - space module to extract the features of each modality, and fusing the omics feature expressions through the omics feature fusion module;

[0036] During the fusion process, each omics feature first enters the omics feature enhancement module to enhance the feature expression of its own omics, and finally, after being input into the selective state - space fusion module (Cross Mamba), it is output to the decoder of its own omics to independently reconstruct the data of each modality, so as to learn general information through self - supervised learning as a pre - trained model.

[0037] More preferably, the omics data operation includes: fine - tuning the overall network structure so that the final result receives the information output from different levels of the selective state - space fusion module. The weights trained in the first stage can more accurately identify the differences in cell types. This process promotes different embeddings of each modality and the global embedding of cell combinations of the two modalities. Then, the clustering index is calculated according to the model clustering coordinates and classification metrics and the true labels, including ARI, NMI, FMI, ASW, AMI.

[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0039] The present invention proposes a multi-modal integration method for single-cell omics based on a linear selective state space, and applies a linear selective state space enhanced fusion network (sc3MAE) to integrate multi-omics data and cell classification tasks. It can achieve a good integration effect on transcriptome data and proteome data, retain unique information while extracting common features of multi-modal data. A new network structure is constructed to distinguish different types of cells, and the best comprehensive effect is achieved compared with existing methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic diagram of the selective state space enhanced fusion network structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In order to make those skilled in the art more clearly understand the technical solutions described in the present invention, the following embodiments are listed for illustration. It should be noted that the following embodiments do not limit the scope of protection required by the present invention.

[0042] Unless otherwise specified, the raw materials, reagents or devices used in the following embodiments can be obtained from conventional commercial channels or can be obtained by existing known methods.

[0043] Embodiment 1

[0044] A multi-modal integration method for single-cell omics based on a linear selective state space.

[0045] The selective state space enhanced fusion network structure (sc3MAE) is as Figure 1 shown.

[0046] It includes the following steps:

[0047] Collect the CITE-seq public datasets from the 10X Genomics platform, including the 10k peripheral blood mononuclear cell dataset (PBMC10k) from healthy donors, the 5k peripheral blood mononuclear cell dataset (PBMC5k) from healthy donors, and the 10k cell dataset (MALT10K) from MALT tumors. Additionally, collect the CITE-seq datasets (SPL111 and SPL206) measuring the spleen and lymph nodes from the literature written by Adam et al. as the original multi-omics data. The above five datasets are all stored in the github repository of Adam et al. (https: / / github.com / YosefLab / totalVI_reproducibility / tree / master / data). The details of the CITE-seq datasets are shown in Table 1. Perform quality control on the original multi-omics data, removing cells with less than 200 gene expressions, removing genes expressed in less than 3 cells, and removing cells with a mitochondrial gene proportion exceeding 20%:

[0048] Perform RPKM normalization on the processed transcriptome data, dividing the number of reads mapped to genes by the number of reads mapped to all genomes; perform CLR normalization on the proteome data, centering logarithmic ratio transformation to make the proteome data conform to the normal distribution rule; and divide it into a training set, a test set, and a validation set according to different stages. Specifically, in the first stage, use a training set:validation set = 3:1 to train the model; in the second stage, use 30% of the annotated data for fine-tuning to improve the pre-trained model. Select 3500 - 4500 genes with high variability and all proteins to form the normalized data matrix of each modality as the model input. The preprocessing process is all executed using the integrated functions of Scanpy.

[0049] Construct an omics feature fusion module to integrate the common features of multi-omics, and construct an omics feature enhancement module to retain and enhance the unique details of each omics; the omics feature fusion module first receives the output of each modality from the adaptive max pooling layer, and then fuses the two modalities and inputs them into the omics feature enhancement module together with each modality. During the enhancement process, the differential features are obtained by subtracting elements from different modality features, the mapping of the differential features is enhanced, and then the differential features are merged with the original features, and the differential features are enriched with additional modality supplementary information. This process effectively extracts and amplifies the inherent common features and unique details in the omics, thus improving the overall fusion performance. Finally, the common features O R and unique details O A of each enhanced modality are input into the selective state space fusion module for integration. The integration formula is:

[0050] F = RNA + ADT

[0051] O R = RNA + LDC(RNA) + F * sigmoid(GAP(RNA - ADT))

[0052] O A = ADT + LDC(ADT) + F * sigmoid(GAP(ADT - RNA));

[0053] Among them, F is the feature of the rough fusion of two modalities, and O R is the common feature and unique details of the transcriptome RNA, and O A is the common feature and unique details of the protein ADT. GAP is the global pooling operation, sigmoid is the activation function operation, LDC is the learnable descriptive convolution operation, RNA is the transcriptomics information, and ADT is the proteomics information.

[0054] Construct a selective state space fusion module to integrate the unique information and common features of each omics; the selective state space fusion module receives the intermediate states of each modality output by the omics feature fusion module, and the enhanced modality undergoes two fusions to obtain the mixed feature. The formula is:

[0055] Among them, O R is the feature of the enhanced transcriptomics, and O A is the feature of the enhanced proteomics. Li is the linear layer, LN is the normalization layer, and F n is the mixed feature.

[0056] The mixed feature is integrated with the original modality and captures the global dependencies between different channels through the channel attention mechanism, and reduces channel redundancy to obtain the integrated result. The formula is:

[0057] O' R = ECA(Li(Li(LN(O R )) * Mamba(F n )))

[0058] O' A = ECA(Li(iLi(LN(O A )) * Mamba(F n )));

[0059] Among them, ECA is through the channel attention mechanism, and O' R is the integrated transcriptomics result, and O' A is the integrated proteomics.

[0060] (5) Construct a selective state space enhanced fusion network, and concatenate each module to form the overall network structure;

[0061] The overall network structure includes the training of the pre - self - masking encoder and the omics data operation.

[0062] The training of the pre - self - masking encoder includes: first receiving the normalized data matrices of each modality, using the RNA self - masker and protein self - masker with embedded linear selective state - space modules to extract the features of each modality respectively, and fusing the omics feature expressions through the omics feature fusion module;

[0063] During the fusion process, each omics feature first enters the omics feature enhancement module to enhance the feature expression of its own omics. Finally, after being input into the selective state - space fusion module (Cross Mamba), it is output to the decoder of its own omics to independently reconstruct the data of each modality, so as to learn general information through self - supervised learning as a pre - trained model.

[0064] The omics data operation includes: fine - tuning the overall network structure so that the final result receives the information output from different levels of the selective state - space fusion module, and the weight trained in the first stage can accurately identify the differences in cell types. This process promotes different embeddings of each modality and the global embedding of cell combinations of the two modalities. Then, the clustering index is obtained by calculating with the model clustering coordinates and classification metrics and the true labels, including adjusted Rand index (ARI), normalized mutual information (NMI), Fowlkes - Mallows index (FMI), average silhouette width (ASW), adjusted mutual information (AMI). In the first stage, that is, the pre - training stage, sc3MAE includes two encoders, two decoders, an omics feature fusion module, and a selective state - space fusion module. During the training process, 20% of the token sequences are randomly masked, and then the un - masked vectors are input into the encoder of the model for learning. The encoded gene expression vectors and proteomics vectors are fused, and the complete token sequence is reconstructed through the decoder. When the mean square error loss function tends to be stable, the pre - trained model is saved. After the pre - training stage, sc3MAE discards one decoder architecture, receives the outputs from different levels of the selective state - space fusion module and enhances them with residuals as the final output of the model. When the cross - entropy loss function region is stable, the fine - tuned model is saved.

[0065] Cluster the cells and compare with the true labels to obtain the clustering performance. Perform clustering analysis on the final output of the model to obtain the final classification and clustering results of different CITE - seq datasets. All clustering and detection results are measured using the adjusted Rand index (ARI), normalized mutual information (NMI), and Fowlkes - Mallows index (FMI).

[0066] Table 1 Details of CITE - seq datasets

[0067] cell RNA protein SPL111 16828 13553 112 SPL206 15820 13553 209 PBMC5K 3994 16581 29 PBMC10K 6855 16727 14 MALT10K 6838 16659 14

[0068] Comparison of evaluation metrics:

[0069] As shown in Table 2, sc3MAE was benchmarked on five CITE-seq datasets: SPL111, SPL206, PBMC5K, PBMC10K, and MALT10K. Ten clustering evaluation metrics were used to evaluate each method multi-dimensionally. The metrics include Adjusted Rand Index (ARI), Normalized Mutual Information (NMI), Fowlkes-Mallows (FMI), F-score (F-measure), Adjusted Mutual Information (AMI), Average Silhouette Width (ASW), Jaccard Index (JI), Silhouette Coefficient (SC), Calinski-Harabasz Index (CHI), and Davies-Bouldin Index (DBI). The methods include TotalVI, CiteFuse, scMM, and SCOIT. Compared with the other four methods, sc3MAE maintains a leading edge in ARI, NMI, FMI, ASW, and AMI. Except for JI, CHI, and DBI, sc3MAE maintains a leading edge in the benchmark tests, indicating that sc3MAE has good generalization and robustness.

[0070] Table 2 Scores of ten evaluation metrics for five methods on each CITE-seq dataset

[0071]

[0072] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, any technical solutions obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art, such as any modifications, equivalent replacements, and improvements, shall fall within the protection scope determined by the claims.

Claims

1. A multimodal integration method for single-cell omics based on linear selective state space, characterized by: The following steps are involved: (1) Perform quality control on raw multi-omics data and remove low-quality samples; (2) normalize the processed multi-omics data and select highly variable genes; (3) Construct an omics feature fusion module to integrate the common features of multiple omics, and construct an omics feature enhancement module to retain and enhance the unique details of each omics; (4) construct a selective state space fusion module to integrate the unique information and common features of each omics; (5) Construct a selective state space enhanced fusion network and connect the modules in series to form an overall network structure; (6) Cluster the cells and compare them with the true labels to obtain the clustering performance.

2. The multimodal integration method according to claim 1, characterized in that: In step (1), the low-quality sample includes at least one of cells in which the number of gene expressions is less than 200, genes expressed in less than 3 cells, and cells in which mitochondrial genes account for more than 20%.

3. The multimodal integration method according to claim 1, characterized in that: In step (2), the normalization is to perform RPKM normalization on the transcriptome data and to perform CLR normalization on the proteome data.

4. The multimodal integration method according to claim 3, characterized in that: The RPKM normalization is to divide the number of reads mapped to the gene by the number of reads mapped to the entire genome.

5. The multimodal integration method according to claim 3, characterized in that: The CLR was normalized to the central logarithmic ratio transformation to make the proteomic data conform to the normal distribution rule.

6. The multimodal integration method according to claim 1, characterized in that: In step (2), the method for selecting highly variable genes includes selecting 3500 to 4500 genes with high variability and a normalized data matrix of each mode formed by all proteins as model input.

7. The multimodal integration method according to claim 1, characterized in that: In step (3), the common features and unique details are input into the selective state space fusion module for integration. The integration formula is: F=RNA+ADT O R =RNA+LDC(RNA)+F*sigmoid(GAP(RNA-ADT)) O A =ADT+LDC(ADT)+F*sigmoid(GAP(ADT-RNA)); Among them, F is the feature of the rough fusion of the two modes, O R The common features and unique details of transcriptome RNA, O A are the common features and unique details of protein ADT, GAP is the global pooling operation, sigmoid is the activation function operation, LDC is the learnable descriptive convolution operation, RNA is the transcriptomics information, and ADT is the proteomics information.

8. The multimodal integration method according to claim 1, characterized in that: In step (4), the selective state space fusion module receives the intermediate states of each modality output by the omics feature fusion module, and the enhanced modality is fused twice to obtain a mixed feature, and the formula is: Among them, O R For the enhanced transcriptomic features, O A is the enhanced proteomics feature, Li is the linear layer, LN is the normalization layer, F n It is a mixed feature.

9. The multimodal integration method according to claim 8, characterized in that: The hybrid features are integrated with the original modality and the global dependencies between different channels are captured through the channel attention mechanism, and channel redundancy is reduced to obtain the integrated result, which is: Among them, ECA is the channel attention mechanism, O′ R is the integrated transcriptomics results, O′ A For integrated proteomics.

10. The multimodal integration method according to claim 1, characterized in that: In step (5), the overall network structure includes the training of the early self-shielding encoder and the omics data operation.