Efficient integration method for single-cell multi-omics matching data

By constructing a graph structure model and graph convolutional network to optimize the embedding space, the problems of alignment accuracy and excessive consumption of computing resources in the integration of single-cell multi-omics data were solved, and efficient and accurate multi-omics data integration was achieved.

CN120690301APending Publication Date: 2025-09-23XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510739966.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

When integrating single-cell multi-omics matching data, existing technologies have problems such as low alignment accuracy, low accuracy of sparse cell subpopulation label migration, and excessive consumption of computing resources, making it difficult to accurately distinguish between real biological differences and technical deviations.

Method used

By constructing a graph structure model of transcriptome and chromatin accessibility, using graph convolutional networks (GCN) for information transmission, and combining contrastive learning to optimize the embedding space, efficient integration of transcriptome and chromatin accessibility data is achieved, reducing computational complexity and improving alignment accuracy.

Benefits of technology

It achieves efficient and accurate multi-omics data integration on large-scale data sets, significantly improves cross-modal alignment accuracy and label transfer accuracy, and reduces computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690301A_ABST
    Figure CN120690301A_ABST
Patent Text Reader

Abstract

The invention particularly relates to an efficient integration method for single-cell multi-omics matching data. The method comprises the following steps: acquiring the single-cell multi-omics matching data; wherein the single-cell multi-omics matching data comprises matched transcriptome data scRNA-seq and chromatin accessibility data scATAC-seq, and the single-cell multi-omics matching data comprises the transcriptome data scRNA-seq and the chromatin Converting peak data of the chromatin accessibility data into a gene activity matrix; performing low-quality filtering on the single-cell multi-omics matching data, deleting abnormal cells, and preprocessing a gene expression matrix of transcriptome data and a gene activity matrix of chromatin accessibility data; respectively constructing a first graph structure and a second graph structure for a gene expression matrix of the preprocessed transcriptome data and a gene activity matrix of the preprocessed chromatin accessibility data; constructing an integrated model of transcriptome and chromatin accessibility; and inputting the first graph structure and the second graph structure into an integration model for training to obtain integrated data. According to the method, biological characteristics are reserved to the maximum extent, the calculation scale is reduced, and the accurate and efficient multi-omics data integration method is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of single-cell multi-omics data integration, and in particular to a method for efficiently integrating single-cell multi-omics matching data. Background Art

[0002] Traditional single-modal sequencing data can only provide a one-sided understanding of cell status and function. This limitation fundamentally restricts our understanding of complex cell behavior and gene regulation mechanisms. With the development of single-cell sequencing technology, the resolution of sequencing data has reached an unprecedented level, which has greatly improved people's ability to analyze cell heterogeneity and made it possible to integrate multimodal single-cell data. There are complex connections and complex noises between different types of sequencing data from the same cell, and the heterogeneity between cells will also interfere with the identification of features. At present, there is a need for a multi-omics matching data integration method that can quickly and accurately integrate different sequencing data and retain biological characteristics. The following methods are currently available in the prior art:

[0003] In 2019, Ilya Korsunsky et al. proposed Harmony, a method for efficient batch correction and integration of single-cell data, in Nature Methods. This method is based on iterative soft clustering and linear correction models, and achieves integration by maximizing the similarity between data sets while retaining biological differences.

[0004] In 2021, Mika Sarkin Jain et al. proposed a method called MultiMap in Genome Biology. This technology jointly reduces the dimensionality of sequencing data and applies optimal transmission theory to achieve accurate mapping of different omics data in a shared low-dimensional space. By applying a cross-modal alignment strategy based on Wasserstein distance, accurate matching between cells can be achieved.

[0005] In 2022, Kai Cao et al. proposed a unified embedding model called uniPort in Nature Communications. Using a hybrid architecture of generative adversarial networks and variational autoencoders, uniPort maps different modalities into a shared latent space, enabling cross-omics cell type alignment, missing modality completion, and multimodal joint dimensionality reduction, significantly improving the robustness of cross-technical data integration.

[0006] In 2022, Lei Xiong et al. proposed a SCALEX alignment method in Nature Communications, which separates biological variation and batch effects through a decoupling learning strategy, effectively eliminating technical bias while retaining true biological differences, and mapping data of different modalities into a unified low-dimensional latent space, thereby achieving data alignment.

[0007] In 2024, Zhen-Hao Guo et al. proposed an algorithm called scCorrector in Briefings in Bioinformatics. This algorithm combines generative adversarial networks (GANs) with contrastive learning to construct a self-supervised learning framework for cell alignment. This method uses GAN adversarial training to eliminate technical noise while preserving meaningful biological differences through contrastive learning. Ultimately, this method achieves data integration, significantly reducing batch-to-batch variability while maintaining cellular heterogeneity.

[0008] Existing technologies such as MultiMap and SCALEX struggle to accurately distinguish true biological differences from technical bias when integrating multi-omics data, leading to misalignment of cell types across modalities and poor label migration performance in rare cell populations. Existing technologies such as uniPort and Harmony, when processing large-scale datasets, suffer from computational framework issues that significantly increase memory usage and runtime, resulting in low integration efficiency.

[0009] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0010] In order to solve the problems of low alignment accuracy, low accuracy of sparse cell subpopulation label migration, and excessive consumption of computing resources in the alignment process when integrating single-cell multi-omics matching data, the present invention provides an efficient integration method for single-cell multi-omics matching data, which retains biological features to the greatest extent and reduces the computing scale, thereby realizing an accurate and efficient multi-omics data integration method.

[0011] Other features and advantages of the present invention will become apparent from the following detailed description, or may be learned in part by practice of the present invention.

[0012] According to a first aspect of the present invention, a method for efficiently integrating single-cell multi-omics matching data is provided, the method comprising:

[0013] Obtain single-cell multi-omics matching data; where single-cell multi-omics matching data includes matching transcriptome data scRNA-seq and chromatin accessibility data scATAC-seq;

[0014] Convert peak data of chromatin accessibility data into gene activity matrix;

[0015] Low-quality single-cell multi-omics matching data were filtered to remove abnormal cells, and the gene expression matrix of transcriptome data and the gene activity matrix of chromatin accessibility data were preprocessed;

[0016] A first graph structure is constructed for the gene expression matrix of the preprocessed transcriptome data, and a second graph structure is constructed for the gene activity matrix of the preprocessed chromatin accessibility data;

[0017] Build integrated models of transcriptome and chromatin accessibility;

[0018] The first graph structure and the second graph structure are input into the integrated model for training. After the training, the first graph structure and the second graph structure are input into the trained model, and the obtained output is the integrated data.

[0019] In some exemplary embodiments, the peak data of the chromatin accessibility data is converted into a gene activity matrix using the following formula:

[0020]

[0021] Among them, X p,c represents the openness of peak p in cell c, d(p,TSS g ) represents the linear distance from peak p to TSS of gene g, P g represents the set of peaks associated with gene g.

[0022] In some exemplary embodiments, the low-quality filtering of the single-cell multi-omics matching data and the pre-processing of the gene expression matrix of the transcriptome data and the gene activity matrix of the chromatin accessibility data include:

[0023] Remove low-quality cells with less than 1,800 sequencing depths or abnormal cells with more than 100,000 sequencing depths in single-cell chromatin accessibility scATAC-seq data;

[0024] Remove low-quality cells with fewer than 1,000 unique molecular identifiers or abnormal cells with more than 25,000 unique molecular identifiers in single-cell transcriptome scRNA-seq data;

[0025] Deleting cells with abnormal chromatin structure whose nucleosome signal is greater than 2, and cells with poor chromatin openness whose transcription start site enrichment score is less than 1;

[0026] High-quality cells were filtered and the gene expression matrix of the filtered transcriptome data and the gene activity matrix of the chromatin accessibility data were normalized to eliminate the systematic bias caused by the difference in sequencing depth between different cells;

[0027] Perform logarithmic transformation on the normalized data to make the data distribution closer to normal distribution, which is convenient for subsequent stable data analysis;

[0028] The top 2000 highly variable genes were screened, which reduced unnecessary computational effort while retaining biological information to the greatest extent possible; the PCA method was used to perform linear dimensionality reduction using the screened highly variable genes as input features.

[0029] In some exemplary embodiments, constructing the first graph structure or constructing the second graph structure includes:

[0030] The reciprocal of the Euclidean distance between cells is calculated as the similarity, and the top 20 similar cells of each cell are taken as adjacent cells. The cell index is recorded as the edge and the cell as the point, and they are stored together as graph structure data.

[0031] In some exemplary embodiments, constructing an integrated model of transcriptome and chromatin accessibility comprises:

[0032] Establish a graph convolutional network (GCN encoder) consisting of a cascade of input layers with the dimension of the reduced data, the first and second hidden layers with 400 dimensions, and the third hidden layer with the dimension of the reduced data. All hidden layers use graph convolution layers followed by ReLU activation functions.

[0033] Establish a projector consisting of a cascade of input dimensions as the reduced data dimension, a first hidden layer with 400 dimensions, and a second hidden layer with the reduced data dimension. Each layer uses a fully connected layer, where the ReLU function is used as the activation function after the first hidden layer, and no activation function is set for the output layer.

[0034] Establish a contrastive loss function by calculating the normalized dot product similarity between low-dimensional embeddings and applying a temperature parameter of 0.1 for exponential scaling to generate a similarity matrix, and construct the negative log-likelihood loss of positive sample pairs based on the diagonal elements;

[0035] The GCN encoder and projector are cascaded to form a graph contrastive learning model. The above-mentioned contrastive loss function is used to optimize the embedding space. The Adam optimizer is used to optimize the network parameters during training, and feature alignment is achieved by maximizing the similarity of positive sample pairs.

[0036] In some exemplary embodiments, inputting the first graph structure and the second graph structure into the integration model for training to obtain integrated data includes:

[0037] Input the first graph structure and the second graph structure into the GCN encoder, and the loss function By calculating the similarity of the low-dimensional latent representations between matching place cells The parameters of the integrated model are adjusted until the value of the loss function drops to the minimum, that is, the cell similarity of the matching position is updated to the maximum, and the integrated potential representation is obtained. The integrated result is visualized by UMAP.

[0038] In some exemplary embodiments, the method further comprises:

[0039] The knn method is used on the integrated latent representation to calculate the distance between the RNA data output by the model and the ATAC data. The first 30 neighbor points of each data record are set, and the data are classified by majority voting to achieve the purpose of label transfer.

[0040] According to a second aspect of the present invention, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for efficiently integrating single-cell multi-omics matching data described in the first aspect is implemented.

[0041] According to a third aspect of the present invention, a computer program product is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for efficiently integrating single-cell multi-omics matching data described in the first aspect is implemented.

[0042] According to a fourth aspect of the present invention, there is provided an electronic device, comprising:

[0043] processor; and

[0044] a memory for storing executable instructions of the processor;

[0045] Wherein, the processor is configured to implement the efficient integration method of single-cell multi-omics matching data described in the first aspect above by executing the executable instructions.

[0046] The efficient integration method for single-cell multi-omics matching data provided by the present invention constructs a graph structure for the data and performs edge-based information transfer within a graph neural network. This enhances the recognition of cell types, enabling the integration to more accurately align similar cells and distinguish heterogeneous cells. Contrastive learning focuses on narrowing the distance between matching cells, reducing computational complexity and significantly reducing computing resources on large datasets.

[0047] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings are incorporated into and constitute a part of this specification, illustrate embodiments consistent with the present invention, and together with the description, serve to explain the principles of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and it is clear that those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0049] Figure 1 A schematic diagram of the overall process of integrating multi-omics matching data in the present invention;

[0050] Figure 2 This is a schematic diagram of the integration principle of the present invention;

[0051] Figure 3 Schematic diagram of the integration results of the present invention and the prior art in the PBMC10k dataset;

[0052] Figure 4 Schematic diagram of label migration results after integration of the present invention and prior art on the PBMC10k dataset;

[0053] Figure 5 Schematic diagram of resource usage during the PBMC10k dataset integration process according to the present invention and the prior art. DETAILED DESCRIPTION

[0054] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0055] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0056] In existing technologies, due to the significant differences in feature dimensions, data distribution, and biological significance of different omics data, existing methods have difficulty accurately capturing cross-modal correspondences between cells, which results in the inability to achieve a high degree of alignment of multi-omics data during the integration process. Existing algorithms overly rely on global similarity metrics and ignore key biological differences in the local neighborhood of cells, resulting in the heterogeneity of rare cell subpopulations being masked and difficulty in accurately migrating labels for sparse cell populations. Existing technologies rely on global cell relationship calculations, and when processing large-scale cell data, memory usage and computation time increase exponentially, resulting in excessive consumption of computing resources.

[0057] In response to the shortcomings and deficiencies of existing technologies, this example embodiment provides an efficient integration method for single-cell multi-omics matching data, which significantly reduces computational complexity while improving cross-modal alignment accuracy and label migration accuracy, thereby achieving efficient and accurate single-cell multi-omics data integration. Figure 1 As shown, the following steps may be specifically included:

[0058] Step S1, obtaining single-cell multi-omics matching data;

[0059] Step S2, converting the peak data of chromatin accessibility into a gene activity matrix;

[0060] Step S3: low-quality filtering of single-cell multi-omics matching data, deletion of abnormal cells, and preprocessing of the gene expression matrix of transcriptome data and the gene activity matrix of chromatin accessibility data;

[0061] Step S4, constructing a first graph structure for the gene expression matrix of the preprocessed transcriptome data, and constructing a second graph structure for the gene activity matrix of the preprocessed chromatin accessibility data;

[0062] Step S5, constructing a model integrating transcriptome and chromatin accessibility based on graph structure;

[0063] In step S6, the first graph structure and the second graph structure are input into the integrated model for training. After the training is completed, the first graph structure and the second graph structure are input into the trained model, and the obtained output is the integrated data.

[0064] Below, each step of the multi-omics matching data integration method in this example embodiment will be described in more detail with reference to the accompanying drawings and examples.

[0065] In step S1, single-cell multi-omics matching data are obtained.

[0066] For example, taking the human peripheral blood mononuclear cell PBMC10k dataset as an example, the matching transcriptome data scRNA-seq and chromatin accessibility data scATAC-seq are obtained from 10xGenomics.

[0067] In step S2, the peak data of chromatin accessibility are converted into a gene activity matrix.

[0068] For example, the GeneActivity method is used for the chromatin accessibility peaks to evaluate the potential transcriptional activity of genes based on chromatin openness, resulting in a matrix representing the activity of each gene in each cell. The gene activity matrix can better link chromatin openness with transcriptome genes, facilitating subsequent feature extraction. The construction process is as follows:

[0069]

[0070] Among them, X p,c represents the openness of peak p in cell c, d(p,TSS g ) represents the linear distance (kb) from peak p to TSS of gene g, P g represents the set of peaks associated with gene g.

[0071] In step S3, low-quality filtering is performed on the single-cell multi-omics matching data to remove abnormal cells, and the gene expression matrix of the transcriptome data and the gene activity matrix of the chromatin accessibility data are preprocessed.

[0072] To reduce noise and effectively capture cell features, low-quality multi-omics matching data is filtered and abnormal cells are deleted. This specifically includes the following sub-steps S3a-S3f:

[0073] Step S3a: Delete low-quality cells with less than 1,800 sequencing depths or abnormal cells with more than 100,000 sequencing depths in the single-cell chromatin accessibility scATAC-seq data.

[0074] Step S3b: Remove low-quality cells with fewer than 1,000 unique molecular identifiers or aberrant cells with more than 25,000 unique molecular identifiers from the single-cell transcriptome scRNA-seq data. This removes cells with abnormal chromatin structure, such as those with a nucleosome signal greater than 2, and cells with poor chromatin accessibility, such as those with a transcription start site enrichment score less than 1, to address cross-omics alignment bias caused by abnormal chromatin structure.

[0075] Step S3c, delete cells with abnormal chromatin structure whose nucleosome signal is greater than 2, and delete cells with poor chromatin openness whose transcription start site enrichment score is less than 1;

[0076] In step S3d, 11,209 cells containing 25,286 genes were filtered and the gene expression matrix of the filtered transcriptome data and the gene activity matrix of the chromatin accessibility data were normalized to eliminate the systematic bias caused by the difference in sequencing depth between different cells. The formula is:

[0077]

[0078] Among them, X raw,ij represents the raw count of gene j in cell i, Represents the total UMI count of cell i.

[0079] Step S3e: Perform logarithmic transformation on the normalized data to make the data distribution closer to normal distribution, which is convenient for subsequent stable data analysis. The formula is as follows:

[0080] Xlog,ij =log2(X norm,ij +1)

[0081] Among them, X norm,ij represents the normalized UMI count of gene j in cell i.

[0082] In step S3f, since stably expressed genes often fail to serve as distinguishing features of cellular heterogeneity, we screen the top 2,000 highly variable genes, minimizing unnecessary computational effort while preserving biological information to the greatest extent possible. We then use the PCA method to perform linear dimensionality reduction using these highly variable genes as input features.

[0083] In step S4, a first graph structure is constructed for the gene expression matrix of the preprocessed transcriptome data, and a second graph structure is constructed for the gene activity matrix of the preprocessed chromatin accessibility data.

[0084] The processed chromatin accessibility and transcriptome data are cells × genes matrices, representing the gene expression in each cell. In order to effectively retain biological information and capture the connection between cells for the subsequent neural network, the inverse of the Euclidean distance between cells is calculated for the two sets of data as the similarity ( Figure 2 Construct a similarity matrix), take the first 20 similar cells of each cell as adjacent cells, record the cell index as the edge, and the cell as the point, and store them together as graph structure data.

[0085] The similarity calculation formula is as follows:

[0086]

[0087] Among them, x i,k represents the distribution of the i-th cell data with dimension k, x j,k Represents the distribution of the j-th cell data with dimension k is the Euclidean distance between cell i and cell j. The Euclidean distance is converted into similarity by adding 1 and taking the inverse, and the range is limited to 0 to 1.

[0088] In step S5, a graph-based model integrating transcriptome and chromatin accessibility is constructed.

[0089] Specifically, it includes the following sub-steps S5a-S5d:

[0090] Step S5a, establish a graph convolutional network (GCN) encoder consisting of a cascade of input layer dimensions of the reduced data dimension, first and second hidden layers of 400 dimensions, and a third hidden layer of the reduced data dimension ( Figure 2 Shared GCN encoder), all hidden layers use graph convolutional layers (GCNConv) followed by ReLU activation functions;

[0091] Step S5b, establish a projector consisting of a cascade of the input dimension of the data after dimensionality reduction, the first hidden layer of 400 dimensions, and the second hidden layer of the data after dimensionality reduction ( Figure 2 Projection head), all layers use fully connected layers (Linear), the ReLU function is used as the activation function after the first hidden layer, and no activation function is set for the output layer;

[0092] Step S5c: Establish a contrast loss function. Generate a similarity matrix by calculating the normalized dot product similarity between low-dimensional embeddings and applying a temperature parameter of 0.1 for exponential scaling. Then construct the negative log-likelihood loss of positive sample pairs based on the diagonal elements. The formula is:

[0093]

[0094] in, and denote the representation of cell i in the transcriptome view and chromatin accessibility view, respectively, sim(·,·) represents the cosine similarity function, and τ is the temperature parameter (used to adjust the sharpness of the similarity distribution);

[0095] In step S5d, the GCN encoder and the projector are cascaded to form a graph contrastive learning model. The above contrastive loss function is used to optimize the embedding space. The Adam optimizer is used to optimize the network parameters during the training process, and feature alignment is achieved by maximizing the similarity of positive sample pairs.

[0096] In step S6, the first graph structure and the second graph structure are input into the integrated model for training. After the training is completed, the first graph structure and the second graph structure are input into the trained model, and the obtained output is the integrated data.

[0097] Specifically, it includes the following sub-steps S6a-S6b:

[0098] Step S6a: Input the graph structure of transcriptome RNA data and chromatin accessibility ATAC data constructed in step S4 into the GCN encoder, and the loss function By calculating the similarity of the low-dimensional latent representations between matching place cells The parameters of the integrated model are adjusted until the value of the loss function drops to the minimum, that is, the cell similarity of the matching position is updated to the maximum, and the integrated potential representation is obtained. The integrated result is visualized by UMAP.

[0099] In step S6b, the knn method (k-nearest neighbor classifier) ​​is used on the integrated latent representation to calculate the distance between the RNA data and the ATAC data output by the model. The first 30 nearest neighbors of each data record are set, and the data is classified by majority voting to achieve the purpose of label transfer. The prediction formula is:

[0100]

[0101] Among them, z A Low-dimensional embedding data representing chromatin accessibility, z R Represents the low-dimensional embedding data of the transcriptome, y is z R cell types, is the nearest neighbor set of chromatin accessibility data, Set for all cell types.

[0102] The best existing technologies, such as uniPort and Harmony, focus on global cell connections, which results in long runtimes, high memory usage, and the tendency to miss underlying cell features, leading to incomplete alignment. This paper constructs a graph structure for the data and performs edge-based information transfer within a graph neural network, enhancing the recognition of cell categories and enabling integration to more accurately align similar cells and distinguish heterogeneous cells. Contrastive learning focuses on closing the distance between matching cells, reducing the computational scale and significantly reducing computing resources on large datasets.

[0103] like Figure 2 As shown, the present invention obtains a graph structure through preprocessing such as gene activity matrix conversion, and inputs the graph structure into the constructed shared GCN encoder. The encoder transfers information between similar cells according to the edges in the graph structure to better identify cell heterogeneity. The two omics data are projected into the same shared low-dimensional space through the projection head. The loss function optimizes the model by calculating the similarity of matching cells and shortens the distance between matching cells to achieve integration.

[0104] The visualization diagram of the integration results based on the present invention and the prior art on the PBMC10k dataset is as follows: Figure 3 As shown, the different colors of the left picture of each part represent the transcriptome RNA and chromatin accessibility ATAC, and the different colors of the right picture represent the cell type label. The cell type label legend is marked in the last picture. It can be seen that the chromatin accessibility of the present invention is highly aligned with the transcriptome, the cluster edges are clear, and the consistency of cell types illustrates accurate label migration. Existing methods such as SCALEX and MultiMap only learn partial connections between matching cells and achieve a certain label transfer (cell type correspondence), but do not align the cells correctly. Although Harmony aligns the cells together, it does not accurately identify the heterogeneity between cells, resulting in different cell types mixed together. The better method uniPort clearly has unaligned parts and discrete edges.

[0105] According to the present invention and the prior art, the label migration results after the integration of the PBMC10k dataset are as follows: Figure 4As shown, the horizontal axis represents the real cell type, the vertical axis represents the predicted cell type, and the meaning of each square is the ratio of the real cell type predicted to the predicted cell type. The darker the square color, the larger the ratio. The average accuracy is calculated based on the weighted proportion of cells in the cell type. The present invention achieves high-accuracy predictions on most cell types and has the highest average accuracy, indicating that biological characteristics are retained and effective integration is achieved. For Harmony, uniPort and scCorrector, the accuracy is high in larger groups, but accurate migration cannot be achieved in sparse groups (the last 9 columns), while MultiMap and SCALEX only achieve accurate migration in a small number of cases.

[0106] According to the present invention and the prior art, the training time and memory usage on the PBMC10k dataset are as follows: Figure 5 As shown in Figure 2, it can be seen that the present invention consumes the least time and occupies relatively less memory. SCALEX and scCorrector, which occupy less memory, are obviously Figure 3 The lack of effective integration in the

[15] and its advantages cannot be demonstrated. Although uniPort occupies less memory, its runtime is as long as 36 minutes. In summary, the present invention achieves a fast and lightweight training process while maintaining efficient integration performance.

[0107] Traditional graph contrast learning is limited to single-modal augmented contrast, but this application addresses the challenge of heterogeneous data integration through RNA-ATAC dual-view contrast. This method demonstrates higher accuracy in integrating single-cell multi-omics matching data, with shorter training time and lower resource usage.

[0108] It should be noted that, as another aspect, the present application also provides a storage medium, which may be included in an electronic device or may exist independently without being incorporated into the electronic device. The storage medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device implements the method described in the following embodiments.

[0109] In one embodiment, the present application provides a computer program product, including a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.

[0110] Furthermore, the above-described figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above-described figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0111] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the invention herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the claims.

[0112] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings and that various modifications and variations can be made without departing from the scope thereof, which is limited only by the appended claims.

Claims

1. An efficient integration method for single-cell multi-omics matching data, characterized by: The method comprises: Obtain single-cell multi-omics matching data; where single-cell multi-omics matching data includes matching transcriptome data scRNA-seq and chromatin accessibility data scATAC-seq; Convert peak data of chromatin accessibility data into gene activity matrix; Low-quality single-cell multi-omics matching data were filtered to remove abnormal cells, and the gene expression matrix of transcriptome data and the gene activity matrix of chromatin accessibility data were preprocessed; A first graph structure is constructed for the gene expression matrix of the preprocessed transcriptome data, and a second graph structure is constructed for the gene activity matrix of the preprocessed chromatin accessibility data; Build integrated models of transcriptome and chromatin accessibility; The first graph structure and the second graph structure are input into the integrated model for training. After the training, the first graph structure and the second graph structure are input into the trained model, and the obtained output is the integrated data.

2. The method according to claim 1, characterized in that The peak data of chromatin accessibility data are converted into gene activity matrix using the following formula: Among them, X p,c represents the openness of peak p in cell c, d(p,TSS g ) represents the linear distance from peak p to TSS of gene g, P g represents the set of peaks associated with gene g.

3. The method according to claim 1, characterized in that The low-quality filtering of single-cell multi-omics matching data and the preprocessing of the gene expression matrix of transcriptome data and the gene activity matrix of chromatin accessibility data are described, including: Remove low-quality cells with less than 1,800 sequencing depths or abnormal cells with more than 100,000 sequencing depths in single-cell chromatin accessibility scATAC-seq data; Remove low-quality cells with fewer than 1,000 unique molecular identifiers or abnormal cells with more than 25,000 unique molecular identifiers in single-cell transcriptome scRNA-seq data; Deleting cells with abnormal chromatin structure whose nucleosome signal is greater than 2, and cells with poor chromatin openness whose transcription start site enrichment score is less than 1; High-quality cells were filtered and the gene expression matrix of the filtered transcriptome data and the gene activity matrix of the chromatin accessibility data were normalized to eliminate the systematic bias caused by the difference in sequencing depth between different cells; Perform logarithmic transformation on the normalized data to make the data distribution closer to normal distribution, which is convenient for subsequent stable data analysis; The top 2000 highly variable genes were screened, which reduced unnecessary computational effort while retaining biological information to the greatest extent possible; the PCA method was used to perform linear dimensionality reduction using the screened highly variable genes as input features.

4. The method according to claim 1, wherein The constructing of the first graph structure or the constructing of the second graph structure includes: The reciprocal of the Euclidean distance between cells is calculated as the similarity, and the top 20 similar cells of each cell are taken as adjacent cells. The cell index is recorded as the edge and the cell as the point, and they are stored together as graph structure data.

5. The method according to claim 1, wherein The method of constructing an integrated model of transcriptome and chromatin accessibility includes: Establish a graph convolutional network (GCN encoder) consisting of a cascade of input layers with the dimension of the reduced data, the first and second hidden layers with 400 dimensions, and the third hidden layer with the dimension of the reduced data. All hidden layers use graph convolution layers followed by ReLU activation functions. Establish a projector consisting of a cascade of input dimensions as the reduced data dimension, a first hidden layer with 400 dimensions, and a second hidden layer with the reduced data dimension. Each layer uses a fully connected layer, where the ReLU function is used as the activation function after the first hidden layer, and no activation function is set for the output layer. Establish a contrastive loss function by calculating the normalized dot product similarity between low-dimensional embeddings and applying a temperature parameter of 0.1 for exponential scaling to generate a similarity matrix, and construct the negative log-likelihood loss of positive sample pairs based on the diagonal elements; The GCN encoder and projector are cascaded to form a graph contrastive learning model. The above-mentioned contrastive loss function is used to optimize the embedding space. The Adam optimizer is used to optimize the network parameters during training, and feature alignment is achieved by maximizing the similarity of positive sample pairs.

6. The method according to claim 1, wherein The step of inputting the first graph structure and the second graph structure into the integration model for training to obtain integrated data includes: Input the first graph structure and the second graph structure into the GCN encoder, and the loss function By calculating the similarity of the low-dimensional latent representations between matching place cells The parameters of the integrated model are adjusted until the value of the loss function drops to the minimum, that is, the cell similarity of the matching position is updated to the maximum, and the integrated potential representation is obtained. The integrated result is visualized by UMAP.

7. The method according to claim 6, characterized in that The method further comprises: The knn method is used on the integrated latent representation to calculate the distance between the RNA data output by the model and the ATAC data. The first 30 neighbor points of each data record are set, and the data are classified by majority voting to achieve the purpose of label transfer.

8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for efficiently integrating single-cell multi-omics matching data according to any one of claims 1 to 7 is implemented.

9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for efficiently integrating single-cell multi-omics matching data according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the method for efficiently integrating single-cell multi-omics matching data according to any one of claims 1 to 7 by executing the executable instructions.