Homogenous and heterogenous data fusion method and system based on sparse subspace clustering and graph convolution filter, electronic equipment and storage medium

Through the homogeneous heterologous data fusion method based on sparse subspace clustering and graph convolution filter, the batch effect correction method in the prior art is solved in terms of applicability, robustness and computational efficiency, and the effect of efficiently removing batch effect and retaining heterogeneity is achieved.

CN120012023APending Publication Date: 2025-05-16SHANSHU TECH (BEIJING) CO LTD +5
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510175472.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing batch effect correction methods have shortcomings in applicability, robustness and computational efficiency, especially when processing high-dimensional single-cell data, it is difficult to effectively remove batch effects.

Method used

The homogeneous and heterologous data fusion method based on sparse subspace clustering and graph convolution filter is adopted. By constructing the homogeneous data source internal graph of each source and the homogeneous data source inter-source graph between each two sources, it is integrated into a global graph, and smoothing and dimensionality reduction operations are performed through graph convolution filters to remove batch effects.

Benefits of technology

This method avoids the limitation of linear assumptions, can preserve heterogeneity among cell types while batch correction, has strong applicability, can capture more global and local features, reduce computing resource consumption, and significantly improve clustering accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012023A_ABST
    Figure CN120012023A_ABST
Patent Text Reader

Abstract

The invention discloses a homogeneous and heterogeneous data fusion method and system based on sparse subspace clustering and a graph convolution filter, electronic equipment and a storage medium, and relates to the field of data processing. The method comprises the following steps: based on homogeneous data of each source, constructing a homogeneous data source inner graph of each source through a sparse subspace clustering method; based on the homogeneous data of every two sources, constructing a homogeneous data inter-source graph between every two sources; integrating all the homogeneous data source inner graphs and all the homogeneous data source inter-graphs, and constructing a homogeneous data global graph; and carrying out smoothing and dimension reduction operation on the global graph of the homogeneous data through a graph convolution filter, removing batch effects of the homogeneous data from different sources, and completing homogeneous and heterogeneous data fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to a method, system, electronic device and storage medium for fusion of homogeneous and heterogeneous data based on sparse subspace clustering and graph convolution filter. Background Art

[0002] "Homogeneous heterogeneous data" refers to data with the same content nature but different sources. Single-cell data is a typical homogeneous heterogeneous data, which usually faces the problem of batch effect. Batch effect is a technical deviation caused by different sample sources, experimental operations, sequencing platforms or technical parameters. This effect may lead to significant differences in the gene expression distribution of the same cell type in different batches, thereby interfering with data integration and the accuracy of clustering results. Especially in multi-center collaborative research or cross-laboratory data analysis, batch effect is a key issue that needs to be solved urgently. In order to solve the impact of batch effect on single-cell data clustering, there are currently a variety of correction methods.

[0003] Existing linear methods such as ComBat1 and Limma2 rely on the assumption that batch effects are independent of biological signals and can handle linear batch effects. However, Combat and limma are based on the assumption of linear models and are difficult to handle complex nonlinear batch effects, especially in high-dimensional single-cell data. In addition, Combat and limma require prior batch labels and assume that the biological distribution within all sources is consistent, which may not be true in single-cell data with large heterogeneity. Combat and limma focus on batch correction and may ignore biological differences between cell types, resulting in information loss.

[0004] Tools based on nonlinear methods (such as MNN (Mutual Nearest Neighbor)3 and the integration function of Seurat4) can improve the correction effect on complex cell distributions. However, the results of MNN matching rely on neighbor relationships and may ignore global structures, resulting in insufficient correction of batch effects. Seurat anchor point selection has a great influence on the effect of batch integration, and parameter adjustment requires more experience. Some biological signals may be lost during batch integration, and the retention effect on rare cell types is limited.

[0005] Deep learning methods (such as scVI5 and DESC6) have introduced a more flexible modeling framework. However, the results of the scVI method are less interpretable, and it is difficult to clearly explain how the model corrects batch effects or retains biological signals. The model performance depends on the network structure and hyperparameter settings, and the adjustment is complex. DESC has high computational complexity, and the training time and resource consumption are large.

[0006] In summary, existing batch effect correction methods have deficiencies in applicability, robustness, and computational efficiency. Summary of the invention

[0007] The embodiments of the present application provide a homogeneous and heterogeneous data fusion method, system, electronic device and storage medium based on sparse subspace clustering and graph convolution filter, which solve the problems in the prior art that batch effect correction methods have insufficient applicability, robustness and computational efficiency.

[0008] On the one hand, this embodiment provides a method for fusing homogeneous and heterogeneous data based on sparse subspace clustering and graph convolution filter, the steps of which are as follows:

[0009] Based on the homogeneous data of each source, a homogeneous data source intra-graph of each source is constructed by sparse subspace clustering method;

[0010] Based on the homogeneous data from every two sources, a homogeneous data source graph between every two sources is constructed;

[0011] Integrate all graphs within and between homogeneous data sources to build a global graph of homogeneous data;

[0012] The global graph of homogeneous data is smoothed and reduced in dimension in sequence through graph convolution filters to remove the batch effect of homogeneous data from different sources, thereby completing the fusion of homogeneous and heterogeneous data.

[0013] On the one hand, this embodiment provides a homogeneous heterogeneous data fusion system based on sparse subspace clustering and graph convolution filter, including:

[0014] A source internal graph construction module is used to construct a homogeneous data source internal graph of each source through a sparse subspace clustering method based on the homogeneous data of each source;

[0015] An inter-source graph construction module is used to construct an inter-source graph of homogeneous data sources between every two sources based on homogeneous data of every two sources;

[0016] The global graph construction module is used to integrate all the intra-graphs and inter-graphs of homogeneous data sources through graph convolution filters to construct a global graph of homogeneous data;

[0017] The smoothing and dimensionality reduction module is used to smooth and reduce the dimensionality of the global graph of homogeneous data through graph convolution filters, remove the batch effect of homogeneous data from different sources, and complete the fusion of homogeneous and heterogeneous data.

[0018] In a possible embodiment, constructing a homogeneous data source internal graph for each source includes:

[0019] Cluster homogeneous data based on sparse subspace clustering method and construct similarity matrix within homogeneous data source;

[0020] Based on the Frobenius norm of the difference between the homogeneous data expression matrix and the corresponding homogeneous data reconstruction matrix, a reconstruction error term is constructed; the homogeneous data reconstruction matrix is ​​a matrix obtained by linearly transforming the homogeneous data expression matrix through the similarity matrix within the homogeneous data source; the homogeneous data expression matrix is ​​a matrix representation of homogeneous data;

[0021] Through the sparsity regularization coefficient and the similarity matrix L within the homogeneous data source 1 norm, constructing a sparse regularization term;

[0022] Taking minimization of the sum of the reconstruction error term and the sparse regularization term as the optimization goal, and imposing a constraint that the diagonal elements of the similarity matrix within the homogeneous data source are zero during the optimization process, the first optimization formula is constructed;

[0023] The first optimization formula is solved to obtain the homogeneous data source intra-graph and the corresponding homogeneous data source intra-similarity matrix.

[0024] In a possible embodiment, constructing a homogeneous data source graph between every two sources includes:

[0025] Based on the mutual nearest neighbor method, the homogeneous data pairs between the two sources that satisfy the mutual nearest neighbor relationship are taken as anchor data pairs to obtain the anchor data pair set;

[0026] Based on the set of anchor data pairs, an unweighted graph between homogeneous data sources and a corresponding similarity matrix between homogeneous data sources are constructed; in the graph between homogeneous data sources, there are edge connections between anchor data, and there are no edge connections between non-anchor data.

[0027] In a possible embodiment, constructing a global graph of homogeneous data includes:

[0028] The similarity matrices within all homogeneous data sources are concatenated with the similarity matrices between all homogeneous data sources to obtain the homogeneous data global similarity matrix and the corresponding homogeneous data global graph.

[0029] In a possible embodiment, constructing a global graph of homogeneous data further includes:

[0030] In the unweighted graph between homogeneous data sources, a weight is assigned to each edge, and a weighted graph between homogeneous data sources and a corresponding similarity matrix between homogeneous data sources are constructed.

[0031] In a possible embodiment, smoothing the homogeneous data global graph includes:

[0032] Input features are input into the graph convolution filter to obtain output features, and then the global graph of homogeneous data is smoothed; the input features are homogeneous data expression matrices;

[0033] The graph convolution filter performs multi-layer convolution on the input features based on the simple graph convolution method, and minimizes the difference between the output features and the input features after convolution as the optimization goal; the number of layers of convolution on the input features is equal to the power of the normalized adjacency matrix of homogeneous data; the normalized adjacency matrix of homogeneous data is obtained based on the normalization of the global similarity matrix of homogeneous data.

[0034] In a possible embodiment, dimension reduction is performed on a global graph of homogeneous data, including:

[0035] The output features are made equal to the input features, and the output features are obtained as the features after dimensionality reduction, thereby reducing the dimensionality of the global graph of homogeneous data.

[0036] On the one hand, this embodiment provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes any of the above-mentioned homogeneous and heterogeneous data fusion methods based on sparse subspace clustering and graph convolution filters.

[0037] On the one hand, this embodiment provides a computer-readable storage medium, including a program code. When the storage medium is run on an electronic device, the program code is used to enable the electronic device to execute any of the above-mentioned homogeneous and heterogeneous data fusion methods based on sparse subspace clustering and graph convolution filters.

[0038] On the one hand, an embodiment of the present application provides a computer program product, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; when a processor of an electronic device reads the computer instructions from the computer-readable storage medium, the processor executes the computer instructions, so that the electronic device executes any of the above-mentioned homogeneous and heterogeneous data fusion methods based on sparse subspace clustering and graph convolution filters.

[0039] The beneficial effects of this application are as follows:

[0040] The homogeneous and heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter provided in this application overcomes the shortcomings of batch effect correction methods in terms of applicability, robustness and computational efficiency.

[0041] 1. It avoids the limitation of linear assumptions. At the same time, it can retain heterogeneity between data such as cell types while performing batch correction, which is highly applicable.

[0042] 2. Ability to capture more global and local features within / outside a source.

[0043] 3. It can provide explainability for the data relationship between two sources and within one source, while reducing the consumption of computing resources.

[0044] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0046] Figure 1 This is a schematic diagram of an application scenario in an embodiment of the present application;

[0047] Figure 2 This is a flowchart of an implementation method of a homogeneous heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter in an embodiment of the present application;

[0048] Figure 3 It is a structural schematic diagram of a homogeneous heterogeneous data fusion system based on sparse subspace clustering and graph convolution filter in an embodiment of the present application;

[0049] Figure 4 The present invention is a schematic diagram of a hardware structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be arbitrarily combined with each other. In addition, although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that here.

[0051] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the invention described herein can be implemented in sequences other than those illustrated or described herein.

[0052] The following is a brief introduction to the design concept of the embodiment of the present application:

[0053] The embodiment of the present application provides a method, system, electronic device and storage medium for fusion of homogeneous and heterogeneous data based on sparse subspace clustering and graph convolution filter. Based on the homogeneous data of each source, the homogeneous data source inner graph of each source is constructed by sparse subspace clustering method; based on the homogeneous data of every two sources, the homogeneous data source inter-graph between every two sources is constructed; all homogeneous data source inner graphs and homogeneous data source inter-graphs are integrated to construct a homogeneous data global graph; the homogeneous data global graph is smoothed and dimensionally reduced by graph convolution filter, the batch effect of homogeneous data from different sources is removed, and the homogeneous and heterogeneous data fusion is completed. The shortcomings of the batch effect correction method in terms of applicability, robustness and computational efficiency are overcome. The limitation of linear assumptions is avoided. At the same time, the heterogeneity of data such as cell types can be retained while batch correction, and the applicability is strong. More global and local features can be captured inside and outside the source. It can provide interpretability for the relationship between sources and within sources, while reducing the consumption of computing resources.

[0054] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. In addition, the embodiments and features in the embodiments of the present application may be combined with each other if there is no conflict.

[0055] like Figure 1 , which is a schematic diagram of an application scenario provided by an embodiment of the present application. In the schematic diagram of the application scenario, a terminal device 101 and a server 102 are included. The terminal device 101 and the server 102 communicate with each other through a communication network.

[0056] The terminal device 101 is an electronic device used by the target object, and the electronic device may be a personal computer, a mobile phone, a tablet computer, a notebook, an e-book reader, a vehicle-mounted terminal, etc. In addition, a client related to the analysis of the natural gas transmission system may be installed on the terminal device 101, and the client may be software (for example, an APP, a browser, etc.), or a web page, a small program, etc. The target object may use the above-mentioned client related to the analysis of the natural gas transmission system through the terminal device 101 to perform operations related to the analysis of the natural gas transmission system.

[0057] The server 102 may be an independent physical server or an edge device 102 in the field of cloud computing. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, cloud functions, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (English name: Content Delivery Network, abbreviated as CDN), as well as big data and artificial intelligence platforms.

[0058] There is no restriction on the number of the terminal devices 101 and / or servers 102 .

[0059] It should be noted that the homogeneous heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter in the embodiment of the present application can be executed by the terminal device 101 or the server 102 alone, or can be executed jointly by the terminal device 101 and the server 102.

[0060] In combination with the above-mentioned application scenarios, the homogeneous and heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter provided by the exemplary embodiment of the present application is described below with reference to the accompanying drawings. It should be noted that the above-mentioned application scenarios are only shown to facilitate the understanding of the spirit and principles of the present application, and the implementation methods of the present application are not limited in this respect.

[0061] refer to Figure 2 , is an implementation flow chart of a homogeneous heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter provided in an embodiment of the present application. Here, the server is used as the execution subject for introduction. The specific implementation process of the method is as follows:

[0062] S201, based on the homogeneous data of each source, constructing a homogeneous data source internal graph of each source by a sparse subspace clustering method;

[0063] S202, based on the homogeneous data from every two sources, construct a homogeneous data source graph between every two sources;

[0064] S203, integrating all intra-graphs and inter-graphs of homogeneous data sources to construct a global graph of homogeneous data;

[0065] S204, smoothing and dimension reduction operations are performed on the global graph of homogeneous data through graph convolution filters to remove the batch effect of homogeneous data from different sources, thereby completing the fusion of homogeneous and heterogeneous data.

[0066] In the embodiment of the present application, a complete process for removing the batch effect of homogeneous data is proposed. Homogeneous data can be single-cell data, and heterogeneous data are different batches of single-cell samples, that is, different sources. Therefore, homogeneous and heterogeneous data are single-cell data from different batches. Each source corresponds to each batch, and the data between each two sources corresponds to the data between each two batches.

[0067] The process of removing batch effects of homogeneous data includes the construction of graphs within homogeneous data sources, the construction of graphs between homogeneous data sources, the construction of global graphs of homogeneous data, and the dimension reduction of homogeneous data features. Finally, the homogeneous data with batch effects removed can significantly improve the accuracy and robustness of clustering.

[0068] The method used in this embodiment can capture more global and local features within and outside the batch; provide explainability for the relationship between and within batches, while reducing the consumption of computing resources; avoid the limitations of linear assumptions; and retain the heterogeneity between cell types while performing batch correction, and has strong applicability.

[0069] Further, executing S201, constructing a homogeneous data source internal graph of each source, including:

[0070] Cluster homogeneous data based on sparse subspace clustering method and construct similarity matrix within homogeneous data source;

[0071] Based on the Frobenius norm of the difference between the homogeneous data expression matrix and the corresponding homogeneous data reconstruction matrix, a reconstruction error term is constructed; the homogeneous data reconstruction matrix is ​​a matrix obtained by linearly transforming the homogeneous data expression matrix through the similarity matrix within the homogeneous data source; the homogeneous data expression matrix is ​​a matrix representation of homogeneous data;

[0072] Through the sparsity regularization coefficient and the similarity matrix L within the homogeneous data source 1 norm, constructing a sparse regularization term;

[0073] Taking minimization of the sum of the reconstruction error term and the sparse regularization term as the optimization goal, and imposing a constraint that the diagonal elements of the similarity matrix within the homogeneous data source are zero during the optimization process, the first optimization formula is constructed;

[0074] The first optimization formula is solved to obtain a homogeneous data source intra-similarity matrix corresponding to the homogeneous data source intra-graph.

[0075] In this embodiment, a homogeneous data source internal graph of each source is constructed mainly based on sparse subspace clustering.

[0076] Specifically, the single-cell batch intra-graph (homogeneous data source intra-graph) in each batch can be constructed based on sparse subspace clustering to generate a sparse single-cell batch intra-similarity matrix (homogeneous data source intra-similarity matrix) S I, which can efficiently identify cell clustering relationships within a batch.

[0077] The optimization goal of sparse subspace clustering is:

[0078]

[0079] in, is the expression matrix of single cells in batch I (homogeneous data expression matrix), n I is the number of cells, d is the number of gene features; S I is a sparse similarity matrix; X I S I is the single-cell data reconstruction matrix (homogeneous data reconstruction matrix); λ is the regularization coefficient, which is used to control the sparsity of the similarity matrix; ∥·∥ F represents the Frobenius norm, ∥·∥ 1 Indicates L 1 Norm. In the objective function Represents the reconstruction error of the data, which belongs to the quadratic form, λ∥S I ∥ 1 is a regularization used to introduce sparsity. In the constraint, diag(S I )=0, the diagonal elements are forced to 0 to prevent autoregression (i.e. a point should not model itself).

[0080] The above equation is a convex optimization problem with sparse regularization, which can be solved using the block coordinate descent method with global convergence. I Different columns (or rows) of are independent. Each time we optimize one variable, fix other variables, and continuously reduce the objective function value to converge to the global optimal solution through iteration. The specific process is as follows:

[0081] 1. Variable decomposition:

[0082] S I Decompose into S by columns I =[s I1 ,s I2 ,…,s IN ], where s Ij YesS I The jth column (corresponding to the jth cell) is updated column by column during optimization.

[0083] 2. Optimization problem decomposition:

[0084] For the jth column s Ij , the optimization objective becomes:

[0085]

[0086] where x Ij Yes XI The jth column of . At this time, the constraint s Ijj = 0, we can directly remove the autoregressive term (that is, the corresponding x Ij From X I (removed from the

[0087] 3. Optimize calculation:

[0088] For each column s Ij For optimization, the fast iterative shrinkage threshold algorithm (FISTA) can be used. The specific method is as follows:

[0089] 3.1. Define the gradient of the sub-problem objective function:

[0090]

[0091] 3.2. Gradient descent update:

[0092]

[0093] 3.3. Apply soft threshold operation:

[0094]

[0095] Among them, the soft threshold operation is defined as:

[0096] soft(z,τ)=sign(z)·max(|z|-τ,0)

[0097] 4. Return to S I As a result, we get the similarity matrix S within the single cell batch I .

[0098] Similarity matrix S within a single cell batch I The corresponding single-cell batch intra-graph is a weighted graph, and its weight is the calculated single-cell batch intra-similarities of two single cells in the batch. Alternatively, based on a preset single-cell batch intra-similarity threshold, only the single-cell batch intra-similarities greater than the single-cell batch intra-similarity threshold are retained as the weights in the single-cell batch intra-graph.

[0099] Further, executing S202, constructing a homogeneous data source graph between every two sources, including:

[0100] Based on the mutual nearest neighbor method, the homogeneous data pairs between the two sources that satisfy the mutual nearest neighbor relationship are taken as anchor data pairs to obtain the anchor data pair set;

[0101] Based on the set of anchor data pairs, an unweighted graph between homogeneous data sources and a corresponding similarity matrix between homogeneous data sources are constructed; in the graph between homogeneous data sources, there are edge connections between anchor data, and there are no edge connections between non-anchor data.

[0102] In this implementation, a graph construction method between homogeneous data sources is mainly based on mutual nearest neighbors.

[0103] Specifically, it can be based on mutual nearest neighbors and effectively capture the correlation between single-cell data between different batches through anchor point matching, and construct a similarity matrix between single-cell batches (similarity matrix between homogeneous data sources) and a graph between single-cell batches (graph between homogeneous data sources).

[0104] For two batches and The calculation process of mutual nearest neighbors (MNN) can be described by the following formula flow:

[0105] 1. Define the nearest neighbors:

[0106] First, for each batch of single-cell data, the nearest neighbor relationship between two batches is calculated. Let x i ∈X I and x j ∈X J These are the single-cell expression vectors in batches I and J, respectively.

[0107] (1) Nearest neighbor search:

[0108] x i ∈X I , find its J Nearest neighbors in :

[0109]

[0110] Similarly, for x j ∈X J , find its I Nearest neighbors in :

[0111]

[0112] (2) Expanded to k nearest neighbors

[0113] If you need to expand to k-nearest neighbors, as follows:

[0114] kNN I (x i ):x i In X J The k nearest neighbors in .

[0115] kNN J (x j ):x j In X I The k nearest neighbors in .

[0116] 2. Define mutual nearest neighbors:

[0117] The core idea of ​​MNN is that only when x i is x j The nearest neighbor of x j Also x i When the nearest neighbor of i and x j are considered anchor cell pairs.

[0118] (1) Basic Definition:

[0119] The mutual nearest neighbor pair is defined as:

[0120] MNN(X I ,X J )={(x i ,x j )∣∣x j ∈NN I (x i )andx i ∈NN J (x j )}

[0121] (2) Extended to k nearest neighbors:

[0122] If k nearest neighbor pairs are defined as:

[0123] MNN k (X I ,X J )={(x i ,x j )∣∣x j ∈kNN I (x i )andx i ∈kNN J (x j )}

[0124] 3. Calculate the distance matrix between batches

[0125] In order to improve the matching efficiency, we can first calculate the batch X I and X J The pairwise distance matrix of

[0126] D ij =∥x i -x j ∥ 2

[0127] Then use the distance matrix to quickly find the nearest neighbor:

[0128] For each i or j, just find the minimum value in the row or column of matrix D.

[0129] 4. Return the MNN result:

[0130] The final output result is a set of anchor cell pairs (anchor data pairs) based on MNN:

[0131] MNN(X I ,X J )={(i,j)|i∈{1,…,n I},j∈{1,…,n J},mutual nearest}

[0132] The single-cell batch similarity matrix A corresponding to the single-cell batch graph inter , then it is defined as:

[0133]

[0134] The above formula indicates that in the single-cell batch graph, there are edge connections between anchor cells (weight value is 1), and there are no edge connections between non-anchor cells (weight value is 0). Since the weight value of the edge between anchor cells is 1, the single-cell batch graph is an unweighted graph.

[0135] Among them, the MNN process identifies anchor cell pairs across batches through nearest neighbor calculation (or k-nearest neighbor expansion), providing biologically relevant contact points for batch effect correction and subsequent integrated analysis. This method can be accelerated using efficient nearest neighbor search algorithms (such as KD trees or Ball trees).

[0136] In some embodiments, other schemes may be used to replace sparse subspace clustering (SSC) and mutual nearest neighbor (MNN) to measure the similarity of single cells within a batch and between batches, for example:

[0137] 1. The similarity measurement of kernel function, through Gaussian kernel, Laplace kernel or multi-kernel learning method, calculates the nonlinear similarity matrix between cells, which can replace the similarity calculation based on linear reconstruction in sparse subspace clustering and generate the similarity matrix within a single cell batch.

[0138] 2. Simple measurement of Euclidean distance or cosine similarity. Directly use low-dimensional representation (such as features after PCA dimensionality reduction) to calculate Euclidean distance or cosine similarity, which can be used to generate similarity matrices within single-cell batches and between single-cell batches.

[0139] 3. The probabilistic graphical model method uses methods such as Gaussian mixture model (GMM) or hidden Markov model (HMM) to construct a relationship network between cells. The similarity matrix between single-cell batches can be generated by replacing the MNN method with probability as the similarity indicator.

[0140] In some embodiments, other schemes may be used to replace the sparse regularization method to optimize the intra-graph similarity matrix of a single-cell batch, for example:

[0141] 1. Distributed optimization: split the optimization objective of sparse subspace clustering into multiple small-scale sub-problems, and achieve parallel solution through distributed computing to improve solution efficiency.

[0142] 2. Graph regularization embedding, by introducing graph regularization constraints (such as Laplace regularization) to optimize the global graph matrix A, to replace the sparse regularization optimization method:

[0143]

[0144] where L is the graph Laplacian matrix.

[0145] 3. Non-convex optimization method: Introduce non-convex optimization models (such as ADMM or greedy optimization) to replace the standard L in subspace clustering 1 -Regularized solution to improve the adaptability to complex data distribution.

[0146] 4. Deep generative model optimization: Convert the graph construction optimization problem into the loss optimization of the deep generative model (such as variational autoencoder), and remove the batch effect through joint distribution learning.

[0147] Further, executing S203, constructing a global graph of homogeneous data, includes:

[0148] The similarity matrices within all homogeneous data sources are concatenated with the similarity matrices between all homogeneous data sources to obtain the homogeneous data global similarity matrix and the corresponding homogeneous data global graph.

[0149] In this implementation, a graph construction method is mainly based on graph aggregation, which integrates the graph within the homogeneous data source and the graph between homogeneous data sources, constructs a global similarity matrix of homogeneous data, and realizes a unified representation of homogeneous data from different sources.

[0150] Specifically, it can be a graph construction method based on graph aggregation, which integrates the single-cell batch intra-graph and the single-cell batch inter-graph, constructs the single-cell global similarity matrix and the single-cell global graph, and realizes the unified representation of single-cell expression data from different batches.

[0151] The single-cell global similarity matrix A is defined as:

[0152]

[0153] The above formula indicates that any element A[u,v] in the matrix A is composed of all the single-cell batch internal graphs (S K [p,q],ifu,v∈batchK) and all single-cell batch graphs (A (I,J) [p,q],ifu∈batchI,v∈batchJ,I≠J) are concatenated.

[0154] Where u,v are indices of the global matrix A, and:

[0155]

[0156] p and q are the local indices of u and v in the batch to which they belong.

[0157] Further, executing S203 and constructing a global graph of homogeneous data further includes:

[0158] In the unweighted graph between homogeneous data sources, a weight is assigned to each edge, and a weighted graph between homogeneous data sources and a corresponding similarity matrix between homogeneous data sources are constructed.

[0159] This implementation method may be a graph construction method based on graph aggregation, which integrates the intra-batch graph of single-cells with the inter-batch graph of single-cells to construct a single-cell global similarity matrix. Before integration, the inter-batch graph of single-cells is converted from an unweighted graph to a weighted graph based on prior experience or other weight calculation methods.

[0160] In the weighted single-cell batch-to-batch graph, the relationship between single-cell pairs between batches can be more finely represented, rather than simply connected or unconnected. It can reflect the true similarity between single-cell pairs, corresponding to the weighted graph of the single-cell batch-intra graph. It is used to filter out noise or unimportant connections and retain more meaningful inter-cell relationships. It helps to identify cell populations more accurately. The final single-cell global graph can more realistically express the relationship between similar single cells within a batch and similar single cells between batches.

[0161] Further, executing S204, smoothing the homogeneous data global graph includes:

[0162] Input features are input into the graph convolution filter to obtain output features, and then the global graph of homogeneous data is smoothed; the input features are homogeneous data expression matrices;

[0163] The graph convolution filter performs multi-layer convolution on the input features based on the simple graph convolution method, and minimizes the difference between the output features and the input features after convolution as the optimization goal; the number of layers of convolution on the input features is equal to the power of the normalized adjacency matrix of homogeneous data; the normalized adjacency matrix of homogeneous data is obtained based on the normalization of the global similarity matrix of homogeneous data.

[0164] In this implementation, it is mainly based on the optimization filter of simple graph convolution. Through the simple graph convolution (SGC) method, the optimization filter is used to perform multi-layer propagation and feature smoothing on the single-cell global similarity matrix (single-cell global graph), thereby generating a high-quality low-dimensional feature representation.

[0165] The optimization filter constructed in this embodiment can smooth the node feature X in the single-cell global graph under the constraint of graph G = (V, E) and achieve dimensionality reduction. The goal of simple graph convolution (SGC) is to optimize the feature distribution through the graph structure, making the features of similar nodes closer, thereby improving the quality of data representation and clustering effect.

[0166] The optimization problem of SGC can be modeled from the perspective of feature smoothing. Similar nodes in the graph G are represented by the single-cell global similarity matrix A. ij The value of represents the similarity between node i and node j. After graph convolution, feature Z can minimize the difference with the initial feature X while maintaining smoothness under the constraints of the graph structure. Based on this, the optimization goal of simple graph convolution can be expressed as:

[0167]

[0168] in is the normalized adjacency matrix of the graph, balancing the influence of node connection strength. This objective function shows that the invention hopes to generate features Z as close as possible to the smooth features after propagation through the graph structure.

[0169] In order to further enhance the utilization of global information, SGC extends the propagation of graph convolution to a high-order operation and uses K layers of graph convolution to capture the high-order neighborhood features of nodes. In this case, the optimization problem can be further expressed as:

[0170]

[0171] Where K represents the order of convolution, that is, the number of propagations. By introducing K layers of convolution, the present invention can aggregate graph structure information in a larger range, thereby making feature Z more global and smooth.

[0172] Further, executing S204, reducing the dimension of the homogeneous data global graph includes:

[0173] The output features are made equal to the input features, and the output features are obtained as the features after dimensionality reduction, thereby reducing the dimensionality of the global graph of homogeneous data.

[0174] In this implementation, the main purpose is to efficiently solve the optimization problem of the simple graph convolution mentioned above.

[0175] SGC can be solved directly using the closed-form solution To generate the reduced-dimensional features. Compared with the conventional graph convolutional network (GCN) method of nonlinear activation through parameter learning, the optimization process of SGC is simplified to matrix multiplication operations. This process removes complex learning parameters and only uses the structural information of the graph for high-order propagation of features, thereby greatly improving the computational efficiency. At the same time, the reduced-dimensional feature Z can be directly used for subsequent clustering or classification tasks, such as K-means or spectral clustering.

[0176] In some embodiments, other methods may be used instead of SGC to perform feature smoothing and dimensionality reduction on the single cell global graph, for example:

[0177] 1. Based on the Graph Attention Network (GAT), the graph attention mechanism is used to capture the importance of different cells within and between batches by assigning different weights, which can replace the fixed adjacency matrix propagation method of SGC.

[0178] 2. The spectral convolution-based network uses Chebyshev polynomials or other spectral domain methods to perform frequency domain convolution processing on the global graph structure, which can replace the time domain convolution operation in SGC.

[0179] 3. Deep Graph Adversarial Network (DGAN), combined with the framework of Generative Adversarial Network (GAN), transforms graph structure features into an adversarial learning process between the discriminator and the generator, further optimizing the effect of feature dimensionality reduction.

[0180] 4. Graph Autoencoder (GAE) and Variational Graph Autoencoder (VGAE), which input the graph structure within and between batches into the autoencoder network, can generate dimensionality reduction features through nonlinear encoding and decoding, replacing the linear propagation method of SGC.

[0181] The following is a summary of the algorithm flow:

[0182] Input: Single cell expression data X 1 ,X 2 ,…,X N (Contains N batches).

[0183] Output: Single-cell data after removing batch effects, and corresponding clustering results.

[0184] To summarize the steps above, the processing process of different batches of single-cell data based on the above method and system is as follows:

[0185] 1. Constructing single-cell batch internal graph based on sparse subspace clustering:

[0186] 1.1 For each batch, the similarity matrix S within the batch is calculated by sparse subspace clustering method I .

[0187] 2. Constructing single-cell batch graph based on mutual nearest neighbor (MNN):

[0188] 2.1 For batch pairs, the MNN method is used to identify single-cell anchor pairs across batches.

[0189] 2.2 Calculate and save the similarity matrix between single-cell batches.

[0190] 3. Graph aggregation: building a single-cell global graph structure:

[0191] 3.1 Concatenate the similarity matrices within all single-cell batches and between single-cell batches to generate a single-cell global similarity matrix.

[0192] 4. Graph filtering: based on simple graph convolution (SGC):

[0193] 4.1 Construct the normalized adjacency matrix.

[0194] 4.2 Perform multiple graph convolutions to generate dimensionality reduction features.

[0195] 5. Clustering:

[0196] 5.1 Based on the dimensionality reduction features, K-means or spectral clustering is used to complete cell clustering analysis.

[0197] In combination with the above specific application scenarios, it can be seen that the homogeneous and heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter provided in this application is:

[0198] 1. We propose a source-intra-graph construction method based on sparse subspace clustering (SSC), which can efficiently identify potential clustering structures within a source in high-dimensional single-cell expression data while ensuring the sparsity of the similarity matrix. Compared with existing clustering algorithms based on principal component analysis (PCA) or other dimensionality reduction methods, this method can more accurately capture the local geometric relationships of cells, thereby improving the accuracy of clustering.

[0199] 2. The proposed source-to-source graph construction method based on mutual nearest neighbor (MNN) can make full use of the anchor information between different sources, effectively avoiding the problem of inaccurate matching caused by data distribution differences in existing methods. Compared with the method of directly aligning all data between different sources globally, the method of the present invention is more flexible and computationally efficient, and it also performs better in some applications, such as capturing real biological relationships across batches.

[0200] 3. By integrating the intra-source and inter-source similarity matrices to construct a global graph structure, and using simple graph convolution (SGC) for multi-layer propagation and smoothing features, compared with existing nonlinear deep models (such as scVI and Harmony), the present invention can more fully retain the graph structure information in the data while reducing the computational complexity, thereby generating more reliable low-dimensional features.

[0201] 4. The proposed complete technical solution has high robustness and adaptability, and can handle complex batch effects and high-dimensional data. This advantage solves the performance degradation problem of existing methods (such as Seurat and DESC) when dealing with diverse data, and significantly improves the accuracy and stability of clustering.

[0202] In summary, this embodiment overcomes the limitations of the prior art in terms of batch effect removal, similarity construction, and feature dimensionality reduction through a series of methods, and provides an efficient and reliable solution for the analysis of single-cell data.

[0203] Based on the same inventive concept, the embodiment of the present application also provides a homogeneous heterogeneous data fusion system based on sparse subspace clustering and graph convolution filter. Figure 3 As shown, it is a schematic diagram of the structure of a system 300 for removing batch effects of homogeneous data, which may include:

[0204] A source internal graph construction module 301 is used to construct a homogeneous data source internal graph of each source by a sparse subspace clustering method based on the homogeneous data of each source;

[0205] The source-to-source graph construction module 302 is used to construct a homogeneous data source-to-source graph between every two sources based on the homogeneous data of every two sources;

[0206] A global graph construction module 303 is used to integrate all homogeneous data source intra-graphs and homogeneous data source inter-graphs to construct a homogeneous data global graph;

[0207] The smoothing and dimensionality reduction module 304 is used to perform smoothing and dimensionality reduction operations on the global graph of homogeneous data in sequence through a graph convolution filter, thereby removing the batch effect of homogeneous data from different sources, and further completing the fusion of homogeneous and heterogeneous data.

[0208] In the embodiment of the present application, the technical effects corresponding to the homogeneous heterogeneous data fusion system based on sparse subspace clustering and graph convolution filter can refer to the above method and will not be repeated here.

[0209] In some possible implementations, the homogeneous heterogeneous data fusion system based on sparse subspace clustering and graph convolution filter according to the present application may include at least a processor and a memory. The memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the homogeneous heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter according to various exemplary embodiments of the present application described in this specification. For example, the processor may execute the following steps: Figure 2 Follow the steps shown in .

[0210] Based on the same inventive concept, an electronic device is also provided in the embodiment of the present application, which can realize the functions of the aforementioned homogeneous heterogeneous data fusion method system based on sparse subspace clustering and graph convolution filter, referring to Figure 4 , the electronic device comprises:

[0211] At least one processor 401, and a memory 402 connected to the at least one processor 401. The specific connection medium between the processor 401 and the memory 402 is not limited in the embodiment of the present application. Figure 4 In the example, the processor 401 and the memory 402 are connected via the bus 400. The bus 400 is Figure 4 The connection between other components is shown by bold lines, and is not intended to be limiting. The bus 400 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 401 can also be called a controller, and there is no limitation on the name.

[0212] In the embodiment of the present application, the memory 402 stores instructions that can be executed by at least one processor 401. The at least one processor 401 can execute the homogeneous heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter discussed above by executing the instructions stored in the memory 402. The processor 401 can implement Figure 3 The functions of each module in the device shown.

[0213] Among them, the processor 401 is the control center of the device, and can use various interfaces and lines to connect the various parts of the entire control device. By running or executing instructions stored in the memory 402 and calling the data stored in the memory 402, the various functions of the device and process data, the device can be monitored as a whole.

[0214] In one possible design, the processor 401 may include one or more processing units, and the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly processes an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the modem processor may not be integrated into the processor 401. In some embodiments, the processor 401 and the memory 402 may be implemented on the same chip, and in some embodiments, they may also be implemented separately on separate chips.

[0215] Processor 401 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the homogeneous heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter disclosed in the embodiments of the present application can be directly embodied as a hardware processor execution, or a combination of hardware and software modules in the processor.

[0216] The memory 402 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 402 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 402 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 402 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0217] By programming the processor 401, the code corresponding to the homogeneous heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter introduced in the above embodiment can be fixed into the chip, so that the chip can execute it when running. Figure 2 The steps of the method for fusion of homogeneous and heterogeneous data based on sparse subspace clustering and graph convolution filter in the illustrated embodiment. How to design and program the processor 401 is a technology well known to those skilled in the art and will not be described in detail here.

[0218] Based on the same inventive concept, an embodiment of the present application also provides a storage medium, which stores computer instructions. When the computer instructions are executed on a computer, the computer executes the homogeneous and heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter discussed above.

[0219] In some possible embodiments, various aspects of the homogeneous heterogeneous data fusion method based on sparse subspace clustering and graph convolution filters provided in the present application can also be implemented in the form of a program product, which includes a program code. When the program product is run on an apparatus, the program code is used to enable the control device to execute the steps of the homogeneous heterogeneous data fusion method based on sparse subspace clustering and graph convolution filters according to various exemplary embodiments of the present application described above in this specification.

[0220] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0221] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0222] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0223] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0224] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A homogeneous and heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter, characterized in that: Here are the steps: Based on the homogeneous data of each source, a homogeneous data source intra-graph of each source is constructed by sparse subspace clustering method; Based on the homogeneous data from every two sources, a homogeneous data source graph between every two sources is constructed; Integrate all graphs within and between homogeneous data sources to build a global graph of homogeneous data; The global graph of homogeneous data is smoothed and dimensionally reduced through a graph convolution filter to remove batch effects of homogeneous data from different sources, thereby completing the fusion of homogeneous and heterogeneous data.

2. According to claim 1, a homogeneous heterogeneous data fusion method based on sparse subspace clustering and graph convolution filter is characterized in that: The construction of a homogeneous data source internal graph for each source includes: Cluster homogeneous data based on sparse subspace clustering method and construct similarity matrix within homogeneous data source; Based on the Frobenius norm of the difference between the homogeneous data expression matrix and the corresponding homogeneous data reconstruction matrix, a reconstruction error term is constructed; the homogeneous data reconstruction matrix is ​​a matrix obtained by linearly transforming the homogeneous data expression matrix through the similarity matrix within the homogeneous data source; the homogeneous data expression matrix is ​​a matrix representation of homogeneous data; The sparse regularization term is constructed through the sparsity regularization coefficient and the L1 norm of the similarity matrix within the homogeneous data source; Taking minimization of the sum of the reconstruction error term and the sparse regularization term as the optimization goal, and imposing a constraint that the diagonal elements of the similarity matrix within the homogeneous data source are zero during the optimization process, the first optimization formula is constructed; The first optimization formula is solved to obtain a homogeneous data source intra-graph and a corresponding homogeneous data source intra-similarity matrix.

3. The method for fusion of homogeneous and heterogeneous data based on sparse subspace clustering and graph convolution filter according to claim 2 is characterized in that: Construct a homogeneous data source graph between every two sources, including: Based on the mutual nearest neighbor method, the homogeneous data pairs between two sources that satisfy the mutual nearest neighbor relationship are taken as anchor data pairs to obtain the anchor data pair set; Based on the set of anchor point data pairs, an unweighted graph between homogeneous data sources and a corresponding similarity matrix between homogeneous data sources are constructed; in the graph between homogeneous data sources, there are edge connections between the anchor point data, and there are no edge connections between the non-anchor point data.

4. The method for fusion of homogeneous and heterogeneous data based on sparse subspace clustering and graph convolution filter according to claim 3 is characterized in that: Build a global graph of homogeneous data, including: The similarity matrices within all homogeneous data sources are concatenated with the similarity matrices between all homogeneous data sources to obtain the homogeneous data global similarity matrix and the corresponding homogeneous data global graph.

5. The method for fusion of homogeneous and heterogeneous data based on sparse subspace clustering and graph convolution filter according to claim 4 is characterized in that: Constructing a global graph of homogeneous data also includes: In the unweighted graph between homogeneous data sources, a weight is assigned to each edge, and a weighted graph between homogeneous data sources and a corresponding similarity matrix between homogeneous data sources are constructed.

6. The method for fusion of homogeneous and heterogeneous data based on sparse subspace clustering and graph convolution filter according to claim 4 or 5, characterized in that: The step of smoothing the global graph of homogeneous data includes: Inputting the input features into the graph convolution filter to obtain the output features, and then smoothing the global graph of homogeneous data; the input features are homogeneous data expression matrices; The graph convolution filter performs multi-layer convolution on the input features based on a simple graph convolution method, and takes minimizing the difference between the output features and the input features after convolution as the optimization goal; wherein the number of layers for convolution on the input features is equal to the power of the normalized adjacency matrix of homogeneous data; the normalized adjacency matrix of homogeneous data is obtained based on the normalization of the global similarity matrix of homogeneous data.

7. The method for fusion of homogeneous and heterogeneous data based on sparse subspace clustering and graph convolution filter according to claim 6, characterized in that: The dimensionality reduction of the global graph of homogeneous data includes: The output features are made equal to the input features, and the output features are obtained as the features after dimensionality reduction, thereby reducing the dimensionality of the global graph of homogeneous data.

8. A homogeneous and heterogeneous data fusion system based on sparse subspace clustering and graph convolution filter, characterized in that: include: A source internal graph construction module is used to construct a homogeneous data source internal graph of each source through a sparse subspace clustering method based on the homogeneous data of each source; An inter-source graph construction module is used to construct an inter-source graph of homogeneous data sources between every two sources based on homogeneous data of every two sources; The global graph construction module is used to integrate all the graphs within and between homogeneous data sources to construct a global graph of homogeneous data; The smoothing and dimensionality reduction module is used to smooth and reduce the dimensionality of the homogeneous data global graph through a graph convolution filter, remove the batch effect of homogeneous data from different sources, and complete the fusion of homogeneous and heterogeneous data.

9. An electronic device, characterized in that: The device comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes any one of the methods in claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The storage medium comprises a program code, and when the storage medium is run on an electronic device, the program code is used to enable the electronic device to execute any one of the methods described in claims 1 to 7.

Citation Information

Cited By

  • Enterprise multi-source heterogeneous data automatic fusion method and system

    CN121117929A