Multi-view financial data clustering integration method and device
The ensemble pool is generated through K-Means clustering and fuzzy membership function, combined with the LSR model of Frobenius norm to optimize the co-correlation matrix, and learn the subspace projection matrix and affinity matrix, which solves the problem of information integration in multi-view clustering and achieves high-quality financial data clustering effect.
Patent Information
- Application Number
- CN202510770181.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-10
AI Technical Summary
When processing multi-source heterogeneous financial data, it is difficult for the existing technology to effectively integrate complementary information of label space and feature space, resulting in the limited application of multi-view clustering in the fields of financial risk assessment, portfolio optimization and market segmentation.
The ensemble pool is generated by the K-Means clustering method, and the binary division matrix is processed by the fuzzy membership function, combined with the LSR model of the Frobenius norm to optimize the initial co-correlation matrix, learn the subspace projection matrix and affinity matrix, and use adaptive weights to optimize the feature co-correlation matrix, and finally obtain the final clustering ensemble result of financial multi-view data through the graph cutting algorithm.
It significantly improves the representation ability of the co-correlation matrix, optimizes the accuracy and stability of the model, improves the clustering effect of financial multi-view data, and enhances the application value in complex financial scenarios.
Smart Images

Figure CN120296456A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a multi-view financial data clustering and integration method and apparatus. Background Art
[0002] The ensemble clustering technology can divide financial data into multiple distinct groups, reflecting different market behaviors and investment characteristics. These characteristics are related to diverse financial performances and may affect investment decisions and risk management effects. Therefore, identifying the similarities and differences in financial data is crucial for financial analysis and investment strategy formulation.
[0003] Multi-view clustering analysis integrates information from different financial data sources, and each source may provide unique insights, which helps to comprehensively understand market dynamics and financial products. However, this method faces multiple challenges: it is necessary to balance the information of each financial data source, evaluate its value, and cope with the noise and complexity brought by high-dimensional financial data.
[0004] Current financial data clustering technologies often have limited ability to represent the co-correlation matrix when dealing with multi-source heterogeneous data, and it is difficult to effectively integrate the complementary information in the label space and feature space of the data, which limits the application prospects of multi-view clustering in fields such as financial risk assessment, portfolio optimization, and market segmentation. In the paper "Multi-view Ensemble Clustering based on BLS-Autoencoder", BLS-Autoencoder is applied for feature extraction, but it fails to effectively integrate the complementary information in the label space and feature space; moreover, the constructed co-correlation matrix is more suitable for processing multimedia data with an obvious hierarchical structure and is limited in dealing with financial data. Summary of the Invention
[0005] The purpose of the present invention is to solve the problems in the prior art.
[0006] The technical solution adopted by the present invention to solve its technical problems is: to provide a multi-view financial data clustering and integration method, including:
[0007] Using the K-Means clustering method to process each subspace of the financial multi-view data to generate an ensemble pool, and then using a fuzzy membership function to process the ensemble pool to generate a binary partition matrix; learning the global structure of the label space of the financial multi-view data through the binary partition matrix to obtain an initial co-correlation matrix; using an LSR model with a Frobenius norm to optimize the initial co-correlation matrix to obtain a label co-correlation matrix; the LSR model represents the least squares method model.
[0008] For each view of the financial multi-view data, the subspace projection matrix, affinity matrix and sample adjacency matrix are learned, and adaptive weights are added according to the learning error to optimize the co-correlation matrix from the feature space to obtain the feature co-correlation matrix;
[0009] The label co-correlation matrix and the feature co-correlation matrix are combined to obtain an optimized co-correlation matrix; the optimized co-correlation matrix is projected into the constraint space, and an undirected graph is constructed based on the projected constraint space; the undirected graph is partitioned using a graph cutting algorithm to obtain the final clustering integration result of the financial multi-view data.
[0010] Preferably, the LSR model with Frobenius norm is used to optimize the initial co-correlation matrix to obtain a label co-correlation matrix, which is expressed as:
[0011] ;
[0012] in, The training data set is a binary partition matrix generated by K-Means clustering and fuzzy membership function. represents the initial co-correlation matrix, represents the Frobenius norm, represents the trade-off parameter of the co-correlation matrix, represents the clustering ensemble result, Representation based on the co-correlation matrix The constructed Laplacian matrix, represents the transpose of a matrix, Represents the rank of the matrix; v represents the view, and V represents the total number of views.
[0013] Preferably, the co-correlation matrix is optimized from the feature space to obtain a feature co-correlation matrix, which is expressed as:
[0014] ;
[0015] in, represents the subspace projection matrix, Representing multiple views of financial data, represents the row-2 norm and column-1 norm, represents adaptive weight; represents the parameter controlling the weight distribution, Represents the affinity matrix.
[0016] Preferably, the label co-correlation matrix and the feature co-correlation matrix are combined to obtain an optimized co-correlation matrix, which is expressed as:
[0017] .
[0018] Preferably, the optimized co - correlation matrix is projected into the constraint space, expressed as:
[0019] ;
[0020] where, represents the element projected into the constraint space; represents the element in the optimized co - correlation matrix , i represents the row index, and j represents the column index.
[0021] Preferably, the undirected graph is constructed based on the projection constraint space, expressed as:
[0022] ;
[0023] ;
[0024] where, represents the undirected bipartite graph, represents the edge weight of the nodes of the undirected bipartite graph, represents the set of sample nodes, represents the optimized co - correlation matrix, represents the fuzzy partition matrix in which, k represents the row index, and l represents the column index; the fuzzy partition matrix is the result obtained by spectral clustering and normalization of the optimized co - correlation matrix ; represents the node composed of samples and clusters, represents the connection weight based on the fuzzy partition matrix , E represents the set of all valid edge connections of the undirected bipartite graph, represents that there is an edge connection between sample node k and cluster node l.
[0025] Preferably, the undirected graph is partitioned by using the graph - cut algorithm to obtain the final clustering integration result of the multi - view financial data, specifically: the normalized cut algorithm is used to cut the undirected bipartite graph into several non - overlapping sub - graphs, each sub - graph contains a subset of the sample node set and the nodes corresponding to the co - correlation matrix, maintaining the properties of the undirected bipartite graph. The sample nodes belonging to the same sub - graph are partitioned into the same cluster, thus forming the final consensus partition , where each sample node is uniquely assigned to a cluster, and the clustering integration result of the multi - view financial data is obtained.
[0026] The present invention also provides a multi - view financial data clustering integration device, including:
[0027] The label space construction module processes each subspace of the financial multi-view data using the K-Means clustering method to generate an ensemble pool, and then uses the fuzzy membership function to process the ensemble pool to generate a binary partition matrix; learns the global structure of the label space of the financial multi-view data through the binary partition matrix to obtain an initial co-correlation matrix; optimizes the initial co-correlation matrix using the LSR model with the Frobenius norm to obtain a label co-correlation matrix;
[0028] The feature space construction module learns the subspace projection matrix, affinity matrix, and sample adjacency matrix for each view of the financial multi-view data, and increases the adaptive weight according to the learning error to optimize the co-correlation matrix from the feature space to obtain a feature co-correlation matrix;
[0029] The data clustering integration module combines the label co-correlation matrix and the feature co-correlation matrix to obtain an optimized co-correlation matrix; projects the optimized co-correlation matrix into the constraint space, constructs an undirected graph based on the projected constraint space; uses the graph cut algorithm to partition the undirected graph to obtain the final clustering integration result of the financial multi-view data.
[0030] The present invention has the following beneficial effects:
[0031] (1) The multi-view clustering integration method based on co-correlation matrix optimization constructed by the present invention obtains a high-quality co-correlation matrix by combining the LSR model and feature space optimization, and uses the trained optimization model to perform feature extraction and relationship mining on the input financial multi-view data, significantly improving the representation ability of the co-correlation matrix and efficiently completing the multi-view clustering analysis of financial data;
[0032] (2) The present invention optimizes the features of the original data through the subspace projection matrix and the adaptive weight mechanism, and at the same time learns the sample adjacency matrix to capture the internal relationship between samples; runs the normalized cut algorithm on the constructed undirected graph to obtain a high-quality consensus partition, and then completes the multi-view clustering integration of financial data, significantly optimizing the accuracy and stability of the model and enhancing its application value in complex financial scenarios;
[0033] (3) This method breaks through the limitation of traditional multi-view clustering algorithms that only learn from a single space, and fully integrates the complementary information of the label space and the feature space. In the co-correlation matrix construction stage, this method innovatively captures the global structure information through the LSR model with the Frobenius norm, and in the matrix optimization stage, by simultaneously learning the subspace projection matrix, affinity matrix, and adaptive weight, effectively reduces the interference of noise features. In the integration stage, a high-quality clustering partition is realized by constructing an undirected graph structure with an accurately optimized co-correlation matrix, thereby enhancing the clustering effect of financial multi-view data.
[0034] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments, but the present invention is not limited to the embodiments. Description of the Drawings
[0035] Figure 1 It is a method step diagram of a multi-view financial data clustering integration method according to an embodiment of the present invention;
[0036] Figure 2 It is a schematic flowchart of a multi-view financial data clustering integration method according to an embodiment of the present invention;
[0037] Figure 3 It is a comparison diagram of the effects of ablation experiments for a multi-view financial data clustering integration method according to an embodiment of the present invention;
[0038] Figure 4 It is a schematic structural diagram of a multi-view financial data clustering integration device according to an embodiment of the present invention. Detailed Embodiments
[0039] Referring to Figure 1 and Figure 2 as shown, it is a method step diagram and a schematic flowchart of a multi-view financial data clustering integration method according to an embodiment of the present invention, including the following steps:
[0040] S101. Use the K-Means clustering method to process each subspace of the financial multi-view data to generate an ensemble pool, and then use the fuzzy membership function to process the ensemble pool to generate a binary partition matrix; learn the global structure of the label space of the financial multi-view data through the binary partition matrix to obtain an initial co-correlation matrix; use the LSR model with the Frobenius norm to optimize the initial co-correlation matrix to obtain a label co-correlation matrix;
[0041] S102. For each view of the financial multi-view data, learn the subspace projection matrix, affinity matrix, and sample adjacency matrix, and increase the adaptive weight according to the learning error to optimize the co-correlation matrix from the feature space to obtain a feature co-correlation matrix;
[0042] S103. Combine the label co-correlation matrix and the feature co-correlation matrix to obtain an optimized co-correlation matrix; project the optimized co-correlation matrix into the constraint space, and construct an undirected graph based on the projected constraint space; use the normalized cut algorithm to partition the undirected graph to obtain the final clustering integration result of the financial multi-view data.
[0043] Specifically, in S101, K-Means clustering is used to extract integrated clustering results from different views to form a diverse ensemble pool; the fuzzy membership function is used to quantify the probabilistic association degree between samples and clusters, generating a binary partition matrix with rich information. The LSR model is optimized using a training dataset, and by minimizing the reconstruction error and regularization term, an accurate estimation of the co-correlation matrix is achieved, effectively reducing the interference of noise and the influence of outliers, enabling the co-correlation matrix to accurately reflect the inherent global structural characteristics of the data in the label space. Among them, the training dataset is a dataset constructed from multi-view financial data. The co-correlation matrix obtained by learning the global structure of the label space of the data through the binary partition matrix is as follows:
[0044] ;
[0045] Among them, The training dataset is a binary partition matrix generated by K-Means clustering and the fuzzy membership function, represents the co-correlation matrix obtained by training with the global structure of the data, represents the Frobenius norm, represents the trade-off parameter of the co-correlation matrix, represents the clustering ensemble result, represents based on the co-correlation matrix The constructed Laplacian matrix, represents the transpose of the matrix, represents the rank of the matrix.
[0046] Specifically, in S102, the adjacency relationship of samples is further mined from the feature layer of the original data. Since there is noise in the original feature space of each view, the subspace projection matrix, affinity matrix, and sample adjacency matrix are learned simultaneously, and an adaptive weight is added according to the learning error to better optimize the co-correlation matrix from the feature space. The co-correlation matrix obtained by learning through the feature space is as follows:
[0047]
[0048] Among them, represents the subspace projection matrix, represents multi-view financial data, represents the row two-norm column one-norm, Adaptive weight; represents the parameter for controlling weight allocation, represents the affinity matrix; represents the sample adjacency matrix.
[0049] Specifically, in S103, it includes the following steps:
[0050] S1031. Integrate the label space and feature space of the data to obtain an overall optimization objective function, expressed as:
[0051] .
[0052] The co - correlation matrix projected into the constraint space is as follows:
[0053] ;
[0054] represents the elements in the optimized co - correlation matrix, i represents the row index, and j represents the column index. in the
[0055] S1032. Construct an undirected bipartite graph based on the co - correlation matrix. The nodes in the undirected bipartite graph use all samples and clusters in the integration pool. The connection weights between the nodes of the undirected bipartite graph correspond to the values in the fuzzy partition matrix. The formula is as follows:
[0056] ;
[0057] ;
[0058] where, represents the undirected bipartite graph, represents the edge weight of the nodes of the undirected bipartite graph, represents the set of sample nodes, represents the optimized co - correlation matrix, represents the fuzzy partition matrix in the elements, k represents the row index, and l represents the column index; the fuzzy partition matrix is the result obtained by spectral clustering and normalization of the optimized co - correlation matrix ; represents the nodes composed of samples and clusters, represents the connection weight based on the fuzzy partition matrix , E represents the set of all valid edge connections of the undirected bipartite graph, represents that there is an edge connection between sample node k and cluster node l.
[0059] S1033. Use the graph - cut algorithm to cut the undirected bipartite graph into several non - overlapping sub - graphs. Each sub - graph contains a subset of the sample node set and the nodes corresponding to the co - correlation matrix, maintaining the properties of the undirected bipartite graph. The sample nodes belonging to the same sub - graph are partitioned into the same cluster, thus forming the final consensus partition , where each sample node is uniquely assigned to a cluster, and the clustering integration result of the multi - view financial data is obtained.
[0060] The embodiments of the present invention are verified as follows: The experimental datasets used mainly consist of a dataset of 100 plant leaves (100 leaves), a facial image dataset (ORL), 474 images (Caltech 7), and 7 different category images (MSRC-v1). For the evaluation of clustering performance, two widely adopted metrics are used, namely, the Normalized Mutual Information (NMI) and the Adjusted Rand Index (ARI). The standard mutual information is normalized, where 0 indicates that the two partitions are completely unrelated and 1 indicates that the two partitions are identical. It is a clustering evaluation metric obtained by adjusting the original Rand Index based on the consistency of sample pair assignments. The results are shown in Table 1 and Table 2.
[0061] Table 1 - Comparative analysis of NMI measurements:
[0062]
[0063] Table 2 - Comparative analysis of ARI measurements:
[0064]
[0065] According to the NMI metric, the method MvCOMOC of the present invention ranks first on three datasets and second on the MSRC-v1 dataset. In terms of ARI, it ranks first on two datasets.
[0066] At the same time, ablation experiments are used to prove the effect of the model constructed by the present invention. As Figure 3 shown, the effectiveness of the model of the present invention is proved by comparing three models. Among the three models, MvCO is a feature space model without optimization of the feature space with binary partitioning; MvCO-SA is a label space model without optimization of the label space at the feature layer; MvCOMOC is a complete model including the label space and the feature space. It can be seen from the figure that the NMI and ARI of MvCOMOC are the highest on different datasets. This phenomenon can be attributed to two main reasons: (1) The multi-view clustering integration method based on the optimization of the co-correlation matrix can synergistically integrate the global structure of the label space of the LSR model and the subspace mapping relationship of the feature space to construct a high-quality co-correlation matrix, thereby effectively reducing the interference of noise features. (2) The co-correlation matrix optimization method based on subspace affinity simultaneously learns the subspace projection matrix and the sample adjacency matrix, introduces an adaptive weight mechanism, and accurately captures the intrinsic relationship between samples, making the integration result more robust.
[0067] In summary, the method of the present invention innovatively combines co - correlation matrix optimization and subspace affinity and applies them to financial multi - view data analysis. The method first uses K - Means clustering to construct a diverse ensemble pool and generates a binary partition matrix through a fuzzy membership function. Then, an LSR model with Frobenius norm is designed to learn and refine the co - correlation matrix to capture the global structural information in the label space. Next, subspace projection matrix and sample adjacency matrix with L2,1 norm constraint are introduced for optimization at the feature space level, and an adaptive weight mechanism is used to dynamically adjust the contributions of each view. Finally, the optimized co - correlation matrix is projected into the constraint space and an undirected graph is constructed, and a high - quality clustering ensemble result is obtained through the NCUT algorithm. A large number of experiments verify the excellent performance of the method in financial multi - view data analysis, prove the effectiveness of the combination of the LSR model and subspace affinity optimization, and show the good application prospect of co - correlation matrix optimization technology in multi - view clustering tasks. The method of the present invention, through the combination of co - correlation matrix optimization and subspace affinity, is particularly suitable for the multi - source heterogeneity and time - varying correlation characteristics of financial data, can more accurately identify market sector rotation, asset risk aggregation and investment style changes, and provide more reliable technical support.
[0068] See Figure 4 As shown, it is a schematic structural diagram of a multi - view financial data clustering integration device according to an embodiment of the present invention, including:
[0069] A label space construction module 401 processes each subspace of financial multi - view data by using the K - Means clustering method to generate an ensemble pool, and then processes the ensemble pool by using a fuzzy membership function to generate a binary partition matrix; learns the global structure of the label space of financial multi - view data through the binary partition matrix to obtain an initial co - correlation matrix; optimizes the initial co - correlation matrix by using an LSR model with Frobenius norm to obtain a label co - correlation matrix;
[0070] A feature space construction module 402 learns a subspace projection matrix, an affinity matrix and a sample adjacency matrix for each view of financial multi - view data, increases the adaptive weight according to the learning error, and optimizes the co - correlation matrix from the feature space to obtain a feature co - correlation matrix;
[0071] A data clustering integration module 403 combines the label co - correlation matrix and the feature co - correlation matrix to obtain an optimized co - correlation matrix; projects the optimized co - correlation matrix into the constraint space, constructs an undirected graph based on the projected constraint space; uses the normalized cut algorithm to partition the undirected graph to obtain the final clustering integration result of financial multi - view data.
[0072] The functional implementation of the modules in the multi-view financial data clustering and integration device is the same as that of a multi-view financial data clustering and integration method, and will not be repeated here.
[0073] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-view financial data clustering integration method, characterized in that Including: Processing each subspace of financial multi-view data using the K-Means clustering method to generate an ensemble pool, and then processing the ensemble pool using a fuzzy membership function to generate a binary partition matrix; Learning the global structure of the label space of financial multi-view data through the binary partition matrix to obtain an initial co-correlation matrix; optimizing the initial co-correlation matrix using an LSR model with a Frobenius norm to obtain a label co-correlation matrix; the LSR model represents the least squares method model; For each view of the financial multi-view data, learning a subspace projection matrix, an affinity matrix, and a sample adjacency matrix, and increasing the adaptive weight according to the learning error to optimize the co-correlation matrix from the feature space to obtain a feature co-correlation matrix; Combining the label co-correlation matrix and the feature co-correlation matrix to obtain an optimized co-correlation matrix; Projecting the optimized co-correlation matrix into a constraint space, constructing an undirected graph based on the projected constraint space; using a graph cut algorithm to partition the undirected graph to obtain the final clustering ensemble result of the financial multi-view data.
2. The multi-view financial data clustering integration method according to claim 1, wherein The optimizing the initial co-correlation matrix using an LSR model with a Frobenius norm to obtain a label co-correlation matrix is expressed as: ; Among them, The training data set generates a binary partition matrix through K-Means clustering and a fuzzy membership function, represents the initial co-correlation matrix, represents the Frobenius norm, represents the trade-off parameter of the co-correlation matrix, represents the clustering ensemble result, represents based on the co-correlation matrix constructed Laplacian matrix, represents the transpose of the matrix, represents the rank of the matrix; v represents the view, and V represents the total number of views.
3. The multi-view financial data clustering integration method according to claim 2, wherein The optimizing the co-correlation matrix from the feature space to obtain a feature co-correlation matrix is expressed as: ; Among them, represents the subspace projection matrix, represents the multi-view financial data, represents the row two-norm and column one-norm, represents the adaptive weight; represents the parameter for controlling weight allocation, represents the affinity matrix.
4. The multi-view financial data clustering integration method according to claim 3, wherein The combining the label co-correlation matrix and the feature co-correlation matrix to obtain an optimized co-correlation matrix is expressed as: 。 5. The multi-view financial data clustering integration method according to claim 1, wherein The projecting the optimized co-correlation matrix into a constraint space is expressed as: ; Among them, represents the element projected into the constraint space; represents the element in the optimized co-correlation matrix where i represents the row index and j represents the column index.
6. The multi-view financial data clustering integration method according to claim 1, wherein The constructing an undirected graph based on the projected constraint space is expressed as: ; ; Among them, represents an undirected bipartite graph, represents the edge weights of the nodes of the undirected bipartite graph, represents the set of sample nodes, represents the optimized co-correlation matrix, represents the fuzzy partition matrix the element in, k represents the row index, and l represents the column index; the fuzzy partition matrix is the result obtained by spectral clustering and normalization of the optimized co-correlation matrix ; represents the nodes composed of samples and clusters, represents based on the fuzzy partition matrix the connection weights, E represents the set of all valid edge connections of the undirected bipartite graph, represents that there is an edge connection between the sample node k and the cluster node l.
7. The multi-view financial data clustering integration method according to claim 6, wherein Partition the undirected graph using the graph cut algorithm to obtain the final clustering ensemble result of the financial multi-view data, specifically: use the normalized cut algorithm to cut the undirected bipartite graph into several non-overlapping subgraphs, each subgraph contains a subset of the sample node set and the nodes corresponding to the co-correlation matrix, maintaining the properties of the undirected bipartite graph, and the sample nodes belonging to the same subgraph are partitioned into the same cluster, thus forming the final consensus partition. , where each sample node is uniquely assigned to a cluster to obtain the clustering ensemble result of the multi-view financial data.
8. A multi-view financial data clustering integration device, characterized in that, Including: A label space construction module that processes each subspace of financial multi-view data using the K-Means clustering method to generate an ensemble pool, and then processes the ensemble pool using a fuzzy membership function to generate a binary partition matrix; learning the global structure of the label space of financial multi-view data through the binary partition matrix to obtain an initial co-correlation matrix; optimizing the initial co-correlation matrix using an LSR model with a Frobenius norm to obtain a label co-correlation matrix; the LSR model represents the least squares method model; A feature space construction module that, for each view of the financial multi-view data, learns a subspace projection matrix, an affinity matrix, and a sample adjacency matrix, and increases the adaptive weight according to the learning error to optimize the co-correlation matrix from the feature space to obtain a feature co-correlation matrix; A data clustering ensemble module that combines the label co-correlation matrix and the feature co-correlation matrix to obtain an optimized co-correlation matrix; Projecting the optimized co-correlation matrix into a constraint space, constructing an undirected graph based on the projected constraint space; using a graph cut algorithm to partition the undirected graph to obtain the final clustering ensemble result of the financial multi-view data.
Citation Information
Patent Citations
Subspace clustering method based on high-dimensional overlapping data analysis
CN107832791A
Multi-view atlas clustering method and system based on graph optimization
CN114782727A
Graph clustering method and device based on multi-order neighbor information transfer fusion clustering network
CN114792113A
Multi-modal adaptive fusion deep clustering model and method based on auto-encoder
US20240095501A1
High-order correlation preserved incomplete multi-view subspace clustering method and system
US20240248960A1