Multi-scale multi-view clustering method and system based on deep learning

Through the multi-scale multi-view clustering method of deep learning, using graph diffusion convolution and high-order contrastive learning, combined with a fine-grained fusion module, the information integration problem of multi-view data is solved and the clustering performance is improved.

CN120747569APending Publication Date: 2025-10-03BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510747251.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies fail to fully utilize the rich information of multi-view data in multi-view clustering and lack effective integration and utilization strategies, resulting in poor clustering performance.

Method used

A multi-scale multi-view clustering method based on deep learning is adopted. The graph diffusion convolution module is used to propagate neighbor information. A high-order contrastive learning module is introduced to capture complex relationships. The importance of samples is adaptively learned through a fine-grained multi-view fusion module, and a hypergraph matrix is ​​constructed for clustering.

Benefits of technology

It improves the deep representation and clustering performance of multi-view data, enhances the integration and utilization of multi-source information, and improves the clustering effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747569A_ABST
    Figure CN120747569A_ABST
Patent Text Reader

Abstract

The invention provides a multi-scale multi-view clustering method and system based on deep learning, and the method comprises the steps: obtaining initial features of a plurality of visual angles, inputting the initial features of each visual angle into a preprocessing module, and enabling the preprocessing module to output and obtain a common structural feature matrix and a common node feature matrix; inputting the common structural feature matrix and the common node feature matrix corresponding to each view angle into a double-graph diffusion convolution module, and calculating the structural feature matrix and the node feature matrix corresponding to each view angle by the double-graph diffusion convolution module based on the common structural feature matrix and the common node feature matrix; obtaining an iteration structure feature matrix and an iteration node feature matrix corresponding to each view angle; the iterative structure feature matrix and the iterative node feature matrix are enhanced through a graph convolutional network of a high-order contrast learning module, and the iterative node feature matrix of each view angle is constructed into a common graph through a fine-grained multi-view fusion module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multi-scale multi-view clustering method and system based on deep learning. Background Art

[0002] Multi-view clustering algorithms typically aim to learn representations for each view and integrate this information to obtain a unified representation for clustering. Earlier, researchers typically used traditional machine learning methods to address multi-view clustering. However, these methods failed to effectively integrate features learned from different views, hindering the final clustering optimization. Consequently, a number of subspace-based learning methods have emerged, enabling multiple views to jointly learn a low-dimensional mapping using complementary information. In addition, traditional methods include graph-based and multi-kernel learning methods. While these methods perform well in specific scenarios and offer strong interpretability, they are limited by their shallow networks and are unable to learn more complex, deep information, resulting in limited clustering results. In recent years, deep learning has emerged as a prominent focus of researchers in the field of multi-view clustering. While its theoretical foundations are similar to those of earlier traditional machine learning methods, its superior feature learning capabilities have become a key factor driving research. Deep learning methods not only extract richer and more abstract features from the latent space but also possess a more powerful representation capability when handling complex data relationships. This trend is particularly pronounced in multi-view clustering tasks. Specifically, deep subspace methods share similar assumptions with their predecessors in machine learning, namely subspace methods. However, deep subspace methods are unique in that they employ deep neural networks to extract data layer by layer, thereby obtaining a deeper representation of the data. The introduction of this deep representation makes the model more adaptable to complex data structures, thereby improving clustering performance.

[0003] However, current research still has shortcomings. Specifically, the inclusion of multi-view data in static multi-view clustering has not been fully considered. Multi-view data contains rich information and can provide observations from different dimensions, but current research has not yet made sufficient use of multi-view data and lacks effective strategies for integrating and utilizing multi-source information. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a multi-scale multi-view clustering method and system based on deep learning to eliminate or improve one or more defects in the prior art.

[0005] One aspect of the present invention provides a multi-scale multi-view clustering method based on deep learning, the method comprising the following steps: Obtaining initial features of multiple perspectives, inputting the initial features of each perspective into a preprocessing module, the preprocessing module outputting a structural feature matrix and a node feature matrix through a first processing module, and aggregating the structural feature matrices and node feature matrices corresponding to the multiple perspectives through a second processing module to obtain a common structural feature matrix and node feature matrix; The common structural feature matrix and node feature matrix corresponding to each perspective are input into the dual graph diffusion convolution module, which calculates the structural feature matrix and node feature matrix corresponding to each perspective based on the common structural feature matrix and node feature matrix to obtain the iterative structural feature matrix and iterative node feature matrix corresponding to each perspective; the iterative structural feature matrix and iterative node feature matrix are enhanced and fused through the graph convolution network of the high-order contrastive learning module to obtain a hypergraph matrix, and the hypergraph matrices of multiple perspectives are constructed into a common graph through the fine-grained multi-view fusion module.

[0006] Adopting the above scheme, this scheme mainly focuses on how to use deep learning methods to efficiently obtain the deep representation of the complex multi-view data in the current Internet and complete the final clustering. Graph diffusion convolution is an iterative method for propagating neighbor information. The diffusion process can make good use of node information, which will ensure the utilization of public information. In addition to diffusion enhancement technology, this scheme also introduces a high-order contrastive learning method. By using hypergraphs to capture the complex relationships between samples for contrastive learning, the model view learning effect is improved. When fusing multiple view features, in order to better utilize the connection information between views, multi-source information can be effectively integrated and utilized.

[0007] In some embodiments of the present invention, the method further comprises processing the public graph by a preset multi-view clustering module, wherein the multi-view clustering module comprises an encoding module and a decoding module, and outputting a final clustering result by the decoding module.

[0008] In some embodiments of the present invention, the first processing module includes a KNN processing layer and a linear layer. In the step in which the preprocessing module outputs the structural feature matrix and the node feature matrix through the first processing module, the node feature matrix is ​​output through the KNN processing layer, and the structural feature matrix is ​​output through the linear layer.

[0009] In some embodiments of the present invention, in the step of aggregating the structural feature matrices and node feature matrices corresponding to multiple perspectives through the second processing module to obtain a common structural feature matrix and node feature matrix, the structural feature matrices corresponding to multiple perspectives are aggregated by means of mean aggregation to obtain a common structural feature matrix; and the node feature matrices corresponding to multiple perspectives are aggregated to obtain a common node feature matrix.

[0010] In some embodiments of the present invention, in the step of calculating the structural feature matrix and the node feature matrix corresponding to each perspective based on the common structural feature matrix and the node feature matrix to obtain the iterative structural feature matrix and the iterative node feature matrix corresponding to each perspective, the structural matrix and the degree matrix corresponding to each of the structural feature matrices and each node feature matrix are calculated, and the iterative structural feature matrix and the iterative node feature matrix corresponding to each perspective are calculated based on the structural matrix and the degree matrix.

[0011] In some embodiments of the present invention, in the step of calculating the iterative structural feature matrix and the iterative node feature matrix corresponding to each perspective based on the structure matrix and the degree matrix, the following formula is used to calculate the iterative structural feature matrix and the iterative node feature matrix: in, represents the final output iterative structure feature matrix, P represents the maximum number of diffusion iterations, k represents the current number of iterations, θ k represents the PageRank coefficient, Indicates the transfer matrix of the structure matrix corresponding to the current number of iterations; b represents the weighting coefficient; Indicates the transfer matrix of the common structural feature matrix corresponding to the current number of iterations; X V represents the node feature matrix of the perspective V, H represents the common node feature matrix, Iteration node feature matrix representing the final output.

[0012] In some embodiments of the present invention, the weighting coefficient is calculated using the following formula: Among them, A v Represents the structural feature matrix of the perspective V, U represents the common structural feature matrix, and F is the step representation in the calculation process.

[0013] In some embodiments of the present invention, the fine-grained multi-view fusion module includes a linear layer, a matrix multiplication layer and an activation function layer. In the step of constructing the hypergraph matrices of multiple perspectives into a common graph through the fine-grained multi-view fusion module, the hypergraph matrices of every two perspectives are constructed into a processing group, the hypergraph matrices in each processing group are processed by a linear layer, and the two processed matrices are sequentially processed by a matrix multiplication layer and an activation function layer to obtain a combined graph corresponding to the processing group, and all the combined graphs are aggregated to obtain a common graph.

[0014] In some embodiments of the present invention, the steps of the method also include pre-training the graph convolutional network of the high-order contrastive learning module. In the pre-training step, a hypergraph adjacency matrix of the hypergraph matrix is ​​constructed, the two hypergraph adjacency matrices of the processing group are compared, and the loss function is calculated, and the graph convolutional network is pre-trained based on the loss function.

[0015] In some embodiments of the present invention, in the step of comparing the two hypergraph matrices based on the processing group and calculating the loss function, the loss function is calculated using the following formula: Among them, L represents the loss function, N represents the number of samples, and N pos Indicates the number of positive samples, N neg ∪N pos Represents the collection of positive and negative samples, Z j Represents the j-th column element sequence of any sample in the positive sample, Z t Represents the t-th column element sequence of any sample in the collection of positive samples and negative samples, It represents the matrix element value obtained from the adjacency matrix of the view v and u. Represents the matrix element value of the i-th row and j-th column in the adjacency matrix of view v, Represents the matrix element value of the i-th row and j-th column in the adjacency matrix of view u, S() represents the calculation of the inner product, The sequence of elements in the i-th row of the hypergraph matrix representing the view v, Represents a normalization operation.

[0016] The second aspect of the present invention also provides a multi-scale multi-view clustering system based on deep learning, which includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps implemented by the method described above.

[0017] The third aspect of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps implemented by the aforementioned multi-scale multi-view clustering method based on deep learning.

[0018] Additional advantages, objects, and features of the present invention will be described in part in the following description and will become apparent to those skilled in the art after studying the following or may be learned by practice of the present invention. The objects and other advantages of the present invention may be particularly pointed out and attained in the description and drawings.

[0019] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other purposes that can be achieved by the present invention will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention.

[0021] Figure 1 This is a schematic diagram of the first implementation of the multi-scale multi-view clustering method based on deep learning in this scheme; Figure 2 This is a schematic diagram of the second implementation of the multi-scale multi-view clustering method based on deep learning in this scheme; Figure 3 Schematic diagram of the processing architecture of the multi-scale multi-view clustering method based on deep learning in this solution. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0023] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, the accompanying drawings only show structures and / or processing steps closely related to the solutions according to the present invention, while other details that are not closely related to the present invention are omitted.

[0024] In the prior art Graph neural networks are models for processing graph-structured data. Each node updates its own representation by considering the information of its neighboring nodes, thereby incorporating the global information of the graph into the representation of each node. This iterative update process enables GNN to gradually propagate and integrate the information in the graph structure, thereby better capturing the structure and characteristics of the graph. Common graph neural networks include GCN and GAT, etc. The difference between these methods lies in the different ways of message interaction. Deep graph neural networks can better capture the hierarchical information and complex relationships in the graph structure. In some multi-view clustering tasks, the introduction of graph-based deep learning methods can help the network enhance its ability to capture associated information, while also facilitating the interaction of information between nodes and making it more conducive to processing multi-view data. Currently, graph-based deep multi-view clustering methods are a research hotspot, and many research works are centered around this direction. ; Prior art 1 The view-specific and cross-view propagation modules proposed in the prior art can capture the consistency and complementary information between multiple views. The designed fusion module not only fuses node attributes and multi-view relationship information, but also simultaneously learns fusion weights on popular datasets.

[0025] Disadvantages of the prior art 1 This method only uses a common graph structure to model the relationships between multi-view data samples, and the constructed graph is not optimized. This causes the model to lose deep sample association information, reduces its ability to learn sample features, and ultimately degrades the final clustering results.

[0026] Prior Art 2 The second existing technique obtains consistent representations from multiple views by aggregating features across samples and views, thereby fully exploiting the complementary information between similar samples. Furthermore, this method uses a structure-guided contrastive learning module to align the consistent representation and the view-specific representation, ensuring that the view-specific representations of different samples are highly correlated and similar.

[0027] Disadvantages of the second prior art The second existing approach, using contrastive learning, fails to carefully consider the interactions between different views and samples. Furthermore, it ignores sample-level correlations during the multi-view fusion phase. These shortcomings can degrade the quality of the resulting features, impacting clustering performance.

[0028] like Figure 1 and 3 As shown, the present invention proposes a multi-scale multi-view clustering method based on deep learning, the steps of the method include: Step S100: Acquire initial features from multiple viewpoints and input the initial features from each viewpoint into a preprocessing module. The preprocessing module outputs a structural feature matrix and a node feature matrix through a first processing module, and aggregates the structural feature matrices and node feature matrices corresponding to the multiple viewpoints through a second processing module to obtain a common structural feature matrix and node feature matrix. In the specific implementation process, graph structures are very effective in representing neighborhood relationships in multi-view data. Given multi-view data containing N samples and V views, this study establishes a graph representation for each view. The graph structure matrix for each view is constructed using the KNN method. In addition, to simplify subsequent processing and enhance the effectiveness of the node representation, a linear layer is used to project the features of each view to the same dimension. Finally, the graphs of all views are averaged.

[0029] Step S200: Inputting the common structural feature matrix and the node feature matrix corresponding to each view into a dual graph diffusion convolution module, wherein the dual graph diffusion convolution module calculates the structural feature matrix and the node feature matrix corresponding to each view based on the common structural feature matrix and the node feature matrix to obtain an iterative structural feature matrix and an iterative node feature matrix corresponding to each view; In some embodiments of the present invention, graph diffusion convolution aims to update node representations by propagating information across the graph structure, based on the perspective of graph signal processing. Compared to conventional graph node information exchange, each node can indirectly obtain information from more distant nodes, thus achieving a larger receptive field. Furthermore, the iterative diffusion mechanism makes the model more stable against noise and local perturbations.

[0030] Specifically, the common structural feature matrix and node feature matrix contain information from all viewpoints. Incorporating this information into each viewpoint can enhance the representation of each viewpoint graph. Graph diffusion convolution is an effective information propagation method. However, existing graph diffusion convolution methods only focus on integrating the structural (edge) information of the graph, ignoring the interaction between node features. In addition, this solution can fully address the diversity problem between viewpoints, using the same graph diffusion convolution for different viewpoints.

[0031] In step S300, the iterative structural feature matrix and the iterative node feature matrix are enhanced and fused through the graph convolutional network of the high-order contrastive learning module to obtain a hypergraph matrix, and the hypergraph matrices of multiple perspectives are constructed into a common graph through the fine-grained multi-view fusion module.

[0032] Adopting the above scheme, this scheme mainly focuses on how to use deep learning methods to efficiently obtain the deep representation of the complex multi-view data in the current Internet and complete the final clustering. Graph diffusion convolution is an iterative method for propagating neighbor information. The diffusion process can make good use of node information, which will ensure the utilization of public information. In addition to diffusion enhancement technology, this scheme also introduces a high-order contrastive learning method. By using hypergraphs to capture the complex relationships between samples for contrastive learning, the model view learning effect is improved. When fusing multiple view features, in order to better utilize the connection information between views, multi-source information can be effectively integrated and utilized.

[0033] like Figure 2 As shown, in some embodiments of the present invention, the steps of the method also include, step S400, the steps of the method also include processing the common graph through a pre-set multi-view clustering module, the multi-view clustering module includes an encoding module and a decoding module, and the final clustering result is output through the decoding module.

[0034] In the specific implementation process, multi-view data often involves complex associations. Therefore, when designing a clustering framework to implement multi-view clustering, it is necessary to consider how to accurately capture valuable information in multiple views in the complex associations. To this end, this study proposes a multi-view clustering framework based on multi-scale, such as Figure 3 As shown in the figure, the proposed framework consists of five key modules: First, the data preprocessing module receives the initial features of different views and outputs the initial common graph and the graphs of each view. These graphs are then fed into the dual graph diffusion convolution module, where complementary information is introduced to enhance each view. Then, the high-order contrastive learning module captures high-order information for better representation learning. Next, in the multi-view fine-grained fusion module, a weight matrix is ​​generated based on the importance of the sample level, and the views are fused accordingly. Finally, the fused features are clustered in the multi-view clustering module.

[0035] In some embodiments of the present invention, the first processing module includes a KNN processing layer and a linear layer. In the step in which the preprocessing module outputs the structural feature matrix and the node feature matrix through the first processing module, the node feature matrix is ​​output through the KNN processing layer, and the structural feature matrix is ​​output through the linear layer.

[0036] In some embodiments of the present invention, in the step of aggregating the structural feature matrices and node feature matrices corresponding to multiple perspectives through the second processing module to obtain a common structural feature matrix and node feature matrix, the structural feature matrices corresponding to multiple perspectives are aggregated by means of mean aggregation to obtain a common structural feature matrix; and the node feature matrices corresponding to multiple perspectives are aggregated to obtain a common node feature matrix.

[0037] In some embodiments of the present invention, in the step of calculating the structural feature matrix and the node feature matrix corresponding to each perspective based on the common structural feature matrix and the node feature matrix to obtain the iterative structural feature matrix and the iterative node feature matrix corresponding to each perspective, the structural matrix and the degree matrix corresponding to each of the structural feature matrices and each node feature matrix are calculated, and the iterative structural feature matrix and the iterative node feature matrix corresponding to each perspective are calculated based on the structural matrix and the degree matrix.

[0038] Using the above scheme, the dual graph diffusion convolution module of this scheme can adaptively integrate the complementary information of edges and nodes in the public graph into each perspective. The module evaluates the diversity of each perspective and adjusts the proportion of the influence of public features on each specific perspective graph through weighting coefficients during the diffusion process.

[0039] In some embodiments of the present invention, in the step of calculating the iterative structural feature matrix and the iterative node feature matrix corresponding to each perspective based on the structure matrix and the degree matrix, the following formula is used to calculate the iterative structural feature matrix and the iterative node feature matrix: in, represents the final output iterative structure feature matrix, P represents the maximum number of diffusion iterations, k represents the current number of iterations, θ k represents the PageRank coefficient, Indicates the transfer matrix of the structure matrix corresponding to the current number of iterations; b represents the weighting coefficient; Indicates the transfer matrix of the common structural feature matrix corresponding to the current number of iterations; X V represents the node feature matrix of the perspective V, H represents the common node feature matrix, Iteration node feature matrix representing the final output.

[0040] Specifically, Among them A v is the structure matrix of the v-th perspective, D v is its corresponding degree matrix, where U is the structure matrix of the public graph.

[0041] In some embodiments of the present invention, the weighting coefficient is calculated using the following formula: Among them, A v Represents the structural feature matrix of the perspective V, U represents the common structural feature matrix, and F is the step representation in the calculation process.

[0042] In some embodiments of the present invention, the fine-grained multi-view fusion module includes a linear layer, a matrix multiplication layer and an activation function layer. In the step of constructing the hypergraph matrices of multiple perspectives into a common graph through the fine-grained multi-view fusion module, the hypergraph matrices of every two perspectives are constructed into a processing group, the hypergraph matrices in each processing group are processed by a linear layer, and the two processed matrices are sequentially processed by a matrix multiplication layer and an activation function layer to obtain a combined graph corresponding to the processing group, and all the combined graphs are aggregated to obtain a common graph.

[0043] In practice, when fusing multiple views, different weights are often assigned to each overall view. This ignores the differences in the importance of individual samples and leads to coarse-grained fusion. The attention mechanism empowers the model to adaptively learn the importance of individual samples based on their inter-sample correlations. This capability enables the fine-grained multi-view fusion module to better capture key patterns and correlations in the data, enabling detailed consideration of the impact of each sample.

[0044] The specific formula is as follows: Among them, Z v represents the output of the final layer of the graph convolutional network at the vth view, W v ∈R m×m and W u ∈R m×m is a learnable parameter. θ represents the gate unit: θ(Y:W g )=Y⊙δ(YW g ), where Y is the input of the gate unit, W g ∈R N×N is a learnable parameter matrix. ⊙ denotes the Hadamard product. Then, our scheme fuses the views into a new common graph for subsequent iterations.

[0045] In some embodiments of the present invention, the steps of the method also include pre-training the graph convolutional network of the high-order contrastive learning module. In the pre-training step, a hypergraph adjacency matrix of the hypergraph matrix is ​​constructed, the two hypergraph adjacency matrices of the processing group are compared, and the loss function is calculated, and the graph convolutional network is pre-trained based on the loss function.

[0046] Using this approach, hypergraphs can connect multiple vertices through a single edge, capturing rich relationships and patterns and more accurately modeling high-order relationships between samples. This approach proposes a high-order contrastive learning module to help the model leverage high-order information for better data representation. This module uses a graph convolutional network (GCN) to enhance node features.

[0047] In some embodiments of the present invention, in the step of comparing the two hypergraph matrices based on the processing group and calculating the loss function, the loss function is calculated using the following formula: Among them, L represents the loss function, N represents the number of samples, and N pos Indicates the number of positive samples, N neg ∪N pos Represents the collection of positive and negative samples, Z jRepresents the j-th column element sequence of any sample in the positive sample, Z t Represents the t-th column element sequence of any sample in the collection of positive samples and negative samples, It represents the matrix element value obtained from the adjacency matrix of the view v and u. Represents the matrix element value of the i-th row and j-th column in the adjacency matrix of view v, Represents the matrix element value of the i-th row and j-th column in the adjacency matrix of view u, S() represents the calculation of the inner product, The sequence of elements in the i-th row of the hypergraph matrix representing the view v, Represents a normalization operation.

[0048] In the specific implementation process, the hypergraph matrix of the perspective v is determined to be a positive sample or a negative sample based on the hypergraph adjacency matrix of the perspective v; Specifically, the distance between the hypergraph adjacency matrix and the standard matrix is ​​calculated, and based on the comparison between the distance and the distance threshold, it is determined to be a positive sample or a negative sample.

[0049] The hypergraph adjacency matrix of view v is constructed as follows: Among them, ρ represents the threshold, B v Represents a normal graph The derived hypergraph adjacency matrix reflects high-order information. If Indicates that the i-th sample and the j-th sample belong to the same hyperedge. In addition, retain As a supplement, it allows subsequent contrastive learning to capture more comprehensive information. Here, this scheme converts the matrix Considered as B v The sub-edge weight matrix of .

[0050] In the specific implementation process, a decoder is used to reconstruct node features and graph structure: Where ψ represents a multi-layer perceptron (MLP), is the reconstructed node feature, is the reconstructed graph structure. δ is the sigmoid activation function.

[0051] A self-trained clustering layer is used to obtain multi-view clustering results based on learned node features. Clustering is achieved using an encoder and decoder network. The encoder is used to obtain the hidden layer representation. Where N is the number of samples, d u is the dimension of the hidden layer. Q∈R N×K represents the clustering result, K is the total number of cluster centers. Then, M is used to derive the clustering result Q: Among them, M i represents the sample features of the i-th row of M, is the jth cluster center that is randomly initialized. β is the degree of freedom of the Student’s t distribution. Q ij Represents the similarity between the sample and the cluster center. The clustering result of sample i can be obtained by: To facilitate unsupervised learning, an auxiliary distribution P is defined to help optimize the cluster centers. The choice of P has been discussed many times. The choice of auxiliary distribution P is defined as follows: Among them, f j =∑ i Q ij Represents the frequency of the jth cluster. This normalization method can prevent the auxiliary distribution bias caused by the existence of larger clusters.

[0052] In summary, to better address the multi-view clustering problem, this paper extracts and integrates information at three different scales: per-sample, per-view, and between-views. This multi-scale strategy enables the present invention to leverage the diverse graph structures and complementary information in multi-view data, achieving superior clustering performance. Specifically, graph diffusion convolution is an iterative method for propagating neighborhood information between nodes. Inspired by this, this proposal proposes a dual-graph diffusion convolution module to enhance each view by incorporating complementary information from the consistency graph. Considering the varying importance of complementary information for each view, this proposal designs a weighting module to adaptively learn the amount of complementary information captured by different views. After enhancing a single view, this proposal constructs a hypergraph and designs a contrastive loss function to capture the richer relational structure between hypergraphs, thereby better modeling complex relationships and providing the model with high-order relational information. Furthermore, to reflect the unique characteristics of different samples, this proposal proposes a fine-grained multi-view fusion module that learns correlations between samples through a self-attention mechanism. This enables the model to exploit the relationships between all samples, thereby improving fusion performance.

[0053] To address the current weakness of multi-view clustering in capturing real-world relationships, this proposal proposes a multi-scale deep multi-view clustering approach. This approach focuses on using deep learning methods to efficiently extract deep representations and perform clustering for the complex multi-view data currently found on the internet. Graph diffusion convolution is an iterative method for propagating neighbor information, and several studies have demonstrated its ability to significantly improve graph quality. The diffusion process effectively leverages node information, ensuring the utilization of public information. In addition to diffusion enhancement techniques, a high-order contrastive learning approach is introduced, which uses a hypergraph to capture complex relationships between samples for contrastive learning, improving the model's view learning performance. To better utilize inter-view connections when fusing features from multiple views, a fine-grained adaptive fusion module is designed based on the self-attention mechanism in the Transformer, enabling the model to more carefully consider the importance of each view. Specifically, a self-attention mechanism is first used in each mid-view to collect correlations with other views as a base score. A gating mechanism is then used to control this base score, allowing the model to autonomously select more important information. By addressing these two issues, the original problem is effectively addressed and clustering performance is improved.

[0054] The beneficial effects of this program include: 1. We introduce a dual graph diffusion convolution module. This module effectively diffuses complementary information between different views, providing global information across different views. Our proposed method surpasses traditional graph diffusion methods by adaptively propagating and diffusing graph structure and node features, resulting in a more comprehensive update of the entire graph. We also incorporate weight coefficients to control the influence of the common graph on different views, enabling the model to better propagate information between views.

[0055] 2. We propose a high-order contrastive learning module. Hypergraphs can connect multiple vertices through a single edge, capturing rich relationships and patterns and more accurately modeling high-order relationships between samples. This module leverages hypergraphs to capture more complex high-order neighborhood relationships between different data samples within a view, enabling better representation learning. Furthermore, incorporating the influence of sub-edge weights within the hypergraph into contrastive learning helps the model more comprehensively understand the correlations between data samples within a view.

[0056] 3. A fine-grained multi-view fusion module is introduced to facilitate adaptive fusion weight learning at the sample scale. When fusing multiple views, samples in each view are typically assigned a uniform weight, which ignores the differences in the importance of individual samples and leads to coarse-grained fusion. This method, however, draws on the computational methods of self-attention and adaptively learns the importance of each sample based on their inter-sample correlations, significantly improving the fusion quality of multi-view data.

[0057] An embodiment of the present invention also provides a multi-scale multi-view clustering system based on deep learning, which includes a computer device, wherein the computer device includes a processor and a memory, wherein the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps implemented by the method described above.

[0058] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps implemented by the aforementioned multi-scale multi-view clustering method based on deep learning. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the art.

[0059] It should be understood by those skilled in the art that the various exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.

[0060] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.

[0061] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.

[0062] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A multi-scale multi-view clustering method based on deep learning, characterized by: The steps of the method include: Obtaining initial features of multiple perspectives, inputting the initial features of each perspective into a preprocessing module, the preprocessing module outputting a structural feature matrix and a node feature matrix through a first processing module, and aggregating the structural feature matrices and node feature matrices corresponding to the multiple perspectives through a second processing module to obtain a common structural feature matrix and node feature matrix; The common structural feature matrix and node feature matrix corresponding to each perspective are input into the dual graph diffusion convolution module, which calculates the structural feature matrix and node feature matrix corresponding to each perspective based on the common structural feature matrix and node feature matrix to obtain the iterative structural feature matrix and iterative node feature matrix corresponding to each perspective; the iterative structural feature matrix and iterative node feature matrix are enhanced and fused through the graph convolution network of the high-order contrastive learning module to obtain a hypergraph matrix, and the hypergraph matrices of multiple perspectives are constructed into a common graph through the fine-grained multi-view fusion module.

2. The multi-scale multi-view clustering method based on deep learning according to claim 1, characterized in that The method further includes processing the public graph through a preset multi-view clustering module, wherein the multi-view clustering module includes an encoding module and a decoding module, and outputting a final clustering result through the decoding module.

3. The multi-scale multi-view clustering method based on deep learning according to claim 1, characterized in that The first processing module includes a KNN processing layer and a linear layer. In the step in which the preprocessing module outputs the structural feature matrix and the node feature matrix through the first processing module, the node feature matrix is ​​output through the KNN processing layer, and the structural feature matrix is ​​output through the linear layer.

4. The multi-scale multi-view clustering method based on deep learning according to any one of claims 1 to 3, characterized in that: In the step where the dual graph diffusion convolution module calculates the structural feature matrix and the node feature matrix corresponding to each perspective based on the common structural feature matrix and the node feature matrix to obtain the iterative structural feature matrix and the iterative node feature matrix corresponding to each perspective, the structural matrix and the degree matrix corresponding to each of the structural feature matrices and each node feature matrix are calculated, and the iterative structural feature matrix and the iterative node feature matrix corresponding to each perspective are calculated based on the structural matrix and the degree matrix.

5. The multi-scale multi-view clustering method based on deep learning according to claim 4, characterized in that In the step of calculating the iterative structural feature matrix and the iterative node feature matrix corresponding to each perspective based on the structure matrix and the degree matrix, the following formula is used to calculate the iterative structural feature matrix and the iterative node feature matrix: in, represents the final output iterative structure feature matrix, P represents the maximum number of diffusion iterations, k represents the current number of iterations, θ k represents the PageRank coefficient, Indicates the transfer matrix of the structure matrix corresponding to the current number of iterations; b represents the weighting coefficient; Indicates the transfer matrix of the common structural feature matrix corresponding to the current number of iterations; X V represents the node feature matrix of the perspective V, H represents the common node feature matrix, Iteration node feature matrix representing the final output.

6. The multi-scale multi-view clustering method based on deep learning according to claim 5, characterized in that The weighting coefficient is calculated using the following formula: Among them, A v Represents the structural feature matrix of the perspective V, U represents the common structural feature matrix, and F is the step representation in the calculation process.

7. The multi-scale multi-view clustering method based on deep learning according to claim 1, characterized in that: The fine-grained multi-view fusion module includes a linear layer, a matrix multiplication layer and an activation function layer. In the step of constructing the hypergraph matrices of multiple perspectives into a common graph through the fine-grained multi-view fusion module, the hypergraph matrices of every two perspectives are constructed into a processing group, the hypergraph matrices in each processing group are processed by a linear layer, and the two processed matrices are sequentially processed by a matrix multiplication layer and an activation function layer to obtain a combined graph corresponding to the processing group, and all the combined graphs are aggregated to obtain a common graph.

8. The multi-scale multi-view clustering method based on deep learning according to claim 1, characterized in that The method also includes pre-training the graph convolutional network of the high-order contrastive learning module. In the pre-training step, a hypergraph adjacency matrix of the hypergraph matrix is ​​constructed, two hypergraph adjacency matrices of the processing group are compared, and a loss function is calculated, and the graph convolutional network is pre-trained based on the loss function.

9. The multi-scale multi-view clustering method based on deep learning according to claim 8, characterized in that In the step of comparing the two hypergraph matrices of the treatment group and calculating the loss function, the loss function is calculated using the following formula: Among them, L represents the loss function, N represents the number of samples, and N pos Indicates the number of positive samples, N neg ∪N pos Represents the collection of positive and negative samples, Z j Represents the j-th column element sequence of any sample in the positive sample, Z t Represents the t-th column element sequence of any sample in the collection of positive samples and negative samples, It represents the matrix element value obtained from the adjacency matrix of the view v and u. Represents the matrix element value of the i-th row and j-th column in the adjacency matrix of view v, Represents the matrix element value of the i-th row and j-th column in the adjacency matrix of view u, S() represents the calculation of the inner product, The sequence of elements in the i-th row of the hypergraph matrix representing the view v, Represents a normalization operation.

10. A multi-scale multi-view clustering system based on deep learning, characterized in that: The system includes a computer device, which includes a processor and a memory. The memory stores computer instructions. The processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps implemented by the method according to any one of claims 1 to 9.