Incremental multi-view data clustering method and system based on cross-time consensus graph

By using an incremental multi-view data clustering method based on a cross-time consensus graph, historical knowledge and current view information are integrated to construct a dynamic consensus graph and perform alternating optimization. This solves the problems of high computational cost and low accuracy of multi-view clustering in incremental environments, and achieves efficient and accurate clustering results.

CN121456530APending Publication Date: 2026-02-03ANHUI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511338450.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing multi-view clustering methods suffer from high computational costs and reduced clustering accuracy in incremental or dynamic environments. In particular, non-incremental methods require full-scale computation, while existing incremental methods suffer from information loss and suboptimal performance due to the decoupling optimization of view fusion and discrete clustering.

Method used

An incremental multi-view data clustering method based on cross-time consensus graphs is adopted. By integrating historical knowledge and current view information through kernel-induced representation, a consensus affinity matrix of time is learned to construct a dynamic consensus graph. Spectral embedding and discrete label learning are then performed, and the matrix in the consensus graph is alternately optimized by updating variables in stages.

Benefits of technology

It improves clustering accuracy, reduces computational costs, enhances temporal stability and robustness, effectively adapts to incremental environments, reduces information loss, and improves computational efficiency and clustering performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456530A_ABST
    Figure CN121456530A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an incremental multi-view data clustering method and system based on a cross-time consensus graph, and belongs to the technical field of artificial intelligence. The method comprises the following steps: integrating historical knowledge with a consensus affinity matrix of current view information and learning time based on kernel induction expression, and constructing a dynamic consensus graph; spectrum embedding and discrete label learning are carried out; and alternately optimizing the consensus affinity matrix, the orthogonal rotation matrix, the consensus spectrum embedding matrix and the discrete clustering label matrix of the moments in the dynamic consensus graph by adopting staged updating variables. The method can efficiently adapt to incremental environment application, and is better in clustering precision, higher in calculation efficiency, higher in time sequence stability and better in robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to an incremental multi-view data clustering method and system based on cross-time consensus graphs. Background Technology

[0002] The rapid growth and diversification of data in terms of dimensionality and volume have presented new challenges and opportunities for data mining and processing. Complex, high-dimensional data is characterized by both structure and multi-source nature. Traditional single-view methods struggle to handle heterogeneous, multi-source data because they cannot integrate complementary information between different modalities. In contrast, multi-view clustering can simultaneously capture shared information from multiple views as well as specific view features, providing more accurate and comprehensive data descriptions and more reliable clustering results. It demonstrates strong practicality and promising prospects in fields such as data mining, multimedia analysis, and recommender systems.

[0003] However, facing increasingly complex and diverse multi-view data, effectively mining deep structural information and developing superior clustering methods remains a significant challenge. Among various methods, graph-based multi-view spectral clustering methods have attracted considerable attention due to their powerful structural modeling capabilities. These methods utilize topological graphs to reveal the local manifold structure and global correlations of data points, thereby improving clustering performance through joint optimization of graph learning and spectral embedding. Some researchers in this field focus on learning better affinity graphs by dynamically weighting different views using multiple features to better capture the intrinsic structure of the data. Others emphasize fully utilizing the structural information of multiple views, simultaneously learning non-negative embedding matrices and spectral embedding matrices to obtain clustering results in one step while fully utilizing the complementary information of different views. Some have incorporated tensor Schatten p-norm constraints to fully utilize consistent and unique information among multiple views. Still others have introduced block diagonal mechanisms and fused the feature similarity matrices of different views to preserve the local clustering structure of each view. Furthermore, some have introduced emerging deep learning methods into clustering, including deep autoencoders, to obtain more comprehensive feature representations while integrating the results of learning multiple views and improving the accuracy of the results. Still others have implemented the networking of traditional optimization processes through the design of partially parameterized and differentiable modules, enhancing the scalability of the model.

[0004] In practical applications, multi-view data often arrives sequentially in a streaming manner. However, while existing graph-based multi-view spectral clustering methods can mine the local manifold structure and global correlations of data through topological graphs, most are designed for static environments and face significant challenges when applied to incremental or dynamic environments. Many methods require complete storage of historical data and repeated joint processing, which leads to a rapid increase in memory and computational costs. Most models adopt a uniform processing strategy without distinguishing between historical and new data, resulting in feature decay, redundancy, and reduced clustering accuracy. Recent research has proposed incremental multi-view spectral clustering methods that can dynamically update the basic kernel and spectral embeddings; others have introduced sparse graph learning to support incremental updates. Still others have explored persistent multi-view clustering methods through late fusion and consensus partitioning. However, these methods still rely on a separate multi-step information extraction and clustering process, which may lead to information loss and unsatisfactory performance.

[0005] Therefore, traditional clustering methods have two major limitations: non-incremental methods require computationally expensive full-scale calculations, while existing incremental methods suffer from error accumulation due to the decoupling optimization of view fusion and discrete clustering. Although there is research on incremental multi-view clustering, such as dynamically updating the base kernel and spectral embedding, and incrementally updating to complete sparse graph learning, these methods still rely on the decoupled process of "feature learning + clustering," which easily leads to information loss and suboptimal performance. Summary of the Invention

[0006] The purpose of this invention is to provide an incremental multi-view data clustering method based on cross-time consensus graphs. This method can efficiently adapt to incremental environment applications, has better clustering accuracy, higher computational efficiency, stronger temporal stability, and better robustness.

[0007] To achieve the above objectives, a first aspect of the present invention provides an incremental multi-view data clustering method based on a cross-time consensus graph, the method comprising: Based on kernel-induced representation, historical knowledge and current view information are integrated, and the learning time... A consensus affinity matrix is ​​used to construct a dynamic consensus graph; Perform spectral embedding and discrete label learning; Using phased variable updates to the dynamic consensus graph The consensus affinity matrix, orthogonal rotation matrix, consensus spectrum embedding matrix, and discrete clustering label matrix at each time step are alternately optimized.

[0008] Preferably, based on kernel-induced representation, historical knowledge and current view information are integrated, and the learning time... t The consensus affinity matrix, and the construction of the dynamic consensus graph, include: The objective function is obtained according to formula (1).

[0009] (1) in, Represents the spectral embedding of the newly arrived view. The obtained kernel matrix, spectral embedding By the Similarity matrix of views Obtained through spectral embedding function Generated by combining the nearest neighbor criterion with a Gaussian kernel function. This indicates that a kernel function is used to map a low-dimensional representation to a high-dimensional feature space, and the matrix... This represents the accumulated historical knowledge stored in the previous consensus kernel. for Consensus affinity matrix at any moment It is by The obtained Laplace matrix, It is an orthogonal rotation matrix. It is a discrete clustering label matrix. , and The parameters that affect the weights of different parts, For the sample size, This represents the number of categories.

[0010] Preferably, spectral embedding and discrete label learning include: According to formula (2), from Derive the normalized Laplace matrix. (2) in, Through Laplace regularization terms Constrained Spectral Embedding Local manifold structure; Using orthogonal rotation matrix Implement spectral embedding To discrete clustering labels The direct mapping to eliminate rotational ambiguity satisfies Each sample belongs to only one class, resulting in direct clustering results.

[0011] Preferably, the phased updating of variables includes: Keep other variables fixed and update according to formulas (3) to (6). ,

[0012] (3) make , ,and Formula (4) is obtained.

[0013] (4) in, ; set up Formula (4) is decomposed into finding Formula (5) for each independent form of a subproblem. (5) If there is express The result after sorting the elements in ascending order, the corresponding Rearranged as ,and, The most One non-zero element; According to formula (6), the optimal solution is obtained by introducing the Lagrange multiplier and applying the KKT conditions. (6).

[0014] Preferably, the phased updating of variables includes: Keep other variables fixed and update according to formulas (7) to (9). , (7) definition Formula (8) is obtained. (8) According to formula (9), using Singular value decomposition We obtain a closed-form solution. , (9).

[0015] Preferably, the phased updating of variables includes: With other variables fixed, update according to formulas (10) and (11). , (10) in, ,make Thus, we obtain formula (11). (11) Formula (11) is solved using a generalized power iteration algorithm, by applying the constructed matrix. Perform singular value decomposition to update step by step Until it converges.

[0016] Preferably, the phased updating of variables further includes: With other variables fixed, update according to formulas (12) and (13). , (12) in, ; Separate rows by selecting column indexes that maximize the contribution of the objective function. Solve this problem, let Thus, we obtain formula (13). (13).

[0017] A second aspect of the present invention provides an incremental multi-view data clustering system based on a cross-time consensus graph, the system comprising: The dynamic consensus graph construction module is used to integrate historical knowledge and current view information based on kernel-induced representation, with a learning time. A consensus affinity matrix is ​​used to construct a dynamic consensus graph; The spectral embedding and discrete label learning module is used for spectral embedding and discrete label learning. The phased variable update module is used to update the variables obtained from the dynamic consensus graph construction module. The consensus affinity matrix, orthogonal rotation matrix, consensus spectrum embedding matrix, and discrete clustering label matrix at each time step are alternately optimized.

[0018] A third aspect of the present invention provides a machine-readable storage medium storing instructions for causing a machine to perform the incremental multi-view data clustering method based on a cross-temporal consensus graph as described above.

[0019] A fourth aspect of the present invention provides a processor for running a program, wherein the program is run to perform the incremental multi-view data clustering method based on cross-time consensus graphs as described above.

[0020] Based on the above technical solution, an incremental multi-view data clustering framework is used to integrate dynamic graph construction, spectral alignment, and discrete label generation into an end-to-end optimization process, simulating an incremental view environment, starting from the first view data and sequentially reaching the next view data. Entering the [time]th [moment] Each view improves clustering accuracy and computational efficiency while adapting to incremental application environments by jointly optimizing dynamic consensus graph construction, spectral embedding learning, and discrete clustering assignment.

[0021] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0022] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of the incremental multi-view data clustering method based on cross-time consensus graph provided by the present invention; Figure 2 The convergence curves of the incremental multi-view data clustering method based on cross-time consensus graph provided by this invention are applied to ORL, Caltech101-7, MSRCV1 and BBC datasets. Figure 3 The incremental multi-view data clustering method based on cross-time consensus graph provided by this invention is applied to BBC, ORL, MSRCV1 and Caltech101-7 multi-view incremental performance curves. Detailed Implementation

[0023] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.

[0024] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0025] See Figure 1 The first aspect of this invention provides an incremental multi-view data clustering method based on a cross-time consensus graph, the method comprising: Based on kernel-induced representation, historical knowledge and current view information are integrated, and the learning time... A consensus affinity matrix is ​​used to construct a dynamic consensus graph; Perform spectral embedding and discrete label learning; Using phased variable updates to the dynamic consensus graph The consensus affinity matrix, orthogonal rotation matrix, consensus spectrum embedding matrix, and discrete clustering label matrix at each time step are alternately optimized.

[0026] First, in this embodiment, historical knowledge and current view information are integrated based on kernel-induced representation, and the learning time... t The consensus affinity matrix, and the construction of the dynamic consensus graph, include: The objective function is obtained according to formula (1).

[0027] (1) in, Represents the spectral embedding of the newly arrived view. The obtained kernel matrix, spectral embedding By the Similarity matrix of views Obtained through spectral embedding function Generated by combining the nearest neighbor criterion with a Gaussian kernel function. This indicates that a kernel function is used to map a low-dimensional representation to a high-dimensional feature space, and the matrix... This represents the accumulated historical knowledge stored in the previous consensus kernel, achieved by minimizing the consensus affinity matrix. Despite the differences between the two cores, this model ensures timing consistency and has good structural adaptability; It is by The obtained Laplace matrix, It is an orthogonal rotation matrix. It is a discrete clustering label matrix. , and The parameters that affect the weights of different parts, For the sample size, This represents the number of categories.

[0028] Secondly, spectral embedding and discrete label learning include: According to formula (2), from Derive the normalized Laplace matrix. (2) in, Through Laplace regularization terms Constrained Spectral Embedding Local manifold structure; Using orthogonal rotation matrix Implement spectral embedding To discrete clustering labels The direct mapping to eliminate rotational ambiguity satisfies Each sample belongs to only one class, resulting in direct clustering results.

[0029] Next, the first step in the phased variable update in the alternating optimization strategy includes: Keep other variables fixed and update according to formulas (3) to (6). ,

[0030] (3) make , ,and Formula (4) is obtained.

[0031] (4) in, ; set up Formula (4) is decomposed into finding Formula (5) for each independent form of a subproblem. (5) If there is express The result after sorting the elements in ascending order, the corresponding Rearranged as ,and, The most One non-zero element; According to formula (6), the optimal solution is obtained by introducing the Lagrange multiplier and applying the KKT conditions. (6).

[0032] Furthermore, the second step of updating variables in stages includes: Keep other variables fixed and update according to formulas (7) to (9). , (7) definition Formula (8) is obtained. (8) According to formula (9), using Singular Value Decomposition (SVD) of We obtain a closed-form solution. , (9).

[0033] Furthermore, the third step of updating variables in stages includes: With other variables fixed, update according to formulas (10) and (11). , (10) in, ,make Thus, we obtain formula (11). (11) The above problem is an orthogonally constrained subproblem, which is solved using the generalized power iteration algorithm (GPI) in formula (11). This is achieved by applying the constructed matrix... Perform singular value decomposition (SVD) to update incrementally. Until it converges.

[0034] Furthermore, the fourth step of the phased variable update includes: With other variables fixed, update according to formulas (12) and (13). , (12) in, ; Separate rows by selecting column indexes that maximize the contribution of the objective function. To solve efficiently, Thus, we obtain formula (13). (13).

[0035] at the same time, Figure 2 The figure shows the convergence curve in the application of the incremental multi-view data clustering method based on cross-time consensus graph provided by this invention. In the figure, the horizontal axis represents the number of iterations, and the vertical axis represents the objective function value. It shows that the objective function value monotonically decreases with the number of iterations and tends to stabilize. Figure 3 The figure shows the multi-view incremental performance curve in the application of the incremental multi-view data clustering method based on cross-time consensus graph provided by the present invention. In the figure, the horizontal axis represents the time step. t (i.e., the first) t Entering the [time]th [moment] t (7 views), with the vertical axis representing the values ​​of 7 clustering indicators. Higher values ​​indicate better performance, showing that clustering accuracy gradually increases with the number of views, indicating that the invention has good adaptability and will not cause catastrophic forgetting of historical data information.

[0036] Therefore, it can be seen that the incremental multi-view data clustering method based on cross-time consensus graph provided by this invention, compared with other multi-step clustering methods, can effectively reduce the impact of noise and other information by eliminating information loss in the decoupling process through end-to-end optimization, and obtain clustering results efficiently and accurately. The ACC, NMI and other indicators on datasets such as COIL20 and ORL are superior to the existing technology.

[0037] The incremental multi-view data clustering method based on cross-time consensus graph provided by this invention uses only one consensus affinity matrix to store historical data information when the views are dynamically added. The incremental update mechanism avoids reprocessing of historical data, reduces storage requirements and computational pressure, and significantly improves running efficiency. In the BBC dataset 4-view scenario, the time taken is only 0.1646 seconds, which is significantly lower than non-incremental methods.

[0038] The incremental multi-view data clustering method based on cross-time consensus graph provided by this invention maintains geometric consistency and local feature structure through cross-time consensus graph, effectively revealing the potential features and deep structure of data. There is no catastrophic forgetting during the multi-view incremental process, and it is highly malleable, achieving a good balance between learning new data and retaining historical data.

[0039] The incremental multi-view data clustering method based on cross-time consensus graph provided by this invention is insensitive to the arrival order of views, exhibits small performance fluctuations under different sequences, and verifies the good adaptability of this invention through various different incremental view update methods.

[0040] The following provides a specific implementation method to illustrate the above-mentioned incremental multi-view data clustering method based on cross-time consensus graphs: First, the experimental environment was selected, with hardware consisting of an AMD Ryzen 7 8845HS @ 3.80GHz CPU and 32GB RAM, and software consisting of MATLAB R2018a. Before clustering, all feature vectors were normalized to ensure consistent scaling across views.

[0041] We used seven public datasets, including: COIL20: Contains 20 different objects, each with images taken from a different perspective, totaling 1440 images, 20 categories, and 3 views; Scene15: Primarily sourced from Google Image Search and personal photos, it includes 15 different scenes (office, kitchen, etc.), totaling 4485 images, 15 categories, and 3 views; ORL is a face dataset containing images divided into 40 different themes, with 10 images in each theme. The images show certain differences in time, expression, etc. There are a total of 400 samples, 40 classes, and 4 views. Caltech101-7 is a subset of the Caltech101 dataset, containing 1474 images across 7 classes and 6 views. Handwritten: A dataset of images containing handwritten digits 0 to 9. Each view is a feature of the original dataset. There are a total of 2000 samples, 10 classes, and 6 views. MSRCV1: A total of 210 images, comprising 7 categories and 6 views; BBC: A dataset of news articles from the BBC News website, containing 685 samples across 5 classes and 4 views.

[0042] At the same time, nine comparison methods were used, including: SMSC: Effectively separates general and specific information in multi-view data and uses it for structured subspace clustering; UOMvSC: Integrates view fusion, similarity matrix construction, and spectral decomposition into a unified framework; AGLLFSR: Jointly optimizes view-specific similarity graphs and cross-view unified representations; MVSC-TLRR: Symmetric low-rank constraints and structured sparse low-rank constraints are applied to the frontal and horizontal slices of the tensor to characterize the relationships within and between views, respectively. SMVSC: Jointly learns a sparse structured consistent similarity matrix by combining the similarity matrices of multiple views; EOMSC-CA: Combines anchor learning and graph construction into a unified framework and imposes graph connectivity constraints, which can directly output cluster labels with the best anchor graph; SCGL: An incremental multi-view spectral clustering method that constructs a similarity graph with sparse and connected properties; CMVC: Introducing incremental learning methods into late-stage fusion multi-view clustering; IMSC: Based on stochastic Fourier features and low-rank approximation methods, it incrementally updates a consensus kernel to complete the incremental multi-view spectrum clustering task.

[0043] Configure the number of nearest neighbors in the parameter configuration. This involves selecting different optimal values ​​within the range {10, 20, 30, 40, 50, 60, 70} for each dataset using a grid search; and tuning the regularization parameters for each dataset using a grid search, with a range of:

[0044] And kernel functions, including comparative experiments of linear kernels and RBF kernels (fixed bandwidth of 1).

[0045] Next, we will proceed with the specific implementation steps: First, initialize the system by inputting data; the number of categories is fixed. Fixed parameters , , ; Calculate the affinity matrix when the first view data arrives. ,Depend on Spectral embeddings are obtained Then solve for the initial labels. ; In the When view data arrives (incremental iteration process): Step 1: When a new view arrives, calculate its similarity matrix. ,Depend on Spectral embeddings are obtained Based on historical consensus affinity matrix Spectral embeddings are obtained ,Depend on Obtain the kernel matrix ,Depend on Obtain the kernel matrix ; Step 2: Update the current consensus affinity matrix ; Step 3: Update sequentially , , Until convergence; Finally, the final discrete label matrix is ​​output. .

[0046] To evaluate the quality of the clustering results, this specific implementation uses seven metrics: accuracy (ACC), normalized mutual information (NMI), adjusted RAND index (ARI), purity, F-score, precision, and recall. It compares the results with nine other methods on seven datasets. This invention uses a linear kernel function and an RBF kernel function with a fixed bandwidth of 1. The comparison results are shown in Table 1 below. Higher values ​​for the metrics indicate better performance (the highest value is the optimal result, and the second highest value is the suboptimal result). Table 1

[0047] As shown in the table above, this invention performs optimally. The method provided by this invention achieves best or near-optimal results in ACC, NMI, ARI, Purity, and F-score on most datasets. Regarding Precision and Recall, this invention performs well on approximately half of the datasets. Notably, this invention performs better on datasets with more views, such as MSRCV1 (6 views), Handwritten (6 views), Caltech101-7 (6 views), and BBC (4 views). Although the risk of information loss increases in incremental view settings, the direct generation mechanism of discrete labels in these methods helps mitigate errors that accumulate over time, demonstrating their strong robustness in handling incrementally arriving multi-view data.

[0048] In addition, a key advantage of the method provided by this invention lies in its incremental learning framework, which results in lower computational costs and memory usage. This invention (using a linear kernel) was compared with nine other methods on the Handwritten and BBC datasets, with the number of views progressively increasing. The specific runtime comparison is shown in Table 2 below, which illustrates the average runtime (in seconds) for each view: Table 2

[0049] As shown in the table above, the method provided in this invention performs very efficiently on the BBC dataset, especially with a small sample size, and exhibits strong robustness to high-dimensional views. On the Handwritten dataset, by embedding a discrete label solving method into the model, this invention reduces computational overhead, outperforming most other methods. Furthermore, compared to non-incremental methods, this invention achieves significant time savings, highlighting the efficiency of incremental learning.

[0050] In summary, the present invention offers a favorable trade-off between accuracy and efficiency, especially in complex, dynamic multi-view environments.

[0051] Furthermore, in this incremental multi-view data clustering method based on cross-time consensus graphs, the kernel function replacement could consider using a multinomial kernel instead of a linear or RBF kernel to adapt to different data distributions and enhance the model's robustness. Regarding optimization strategy adjustment, in large-scale data scenarios, stochastic gradient descent can be used instead of alternating optimization to accelerate convergence and reduce time complexity. Adaptive weight learning can also be combined to incorporate weights balancing historical and new view information into the optimization, fine-tuning based on real-time clustering quality to improve consensus graph quality and replace fixed parameters. and The nearest neighbor strategy.

[0052] In addition, a second aspect of the present invention provides an incremental multi-view data clustering system based on a cross-time consensus graph, the system comprising: The dynamic consensus graph construction module is used to integrate historical knowledge and current view information based on kernel-induced representation, with a learning time. A consensus affinity matrix is ​​used to construct a dynamic consensus graph; The spectral embedding and discrete label learning module is used for spectral embedding and discrete label learning. The phased variable update module is used to update the variables obtained from the dynamic consensus graph construction module. The consensus affinity matrix, orthogonal rotation matrix, consensus spectrum embedding matrix, and discrete clustering label matrix at each time step are alternately optimized.

[0053] Meanwhile, a third aspect of the present invention provides a machine-readable storage medium storing instructions for causing a machine to execute the incremental multi-view data clustering method based on a cross-time consensus graph as described above.

[0054] Furthermore, a fourth aspect of the present invention provides a processor for running a program, wherein the program is run to execute the incremental multi-view data clustering method based on cross-time consensus graphs as described above.

[0055] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0056] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0057] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0058] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0059] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0060] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0061] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0062] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0063] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. An incremental multi-view data clustering method based on cross-time consensus graphs, characterized in that, The method includes: Based on kernel-induced representation, historical knowledge and current view information are integrated, and the learning time... A consensus affinity matrix is ​​used to construct a dynamic consensus graph; Perform spectral embedding and discrete label learning; Using phased variable updates to the dynamic consensus graph The consensus affinity matrix, orthogonal rotation matrix, consensus spectrum embedding matrix, and discrete clustering label matrix at each time step are alternately optimized.

2. The incremental multi-view data clustering method based on cross-time consensus graphs according to claim 1, characterized in that, The learning time is based on the integration of historical knowledge and current view information using kernel-induced representation. t The consensus affinity matrix, and the construction of the dynamic consensus graph, include: The objective function is obtained according to formula (1). (1) in, Represents the spectral embedding of the newly arrived view. The obtained kernel matrix, spectral embedding By the Similarity matrix of views Obtained through spectral embedding function Generated by combining the nearest neighbor criterion with a Gaussian kernel function. This indicates that a kernel function is used to map a low-dimensional representation to a high-dimensional feature space, and the matrix... This represents the accumulated historical knowledge stored in the previous consensus kernel. for Consensus affinity matrix at any moment It is by The obtained Laplace matrix, It is an orthogonal rotation matrix. It is a discrete clustering label matrix. , and The parameters that affect the weights of different parts, For the sample size, This represents the number of categories.

3. The incremental multi-view data clustering method based on cross-time consensus graphs according to claim 1, characterized in that, The spectral embedding and discrete label learning process includes: According to formula (2), from Derive the normalized Laplace matrix. ,(2) in, Through Laplace regularization terms Constrained Spectral Embedding Local manifold structure; Using orthogonal rotation matrix Implement spectral embedding To discrete clustering labels The direct mapping to eliminate rotational ambiguity satisfies Each sample belongs to only one class, resulting in direct clustering results.

4. The incremental multi-view data clustering method based on cross-time consensus graphs according to claim 1, characterized in that, The phased variable update includes: Keep other variables fixed and update according to formulas (3) to (6). , ,(3) make , ,and Formula (4) is obtained. ,(4) in, ; set up Formula (4) is decomposed into finding Formula (5) for each independent form of a subproblem. ,(5) If there is express The result after sorting the elements in ascending order, the corresponding Rearranged as ,and, The most in One non-zero element; According to formula (6), the optimal solution is obtained by introducing the Lagrange multiplier and applying the KKT conditions. (6)。 5. The incremental multi-view data clustering method based on cross-time consensus graphs according to claim 4, characterized in that, The phased variable update includes: Keep other variables fixed and update according to formulas (7) to (9). , ,(7) definition Formula (8) is obtained. ,(8) According to formula (9), using Singular value decomposition We obtain a closed-form solution. ,(9)。 6. The incremental multi-view data clustering method based on cross-time consensus graphs according to claim 5, characterized in that, The phased variable update includes: With other variables fixed, update according to formulas (10) and (11). , ,(10) in, ,make Thus, we obtain formula (11). ,(11) Formula (11) is solved using a generalized power iteration algorithm, by applying the constructed matrix. Perform singular value decomposition to update step by step Until it converges.

7. The incremental multi-view data clustering method based on cross-time consensus graphs according to claim 6, characterized in that, The phased variable update also includes: With other variables fixed, update according to formulas (12) and (13). , ,(12) in, ; Separate rows by selecting column indexes that maximize the contribution of the objective function. Solve this problem, let Thus, we obtain formula (13). (13)。 8. An incremental multi-view data clustering system based on cross-time consensus graphs, characterized in that, The system includes: The dynamic consensus graph construction module is used to integrate historical knowledge and current view information based on kernel-induced representation, with a learning time. A consensus affinity matrix is ​​used to construct a dynamic consensus graph; The spectral embedding and discrete label learning module is used for spectral embedding and discrete label learning. The phased variable update module is used to update the variables obtained from the dynamic consensus graph construction module. The consensus affinity matrix, orthogonal rotation matrix, consensus spectrum embedding matrix, and discrete clustering label matrix at each time step are alternately optimized.

9. A machine-readable storage medium storing instructions for causing a machine to perform an incremental multi-view data clustering method based on a cross-temporal consensus graph as described in any one of claims 1-7.

10. A processor, characterized in that, Used to run a program, wherein the program is run to perform the incremental multi-view data clustering method based on cross-time consensus graph as described in any one of claims 1-7.