Government affair data multi-view feature selection method and system based on consensus clustering

Through the multi-view feature selection method of government data based on consensus clustering, the problem of heavy load and poor results in the processing of large-scale government data is solved, and efficient and accurate feature extraction and data integration are achieved.

CN120123718AActive Publication Date: 2025-06-10HUAQIAO UNIVERSITY

Patent Information

Application Number
CN202510583447.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-10
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The existing technology has heavy calculation load when processing large-scale government data, and its effect is poor, limiting the application prospects of multi-perspective feature selection in government data analysis.

Method used

The multi-view feature selection method of government data based on consensus clustering is adopted, and the co-correlation matrix is ​​generated through multiple clustering, and the global consensus matrix is ​​adaptively constructed, combining the regularized feature selection model of L2, 1 norms and Frobenius norms to efficiently integrate multi-view government data information.

Benefits of technology

It significantly improves the accuracy and efficiency of feature extraction, reduces data redundancy and computational complexity, making this method suitable for large-scale government data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123718A_ABST
    Figure CN120123718A_ABST
Patent Text Reader

Abstract

The invention discloses a government affair data multi-view feature selection method and system based on consensus clustering, and relates to the technical field of machine learning, and the method comprises the steps: creating a differentiated basic clustering set for each government affair view through adjusting the initialization strategy and distance measurement parameters of a K-means algorithm, and constructing a co-correlation matrix; analyzing the matrixes by using an adaptive weight mechanism to extract key government affair information, and further constructing and normalizing a global consensus matrix; a multi-view feature selection model is optimized by adopting least square regression, L2, 1 norm regularization and Frobenius norm constraint, and feature selection guidance is carried out through self-paced learning; and combining the optimized model with a global consensus matrix to establish a feature selection objective function, sorting government affair view data, and selecting key features ranking the top k. According to the method, the multi-view feature selection model after self-paced learning is combined with the global consensus matrix, the feature selection target function guided by the consensus matrix is constructed, and key feature selection of government affair view data is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and particularly to a multi-view feature selection method and system for government affairs data based on consensus clustering. Background Art

[0002] Data of government departments is often scattered in multiple systems (such as economy, population, environment, social security, etc.); and has various formats (structured tables, text reports, geographical information, etc.); and these data have privacy and security issues, making it difficult to obtain their available labels. Through multi-view unsupervised feature selection, core correlation features across domains can be extracted to provide data support for decision-making;

[0003] Due to the sensitivity of government affairs data, it may be very difficult to obtain labeled data. Unsupervised learning, especially multi-view unsupervised feature selection, can discover potential patterns in data in the absence of explicit labels, which is also applicable to image recognition, especially in some application scenarios that require automatic recognition but are difficult to annotate (such as large-scale remote sensing image analysis). Using this method can help reduce the need for manual annotation.

[0004] Multi-view feature selection can extract information from different data sources, and each source may provide unique insights, which helps to comprehensively identify and analyze the core factors affecting policies. However, this method faces multiple challenges: it is necessary to balance the information of each data source, evaluate its value, and cope with the complexity brought by high-dimensional data; multi-view feature selection can extract information from different data sources, just like multi-view methods in image recognition can process image data from different perspectives or sensors. For example, in government affairs data analysis, data from a geographic information system (GIS) is parsed through image recognition technology to extract information about land use, environmental changes, etc., and then combined with other data such as economy and population data to comprehensively analyze the impact of policies.

[0005] Current feature selection technologies often have heavy computational loads and poor effects when dealing with large-scale data, which limits the application prospects of multi-view feature selection in government affairs data analysis. Summary of the Invention

[0006] In order to solve the above problems, the present invention proposes a multi-view feature selection method and system for government affairs data based on consensus clustering. By generating a co-correlation matrix through multiple clusterings and adaptively constructing a global consensus matrix, combined with a feature selection model regularized by the L2,1 norm and the Frobenius norm, the multi-view government affairs data information is efficiently integrated, the accuracy and efficiency of feature extraction are improved, and data redundancy and computational complexity are effectively reduced.

[0007] The specific solutions are as follows:

[0008] On the one hand, a multi-view feature selection method for government affairs data based on consensus clustering includes:

[0009] S1. For each government affairs data view, by optimizing the initialization strategy of the clustering center and adjusting the distance metric parameters, generate a set of basic clusters with differences, perform b times of K-means clustering on the set of basic clusters to obtain the clustering results, integrate the clustering results to obtain the basic partition matrix, perform inner product multiplication on the basic partition matrix, and construct the co-correlation matrix of each government affairs view;

[0010] S2. Through the adaptive weight mechanism, learn the co-correlation matrix of each government affairs view to obtain key government affairs information, construct the global consensus matrix based on the key government affairs information, and perform normalization processing on the global consensus matrix to obtain the normalized global consensus matrix;

[0011] S3. Construct a multi-view feature selection model through the least squares regression algorithm, optimize the feature selection matrix in the multi-view feature selection model through L2,1 norm regularization and Frobenius norm constraint to obtain the optimized multi-view feature selection model, and perform feature selection guidance on the optimized multi-view feature selection model through self-paced learning to obtain the multi-view feature selection model after self-paced learning;

[0012] S4. Combine the multi-view feature selection model after self-paced learning with the global consensus matrix to construct a feature selection objective function guided by the consensus matrix;

[0013] S5. Combine the data of each government affairs view through the feature selection objective function, sort the combined government affairs view data in descending order, and select the top k features.

[0014] Furthermore, the basic partition matrix ; where n is the number of data samples; b is the number of clusters; is the cluster generated by clustering;

[0015] The calculation formula of the co-correlation matrix is as follows:

[0016] ;

[0017] Among them, represents the obtained co-correlation matrix.

[0018] Furthermore, in S2, the co-correlation matrix of each government affairs view is learned through the adaptive weight mechanism to obtain key government affairs information, and the global consensus matrix is constructed based on the key government affairs information. The calculation formula is as follows:

[0019] ;

[0020] ;

[0021] Among them, M represents the global consensus matrix; represents the number of government affairs views; represents the number of government affairs views in each cycle; represents the square of the Frobenius norm; represents the global consensus matrix the minimum value of; , represents the adaptive parameter; represents the sum of the elements in each row of the global consensus matrix M; represents the constraint condition.

[0022] Furthermore, in S3, a multi-view feature selection model is constructed through the least squares regression algorithm, and the feature selection matrix in the multi-view feature selection model is optimized by L2,1 norm regularization and Frobenius norm constraint. The calculation formula is as follows:

[0023] ;

[0024] Among them, , represents the data of each government affairs view, n represents the number of government affairs view samples, represents the dimension of each government affairs view sample; represents the feature selection matrix, , used to learn the relationship between features and clusters; represents the clustering indicator matrix, , k is the projection dimension; is a hyperparameter used to control the regularization term; is used to control the trade-off between data fitting and regularization; represents the feature selection matrix ; represents the minimum value of.

[0025] Furthermore, in S3, feature selection guidance is performed on the optimized multi-view feature selection model through self-paced learning to obtain the multi-view feature selection model after self-paced learning, specifically as follows:

[0026] ;

[0027] s ;

[0028] Among them, represents the self-paced learning coefficient, , used to control the number of government affairs data in the next iteration; Denote the selection vector ; The quantity coefficient used to control the government affairs data for the next iteration

[0029] Furthermore, in S4, the multi-view feature selection model after self-paced learning is combined with the normalized global consensus matrix to construct a consensus matrix-guided feature selection objective function, which is specifically as follows:

[0030] Replace the pseudo-constraint with the global consensus matrix M , and construct a consensus matrix-guided feature selection objective function, which is specifically as follows:

[0031] ;

[0032] ; ;

[0033] Among them, , Denote the number of government affairs view samples is the projection dimension; Denote the constraint condition; Denote the square of the Frobenius norm

[0034] On the other hand, a multi-view feature selection method for government affairs data based on consensus clustering includes:

[0035] A co-correlation matrix construction module, which is used for each government affairs data view to generate a differential basic clustering set by optimizing the initialization strategy of the clustering center and adjusting the distance metric parameters, performing b times of K-means clustering on the basic clustering set to obtain a clustering result, integrating the clustering results to obtain a basic partition matrix, and performing inner product multiplication on the basic partition matrix to construct the co-correlation matrix of each government affairs view;

[0036] A global consensus matrix construction module, which is used to learn the co-correlation matrix of each government affairs view through an adaptive weight mechanism, obtain key government affairs information, construct a global consensus matrix based on the key government affairs information, and perform normalization processing on the global consensus matrix to obtain a normalized global consensus matrix;

[0037] A multi-view feature selection model self-paced learning module, which is used to construct a multi-view feature selection model through the least squares regression algorithm, optimize the feature selection matrix in the multi-view feature selection model through L2,1 norm regularization and Frobenius norm constraint to obtain an optimized multi-view feature selection model, and perform feature selection guidance on the optimized multi-view feature selection model through self-paced learning to obtain a multi-view feature selection model after self-paced learning;

[0038] A feature selection objective function construction module, which is used to combine the multi-view feature selection model after self-paced learning with the normalized global consensus matrix to construct a consensus matrix-guided feature selection objective function;

[0039] A feature selection module, which is used to combine each government affairs view data through the feature selection objective function, sort the combined government affairs view data in descending order, and select the top k features.

[0040] The present invention adopts the above technical solutions and has the following beneficial effects:

[0041] (1) The present invention generates a co-correlation matrix through multiple clusterings and adaptively constructs a global consensus matrix, effectively capturing the deep correlation information between multi-view data, thereby significantly improving the accuracy and efficiency of feature extraction;

[0042] (2) The present invention adopts a feature selection model regularized by L2,1 norm and Frobenius norm to ensure the sparsity of the selected features, reduce redundant data and at the same time reduce the computational complexity, making this method particularly suitable for large-scale government affairs data processing;

[0043] (3) The present invention combines the multi-view feature selection model after self-paced learning with the global consensus matrix to construct a consensus matrix-guided feature selection objective function, optimizing the key feature selection of government affairs view data. Description of the Drawings

[0044] Figure 1 It is a flowchart of the multi-view feature selection method for government affairs data based on consensus clustering according to an embodiment of the present invention;

[0045] Figure 2(a) is a convergence graph on the Handwritten dataset according to an embodiment of the present invention;

[0046] Figure 2(b) is a convergence graph on the MSRCV1 dataset according to an embodiment of the present invention;

[0047] Figure 2(c) is a convergence graph on the ORL dataset according to an embodiment of the present invention;

[0048] Figure 3 It is a system diagram of the multi-view feature selection for government affairs data based on consensus clustering according to an embodiment of the present invention. Detailed Embodiment

[0049] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.

[0050] As Figure 1 shown, the multi-view feature selection method for government affairs data based on consensus clustering of the present invention includes:

[0051] S1. For each government affairs data view, by optimizing the initialization strategy of the clustering center and adjusting the distance metric parameters, a basic clustering set with differences is generated. Perform b - times K - means clustering on the basic clustering set to obtain the clustering results, integrate the clustering results to obtain the basic partition matrix, and perform inner - product multiplication on the basic partition matrix to construct the co - correlation matrix of each government affairs view.

[0052] Specifically, the basic partition matrix ; where n is the number of data samples; b is the number of clusters; is the cluster generated by clustering;

[0053] The calculation formula of the co - correlation matrix is as follows:

[0054] ;

[0055] Among them, represents the obtained co - correlation matrix.

[0056] In this embodiment, by adjusting the clustering center initialization strategy and distance metric parameters to generate a basic clustering set with differences and performing b - times kmeans clustering, and then integrating the clustering results to obtain the basic partition matrix, it is to better obtain the connection between government affairs data points.

[0057] S2. Through the adaptive weight mechanism, learn the co - correlation matrix of each government affairs view, obtain the key government affairs information, construct the global consensus matrix based on the key government affairs information, and perform normalization processing on the global consensus matrix to obtain the normalized global consensus matrix.

[0058] Specifically, through the adaptive weight mechanism, learn the co - correlation matrix of each government affairs view to obtain the key government affairs information, and construct the global consensus matrix based on the key government affairs information. The calculation formula is as follows:

[0059] ;

[0060]

[0061] Among them, M represents the global consensus matrix; represents the number of government affairs views; represents the number of government affairs views in each cycle; represents the square of the Frobenius norm; represents the global consensus matrix the minimum value of; , represents the adaptive parameter; represents the sum of each row of the global consensus matrix M; Represents a constraint condition.

[0062] S3. Construct a multi-view feature selection model through the least squares regression algorithm. Optimize the feature selection matrix in the multi-view feature selection model through L2,1 norm regularization and Frobenius norm constraint to obtain an optimized multi-view feature selection model. Conduct feature selection guidance on the optimized multi-view feature selection model through self-paced learning to obtain the multi-view feature selection model after self-paced learning.

[0063] Specifically, construct a multi-view feature selection model through the least squares regression algorithm. Optimize the feature selection matrix in the multi-view feature selection model through L2,1 norm regularization and Frobenius norm constraint. The calculation formula is as follows:

[0064] ;

[0065] Among them, , represents the data of each government affairs view, n represents the number of government affairs view samples, represents the dimension of each government affairs view sample; represents the feature selection matrix, , used to learn the relationship between features and clusters; represents the clustering indicator matrix, , k is the projection dimension; is a hyperparameter, used to control the regularization term; is used to control the trade-off between data fitting and regularization; represents the for the feature selection matrix ; represents the minimum value of.

[0066] Specifically, conduct feature selection guidance on the optimized multi-view feature selection model through self-paced learning to obtain the multi-view feature selection model after self-paced learning, as follows:

[0067] ;

[0068] s ;

[0069] Among them, represents the self-paced learning coefficient, , used to control the number of government affairs data in the next iteration; represents the selection vector, ; is the coefficient used to control the number of government affairs data in the next iteration. If the government affairs data sample is in the projection matrix and the pseudo-label If they are consistent, the pseudo-labels are considered credible. Conversely, if the gap between the two is large, it indicates that the credibility of the pseudo-labels is low and they are not suitable for participating in the next iteration.

[0070] S4. Combine the multi-view feature selection model after self-paced learning with the global consensus matrix to construct a consensus matrix-guided feature selection objective function.

[0071] Specifically, combine the multi-view feature selection model after self-paced learning with the normalized global consensus matrix to construct a consensus matrix-guided feature selection objective function as follows:

[0072] Replace the pseudo-constraint with the global consensus matrix M and construct a consensus matrix-guided feature selection objective function as follows:

[0073] ;

[0074] ; ;

[0075] where , represents the number of government affairs view samples, is the projection dimension; represents the constraint condition; represents the square of the Frobenius norm.

[0076] S5. Combine each government affairs view data through the feature selection objective function, sort the combined government affairs view data in descending order, and select the top k features.

[0077] Specifically, the multi-view unsupervised feature selection method for government affairs data based on consensus clustering of the present invention realizes the efficient selection of multi-source data through the application of consensus clustering learning technology, constructs a model with excellent performance and high operation efficiency. Effectively integrates the similarity information from multiple views, while avoiding the deviation of a single view; combines L2,1 norm and Frobenius norm regularization to enhance the sparsity of features, prevent overfitting, and improve the stability of the model.

[0078] The experimental data sets used in the present invention mainly consist of a handwritten data set (Handwritten), a facial image data set (ORL), a Yale face data set (Yale), an outdoor scene data set (Outdoor Scene), a web data (WebKB), and 7 different category images (MSRCV1). For the evaluation of clustering performance, two widely used metrics are adopted, namely, normalized mutual information (NMI) and accuracy (ACC). NMI [0,1] normalizes the standard mutual information, where 0 indicates that the two partitions are completely unrelated and 1 indicates that the two partitions are consistent. ACC [0,1] is a metric for evaluating classification models.

[0079] Table 1 Comparative analysis of normalized mutual information measurements;

[0080]

[0081] According to the results in Table 1, the method of the present invention ranks first in the evaluation based on the NMI metric on six datasets (MSRVCV1, Yale, Handwritten, Outdoor scene, ORL, and WebKB), and has excellent performance and stability in the multi-view clustering task, and can effectively extract key information and perform feature selection.

[0082] Table 2 Comparative analysis of accuracy measurements;

[0083]

[0084] According to the results in Table 2, the corresponding method of the present invention performs excellently in the evaluation based on specific metrics on six datasets (MSRVCV1, Yale, Handwritten, Outdoor scene, ORL, and WebKB), and obtains the highest average value. Specifically, in the MSRVCV1, Handwritten, and WebKB datasets, the corresponding methods of the present invention achieve excellent results of 83.14 ± 0.79, 92.57 ± 0.63, and 70.15 ± 1.38 respectively; in the Yale, Outdoorscene, and ORL datasets, high scores of 57.82 ± 1.85, 61.76 ± 0.66, and 62.80 ± 1.66 are also obtained respectively. This shows that the method of the present invention has significant advantages in the multi-view feature selection task and can effectively improve the accuracy and reliability of data processing.

[0085] In addition, to further study the efficiency of the method, the present invention conducts a convergence study. Convergence graphs of the method on three benchmark datasets are shown: Handwritten, MSRCV1, and WebKB; as shown in Figures 2(a), 2(b), and 2(c) respectively. As shown, the curves all demonstrate the convergence process of the optimization algorithm on different datasets as the number of iterations increases, where the objective function value rapidly decreases and gradually levels off: in the initial few iterations, a larger step size or an effective search strategy is used, and the function value quickly drops to a fraction of its original value from a relatively high position; subsequently, it enters a refinement stage where the descent speed slows down and is accompanied by small oscillations, indicating that the algorithm is adaptively adjusting the step size to accurately approximate the optimum; finally, after dozens of iterations, it basically stops decreasing, meaning that the algorithm has converged and obtained a stable solution. This confirms that the method corresponding to the present invention is computationally effective and can effectively handle large-scale datasets.

[0086] In summary, the present invention proposes a multi-view unsupervised feature selection method for government affairs data based on consensus learning. It effectively integrates similarity information from multiple views through consensus clustering and learning for each view, and constructs a unified global consensus matrix through an adaptive mechanism to capture the deep correlation information between multi-views. At the same time, it avoids the bias of a single view. In addition, we combine L2,1 norm and Frobenius norm regularization to enhance the sparsity of features, prevent overfitting, and improve the stability of the model. A large number of experiments verify the superior performance of this method, prove the effectiveness of this method, and demonstrate the good prospect of the innovative application of consensus learning in the multi-view feature unsupervised selection task.

[0087] Generally speaking, a multi-view feature selection method and system for government affairs data based on consensus clustering. The core lies in realizing the effective analysis and feature extraction of government affairs data through a series of steps to improve the quality of decision-making support. First, for each government affairs data view, an optimized K-means clustering strategy is adopted to generate a basic partition matrix, and the co-correlation matrix is calculated to capture the internal relationship between different views (S1). Then, an adaptive weight mechanism is used to learn key information from the co-correlation matrices of each view, construct and normalize the global consensus matrix, strengthening the cross-view information integration ability (S2). Next, the least squares regression algorithm is combined with L2,1 norm regularization and Frobenius norm constraint to optimize the feature selection model, and a self-paced learning mechanism is introduced to guide the feature selection process, improving the robustness and generalization ability of the model (S3). After that, the model after self-paced learning is combined with the global consensus matrix to construct a feature selection objective function, achieving more accurate feature selection (S4). Finally, the government affairs view data is sorted and selected according to the constructed objective function to ensure that the most representative features are finally selected (S5). By accurately calculating the basic partition matrix and co-correlation matrix to enhance the expression of differences and correlations between different view data, using the adaptive weight mechanism to optimize the construction of the global consensus matrix to accurately capture key government affairs information, adopting the least squares regression combined with L2,1 norm and Frobenius norm constraints to optimize the feature selection model, and guiding the feature selection process step by step through the self-paced learning mechanism, thereby ensuring the robustness and effectiveness of the model when dealing with multi-view government affairs data. Finally, the model after self-paced learning is combined with the normalized global consensus matrix to form a feature selection objective function guided by the consensus matrix, realizing the efficient and accurate selection of government affairs view data. Overall, these technical effects work together to improve the accuracy, efficiency of feature selection in government affairs data analysis and the generalization ability of the model.

[0088] As Figure 3 shown, this embodiment also discloses a multi-view unsupervised feature selection system for government affairs data based on consensus learning, including:

[0089] A co-correlation matrix construction module 31, for each government affairs data view, by optimizing the initialization strategy of the clustering center and adjusting the distance metric parameters, generating a basic clustering set with differences, performing b times of K-means clustering on the basic clustering set to obtain a clustering result, integrating the clustering results to obtain a basic partition matrix, and performing inner product multiplication on the basic partition matrix to construct the co-correlation matrix of each government affairs view;

[0090] The global consensus matrix construction module 32 is used to learn the co - correlation matrix of each government affairs view through an adaptive weight mechanism, obtain key government affairs information, construct a global consensus matrix based on the key government affairs information, and perform normalization processing on the global consensus matrix to obtain a normalized global consensus matrix;

[0091] The multi - view feature selection model self - paced learning module 33 is used to construct a multi - view feature selection model through the least - squares regression algorithm, optimize the feature selection matrix in the multi - view feature selection model through L2,1 - norm regularization and Frobenius - norm constraint to obtain an optimized multi - view feature selection model, and perform feature selection guidance on the optimized multi - view feature selection model through self - paced learning to obtain a multi - view feature selection model after self - paced learning;

[0092] The feature selection objective function construction module 34 is used to combine the multi - view feature selection model after self - paced learning with the normalized global consensus matrix to construct a feature selection objective function guided by the consensus matrix;

[0093] The feature selection module 35 is used to combine the data of each government affairs view through the feature selection objective function, sort the combined government affairs view data in descending order, and select the top k features.

[0094] Although the present invention has been specifically shown and described with reference to the preferred embodiments, those skilled in the art should understand that various changes in form and detail can be made to the present invention without departing from the spirit and scope of the present invention as defined by the appended claims, and all such changes are within the scope of protection of the present invention.

Claims

1. A multi-view feature selection method for government data based on consensus clustering, characterized in that: include: S1, for each government data view, generate a basic clustering set with differences by optimizing the initialization strategy of the cluster center and adjusting the distance measurement parameters, perform K-means clustering on the basic clustering set b times to obtain the clustering results, integrate the clustering results to obtain the basic partition matrix, perform inner product multiplication on the basic partition matrix, and construct the co-correlation matrix of each government view; S2, learn the correlation matrix of each government view through the adaptive weight mechanism, obtain key government information, build a global consensus matrix based on the key government information, normalize the global consensus matrix, and obtain the normalized global consensus matrix; S3, construct a multi-view feature selection model through the least squares regression algorithm, optimize the feature selection matrix in the multi-view feature selection model through L2,1 norm regularization and Frobenius norm constraint to obtain an optimized multi-view feature selection model, guide feature selection of the optimized multi-view feature selection model through self-paced learning, and obtain a multi-view feature selection model after self-paced learning; S4, combining the multi-view feature selection model after self-paced learning with the global consensus matrix to construct a feature selection objective function guided by the consensus matrix; S5, combining each government affairs view data through the feature selection objective function, sorting the combined government affairs view data in descending order, and selecting the top k features.

2. The consensus clustering-based multi-view feature selection method for government data according to claim 1 is characterized in that: In S1, the basic partition matrix ; Where n is the number of data samples; b is the number of clusters; Clusters generated for clustering; The calculation formula of the co-correlation matrix is ​​as follows: ; in, represents the obtained co-correlation matrix.

3. The consensus clustering-based multi-view feature selection method for government data according to claim 2 is characterized in that: In S2, the correlation matrix of each government view is learned through the adaptive weight mechanism to obtain key government information, and a global consensus matrix is ​​constructed based on the key government information. The calculation formula is as follows: ; ; Where M represents the global consensus matrix; Indicates the number of government views; Indicates the number of government views per cycle; represents the square of the Frobenius norm; Represents the global consensus matrix The minimum value of , represents the adaptive parameters; represents the sum of each row of the global consensus matrix M; Represents a constraint.

4. The consensus clustering-based multi-view feature selection method for government data according to claim 1 is characterized in that: In S3, a multi-view feature selection model is constructed by the least squares regression algorithm. The feature selection matrix in the multi-view feature selection model is optimized by L2,1 norm regularization and Frobenius norm constraint. The calculation formula is as follows: ; in, , represents the data of each government view, n represents the number of government view samples, Represents the sample dimension of each government view; represents the feature selection matrix, , used to learn the relationship between features and clusters; represents the cluster indicator matrix, , k is the projection dimension; is a hyperparameter used to control the regularization term; Used to control the trade-off between data fitting and regularization; Represents the feature selection matrix conduct ; express The minimum value of .

5. The consensus clustering-based multi-view feature selection method for government data according to claim 4 is characterized in that: In S3, the optimized multi-view feature selection model is guided by feature selection through self-paced learning to obtain the multi-view feature selection model after self-paced learning, as follows: ; s ; in, represents the self-paced learning coefficient, , used to control the amount of government data in the next iteration; represents the selection vector, ; Used to control the quantity coefficient of government data in the next iteration.

6. The consensus clustering-based multi-view feature selection method for government data according to claim 5 is characterized in that: In S4, the multi-view feature selection model after self-paced learning is combined with the normalized global consensus matrix to construct a feature selection objective function guided by the consensus matrix, as follows: Replacing pseudo-constraints with a global consensus matrix M , construct the feature selection objective function guided by the consensus matrix, as follows: ; ; ; in, , Indicates the number of government view samples, is the projection dimension; Indicates constraints; Represents the square of the Frobenius norm.

7. A multi-view feature selection system for government data based on consensus clustering, characterized in that: include: The co-correlation matrix construction module is used to generate a basic clustering set with differences for each government data view by optimizing the initialization strategy of the cluster center and adjusting the distance measurement parameters, perform K-means clustering on the basic clustering set b times to obtain the clustering results, integrate the clustering results to obtain the basic partition matrix, perform inner product multiplication on the basic partition matrix, and construct the co-correlation matrix of each government view; The global consensus matrix construction module is used to learn the correlation matrix of each government view through an adaptive weight mechanism, obtain key government information, build a global consensus matrix based on the key government information, normalize the global consensus matrix, and obtain the normalized global consensus matrix; The multi-view feature selection model self-paced learning module is used to construct a multi-view feature selection model through a least squares regression algorithm, optimize the feature selection matrix in the multi-view feature selection model through L2,1 norm regularization and Frobenius norm constraint to obtain an optimized multi-view feature selection model, and guide feature selection of the optimized multi-view feature selection model through self-paced learning to obtain a multi-view feature selection model after self-paced learning; A feature selection objective function construction module is used to combine the multi-view feature selection model after self-paced learning with the normalized global consensus matrix to construct a feature selection objective function guided by the consensus matrix; The feature selection module is used to combine each government affairs view data through the feature selection objective function, sort the combined government affairs view data in descending order, and select the top k features.

Citation Information

Patent Citations

  • Unsupervised multi-view feature selection method and system based on low-rank tensor learning

    CN114549916A

  • Semi-supervised multi-view clustering integration method and system based on width learning

    CN119479047A

  • Pneumovirus gene data multi-view clustering integration method and device and electronic equipment

    CN119513631A

  • Consensus graph learning-based multi-view clustering method

    US20240143699A1

Cited By

  • Financial data multi-source feature selection method and system based on self-paced tensor learning

    CN120705533A