Compressed anchor diagram-based incomplete multi-view semi-supervised classification method and system

By using compressed anchor maps in the incomplete multi-view semi-supervised classification method, the problems of high computational complexity and difficulty in dealing with missing views in the prior art are solved, and efficient label prediction and classification accuracy are achieved.

CN120107641APending Publication Date: 2025-06-06SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311656211.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle incomplete multi-view data in the case of missing arbitrary views under the semi-supervised learning framework, and the calculation complexity is high, especially in dense matrix inversion operations.

Method used

A complete multi-view semi-supervised classification method based on compressed anchor point map is proposed. By extending the anchor point map matrix, the average is obtained, and the initial fusion anchor point map is obtained, and the class label is used to compress it, and iteratively predict labels and anchor point map compression is avoided.

Benefits of technology

Effectively dealing with incomplete multi-view data in the case of missing arbitrary views improves computing efficiency and achieves rapid label prediction and classification accuracy improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107641A_ABST
    Figure CN120107641A_ABST
Patent Text Reader

Abstract

The invention discloses an incomplete multi-view semi-supervised classification method and system based on a compressed anchor point diagram, which are applied to the technical field of artificial intelligence, and the method comprises the following steps: calculating an average value based on an extended anchor point diagram matrix of original multi-view data to obtain an initial fusion anchor point diagram; compressing the initial fusion anchor point diagram based on a class label to obtain a compressed anchor point diagram; predicting class labels and confidence scores of the unlabeled samples based on the compressed anchor point diagram, and combining the class labels of the unlabeled samples with high confidence scores with the class labels of the known label samples to serve as known class labels; carrying out label prediction and anchor point diagram compression learning through iteration to judge an anchor point diagram; calculating the classification precision on the label-free sample; the invention provides a fast label prediction strategy based on a compressed anchor diagram. And through iterative label prediction and anchor point diagram compression, the learned anchor point diagram has better discrimination, and the classification precision of the model can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an incomplete multi-view semi-supervised classification method and system based on a compressed anchor graph. Background Art

[0002] Anchor graphs are widely used in single-view and multi-view clustering tasks [1,2], but are rarely used in semi-supervised multi-view learning problems. Recently, researchers have proposed some anchor graph-based models to accelerate the computational efficiency of semi-supervised classification models [3,4]. However, these models all use k-means clustering to generate anchors [5] and are unable to use partial labeled data to learn discriminative anchor graphs. Current graph-based or anchor graph-based models usually need to calculate the inverse matrix of dense matrices during the optimization process [6,7], and the computational complexity of this operation grows rapidly with the increase in the number of data samples. Existing methods cannot flexibly deal with incomplete multi-view data with any missing views in a semi-supervised learning framework.

[0003] Nie et al. proposed a semi-supervised multi-view classification method based on low-rank constrained similarity graph [7]. While learning the similarity matrix, the similarity matrix is ​​low-rank constrained so that the similarity matrix has exactly the number of connected components of the category. When predicting the category of unlabeled samples, this method needs to invert the dense Laplacian matrix corresponding to the unlabeled data. As the number of unlabeled samples increases, the computational complexity increases rapidly. Based on reference [7], Shi et al. [6] weighted the sample features during the learning process of the similarity graph. This method also needs to calculate the inverse matrix of the dense Laplacian matrix, and because the inverse matrix needs to be calculated multiple times during the iteration process, the computational complexity is higher than the method in reference [7]. Zhang et al. proposed a semi-supervised multi-view classification algorithm based on anchor graph [3]. This method constrains the consistent anchor graph to be a linear combination of multiple private anchor graphs, and constructs a bipartite graph based on anchors and sample data. At the same time, the bipartite graph is low-rank constrained, and supervised constraints are imposed on some labeled data. This method predicts the category labels of anchor points and original data at the same time. During the iteration process, it is necessary to calculate the inverse matrix of the dense matrix multiple times, so the computational complexity is still very high. At the same time, the above method needs to use a clustering algorithm to generate an anchor graph, and usually in order to ensure the prediction accuracy, it is necessary to ensure that the number of anchor points is large enough. This is related to the fact that the anchor graph itself does not have sufficient discriminability, that is, the label information is not used in the anchor graph construction process. In addition, the increase in the number of anchor points means higher computational complexity. Wang et al. proposed a semi-supervised classification algorithm based on bipartite graphs [8]. This model also imposes low-rank constraints on the bipartite graph based on anchor points and sample data, and uses the same supervised constraint terms. This model is mainly aimed at single-view data. Its shortcomings are similar to the method in reference [3]. In reference [9], Wang et al. proposed a kernel-based label propagation semi-supervised multi-view classification model, which reduces the computational complexity from O(N 3 ) is reduced to O(N 2 ), however, its computational efficiency is still limited by the number of iterations due to the need for multiple iterations. Reference

[10] proposed a semi-supervised multi-view classification model based on multi-image fusion linear regression. During the optimization process, it is necessary to solve the inverse matrix of a dense matrix with a scale equal to the size of the dataset, and the computational complexity is O(N 3 He et al. proposed a fast semi-supervised learning model based on anchor graph for hyperspectral classification

[11] . However, this method can only process single-view data and cannot process multi-view data. Currently, there is no semi-supervised multi-view classification model that can effectively process incomplete multi-view data with any missing views.

[0004] [1]Xinlei Chen and Deng Cai,“Large scale spectral clustering withlandmark-based representation,”in Proceedings of the AAAI Conference onArtificial Intelligence,2011,vol.25,pp.313-318.

[0005] [2]Suyuan Liu,Siwei Wang,Pei Zhang,Kai Xu,Xinwang Liu,ChangwangZhang,and Feng Gao,“Efficient onepass multi-view subspace clustering withconsensus anchors,”in Proceedings of the AAAI Conference on ArtificialIntelligence,2022,vol.36,pp.7576-7584.

[0006] [3]Bin Zhang,Qianyao Qiang,Fei Wang,and Feiping Nie,“Fast multi-viewsemi-supervised learning with learned graph,”IEEE Transactions on Knowledgeand Data Engineering,vol.34,no.1,pp.286-299,2020.

[0007] [4]Senhong Wang,Jiangzhong Cao,Fangyuan Lei,Qingyun Dai,ShangsongLiang,and Bingo WingKuen Ling,“Semi-supervised multi-view clustering withweighted anchor graph embedding,”Computational Intelligence and Neuroscience,vol.2021,2021.

[0008] [5]Xiao Yu,Hui Liu,Yuxiu Lin,Yan Wu,and Caiming Zhang,“Auto-weightedsample-level fusion with anchors for incomplete multi-view clustering,”Pattern Recognition,vol.130,pp.108772,2022.

[0009] [6]Shaojun Shi,Feiping Nie,Rong Wang,and Xuelong Li,“Semi-supervisedlearning based on intra-view heterogeneity and inter-view compatibility forimage classification,”Neurocomputing,vol.488,pp.248-260,2022.

[0010] [7]Feiping Nie,Guohao Cai,Jing Li,and Xuelong Li,“Auto-weightedmulti-view learning for image clustering and semi-supervised classification,”IEEE Transactions on Image Processing,vol.27,no.3,pp.1501-1511,2017.

[0011] [8]Zhen Wang,Long Zhang,Rong Wang,Feiping Nie,and Xuelong Li,“Semi-supervised learning via bipartite graph construction with adaptiveneighbors,”IEEE Transactions on Knowledge and Data Engineering,vol.35,pp.5257-5268,2023.

[0012] [9]Shiping Wang, Zhewen Wang, and Wenzhong Guo, "Accelerated manifoldembedding for multi-view semisupervised classification," Inf.Sci., vol.562, pp.438-451, 2021.

[0013]

[10] Zhongheng Li, Qianyao Qiang, Bin Zhang, Fei Wang, and Feiping Nie, "Flexible multi-view semi-supervised learning with unified graph," Neuralnetworks: the official journal of the International Neural Network Society, vol.142, pp.92-104, 2021.

[0014]

[11] Fang He, Rong Wang, and W.Jia, "Fast semi-supervised learning with anchor graph for large hyperspectral images," Pattern Recognit. Lett., vol.130, pp.319-326, 2020.

[0015] To address the above problems, this application proposes an incomplete multi-view semi-supervised classification method and system based on compressed anchor graph, which can handle incomplete multi-view data with any view missing, and avoids dense matrix inversion operations in the prediction process, effectively improving the computational efficiency of the model. Summary of the invention

[0016] The purpose of this application is to provide an incomplete multi-view semi-supervised classification method and system based on compressed anchor graph, aiming to solve the above-mentioned problems.

[0017] To achieve the above objectives, this application provides the following technical solutions:

[0018] This application provides an incomplete multi-view semi-supervised classification method based on a compressed anchor graph, including:

[0019] The expanded anchor map matrix of the original multi-view data is averaged to obtain the initial fused anchor map;

[0020] Compressing the initial fused anchor point graph based on the class label to obtain a compressed anchor point graph;

[0021] Predicting the class labels and confidence scores of the unlabeled samples based on the compressed anchor graph, and combining the class labels of the unlabeled samples with high confidence scores with the class labels of the known label samples as known class labels;

[0022] Learn the discriminative anchor map by iteratively performing label prediction and anchor map compression;

[0023] Calculate the classification accuracy on unlabeled samples.

[0024] Furthermore, in the step of calculating the average of the extended anchor map matrix based on the original multi-view data to obtain the initial fused anchor map, the following steps are specifically included:

[0025] Calculate the rows and columns of the expanded anchor map matrix:

[0026]

[0027] Among them G v P v G vT To transform the matrix P v The number of rows and columns of is expanded from the number of samples in each view to the size n of the complete data;

[0028] Based on the average of the extended anchor map matrix, the initial fused anchor map is obtained:

[0029]

[0030] Where V is the number of views.

[0031] Furthermore, in the step of compressing the initial fused anchor graph based on the class label to obtain a compressed anchor graph, the following steps are specifically included:

[0032] Based on the class label, the initial fusion anchor map P ave To compress, the shape of the anchor map matrix changes from n v ×n v becomes n v ×C, that is:

[0033] P=S k (P ave )

[0034] Where P = [P 1 ,...,P c ,...,P C ], C is the number of categories, c = 1, 2, ..., C;

[0035] Specifically,

[0036] Where N(c) is the sample index set of the cth class; |N(c)| is the set size; (P ave ) ·j is the matrix P ave The jth column of

[0037] The first iteration is recorded as t = 0, and the initial fusion anchor graph P is ave Compress to obtain the compressed anchor point map P 0 ;

[0038] If t≠0, update r=min(r0+2×(t-1),100) and update the initial fusion anchor graph P based on the class labels generated by the model. ave Compress to obtain the compressed anchor point map P t .

[0039] Furthermore, in the step of predicting the class label and confidence score of the unlabeled sample based on the compressed anchor graph, combining the class label of the unlabeled sample with a high confidence score with the class label of the known label sample as the known class label, the following steps are specifically included:

[0040] The compressed anchor graph is divided into sub-anchor graphs P corresponding to labeled samples lm And the sub-anchor graph P corresponding to the unlabeled sample um ;

[0041] Based on the sub-anchor graph P lm and label matrix Q l Calculate the anchor soft label matrix F m ;

[0042] Based on the sub-anchor graph P um and anchor soft label matrix F m Calculate the soft label matrix F for unlabeled data u , and get the corresponding confidence score s i With label y ur ;

[0043] The confidence score s i Sorting is performed, and the class labels of unlabeled samples with high confidence scores are combined with the class labels of known label samples as known class labels.

[0044] Furthermore, based on the sub-anchor graph P lm and label matrix Q l Calculate the anchor soft label matrix F m The steps specifically include the following steps:

[0045] Assume that the first l samples of the complete data are known label samples, and the last u samples are unknown label samples, n = l + u; the compressed anchor point graph is recorded as P = [P lm ;P um ]; the bipartite graph with labeled data and anchor points is:

[0046]

[0047] Minimize the following issues:

[0048]

[0049] Among them, F s1 =[F l ; F m ]; L s1 is the Laplace matrix;

[0050] The solution to the minimization problem is:

[0051]

[0052] Where Q l is the label matrix of known data, D lm_r and D lm_c is a diagonal matrix.

[0053] Furthermore, based on the sub-anchor graph P um and anchor soft label matrix F m Calculate the soft label matrix F for unlabeled data u , and get the corresponding confidence score s i With label y ur The steps specifically include the following steps:

[0054] The bipartite graph of unlabeled data and anchor points is:

[0055]

[0056] Minimize the following issues:

[0057]

[0058] Among them, F s2 =[F u ; F m ]; L s2 is the Laplace matrix;

[0059] The solution to the minimization problem is:

[0060]

[0061] Where D um_r and D um_cis a diagonal matrix;

[0062] The calculation formula for the label of unlabeled data is:

[0063]

[0064] in is the confidence score of the unlabeled data; and F = [F l ; F u ],y=[y l ;y u_pre ].

[0065] Furthermore, when the confidence score s i The step of sorting and combining the class labels of the unlabeled samples with high confidence scores with the class labels of the known label samples as the known class labels specifically includes the following steps:

[0066] Sort the confidence scores of the unlabeled data and retain the data labels y with the top r% confidence scores ur , combined with the label of the known labeled sample as the known class label, update t=t+1 until r=100.

[0067] This application proposes an incomplete multi-view semi-supervised classification system based on compressed anchor graph, including:

[0068] Calculation module: Calculate the average of the extended anchor map matrix based on the original multi-view data to obtain the initial fused anchor map;

[0069] Compression module: compressing the initial fusion anchor map based on the class label to obtain a compressed anchor map;

[0070] Judgment module: predicting the class label and confidence score of the unlabeled sample based on the compressed anchor point graph, and combining the class label of the unlabeled sample with a high confidence score with the class label of the known label sample as the known class label;

[0071] Iteration module: learns to discriminate anchor graphs through iterative label prediction and anchor graph compression;

[0072] Classification module: Calculates the classification accuracy on unlabeled samples.

[0073] The present application provides a device, which includes a processor and a memory coupled to the processor, wherein the memory stores program instructions for implementing an incomplete multi-view semi-supervised classification method based on a compressed anchor graph; the processor is used to execute the program instructions stored in the memory to implement incomplete multi-view semi-supervised classification based on a compressed anchor graph.

[0074] The present application provides a storage medium storing program instructions executable by a processor, wherein the program instructions are used to execute an incomplete multi-view semi-supervised classification method based on a compressed anchor graph.

[0075] The present application provides an incomplete multi-view semi-supervised classification method and system based on a compressed anchor graph, which has the following beneficial effects:

[0076] Based on all samples of each incomplete view data, an anchor map matrix is ​​constructed, and the rows and columns of the anchor map matrix are expanded to achieve alignment of anchor maps of different views at the sample level. The initial fused anchor map is obtained by averaging; the anchor map is compressed, and label prediction and anchor map compression are iteratively performed to learn a discriminative anchor map; the prediction of unlabeled samples is split into two steps. The prediction of unlabeled samples is split into two steps. First, the soft label matrix F of the anchor point is calculated. m , and then based on the soft label matrix F m Calculate the soft label matrix F of the unlabeled samples u And get the predicted label y u_pre , thus avoiding the calculation of the inverse matrix of the dense matrix during the label prediction process and achieving fast label prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 This is a flow chart of an incomplete multi-view semi-supervised classification method based on a compressed anchor graph according to Example 1 of the present application;

[0078] Figure 2 This is a structural diagram of an incomplete multi-view semi-supervised classification system based on a compressed anchor graph according to Example 2 of the present application;

[0079] Figure 3 This is a schematic diagram of the device structure of Example 3 of the present application;

[0080] Figure 4 This is a schematic diagram of the storage medium structure of Example 4 of the present application. DETAILED DESCRIPTION

[0081] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0082] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0083] Example 1

[0084] See also Figure 1 , is a flow chart of an incomplete multi-view semi-supervised classification method based on a compressed anchor graph according to Example 1 of the present application; the specific steps include:

[0085] S1: The expanded anchor map matrix of the original multi-view data is averaged to obtain the initial fused anchor map.

[0086] In this embodiment, the data matrix of the vth view in the original multi-view data is (v=1,2,...V, V is the number of views), m v is the dimension of the v-th view sample, n v is the number of samples of the vth view;

[0087] Each view data itself is used as the anchor point set, that is, the anchor point matrix of the view is The anchor point graph matrix P between the data and the anchor points is obtained by minimizing the objective function v , the matrix size is n v ×n v ; The specific formula is:

[0088]

[0089] Further constrain each row to have only k non-zero elements, and solve it row by row as follows

[0090]

[0091] in,

[0092] Define the corresponding missing indicator matrix G according to the position of the missing samples in each view v , if the i-th sample in the complete data is the j-th sample in the missing data, then otherwise

[0093] In order to align the anchor map matrices of different views at the sample level, the rows and columns of the extended anchor map matrix are calculated:

[0094]

[0095] Among them G v P v G vT To transform the matrix P v The number of rows and columns of is expanded from the number of samples in each view to the size n of the complete data;

[0096] Based on the average of the extended anchor map matrix, the initial fused anchor map is obtained:

[0097]

[0098] Where V is the number of views.

[0099] S2: compressing the initial fusion anchor map based on the class label to obtain a compressed anchor map.

[0100] In this embodiment, the initial fusion anchor graph P is converted into ave To compress, the shape of the anchor map matrix changes from n v ×n v becomes n v ×C, that is:

[0101] P=S k (P ave )

[0102] Where P = [P 1 ,...,P c ,...,P C ], C is the number of categories, c = 1, 2, ..., C;

[0103] Specifically,

[0104] Where N(c) is the sample index set of the cth class; |N(c)| is the set size; (P ave ) ·j is the matrix P ave The jth column of

[0105] The first iteration is recorded as t = 0, and the initial fusion anchor graph P is ave Compress to obtain the compressed anchor point map P 0 ;

[0106] If t≠0, update r=min(r0+2×(t-1),100) and update the initial fusion anchor graph P based on the class labels generated by the model. ave Compress to obtain the compressed anchor point map P t .

[0107] S3: Predicting the class label and confidence score of the unlabeled sample based on the compressed anchor point graph, and combining the class label of the unlabeled sample with a high confidence score with the class label of the known label sample as the known class label.

[0108] In this embodiment, in the step of predicting the class label and confidence score of the unlabeled sample based on the compressed anchor point map, combining the class label of the unlabeled sample with a high confidence score with the class label of the known label sample as the known class label, specifically includes steps S31 to S34, specifically including the following contents.

[0109] S31: Divide the compressed anchor graph into sub-anchor graphs P corresponding to labeled samples lm And the sub-anchor graph P corresponding to the unlabeled sample um .

[0110] S32: Based on the sub-anchor graph P lm and label matrix Q l Calculate the anchor soft label matrix F m .

[0111] Assume that the first l samples of the complete data are known label samples, and the last u samples are unknown label samples, n = l + u; the compressed anchor point graph is recorded as P = [P lm ;P um ]; the bipartite graph with labeled data and anchor points is:

[0112]

[0113] Minimize the following issues:

[0114]

[0115] Among them, F s1 =[F l ; F m ]; L s1 is the Laplace matrix, specifically:

[0116]

[0117] in is the degree matrix, D lm_r is a diagonal matrix, and (D lm_r ) ii =∑ j (P lm ) ij , D lm_c is a diagonal matrix, and (D lm_c ) jj =∑ i (P lm ) ij .

[0118] The solution to the minimization problem is:

[0119]

[0120] Where Q l is the label matrix of known data, since D lm_r and D lm_c is a diagonal matrix, the soft label matrix F of the anchor graph m Can be calculated quickly.

[0121] S33: Based on the sub-anchor graph P um and anchor soft label matrix F m Calculate the soft label matrix F for unlabeled data u , and get the corresponding confidence score s i With label y ur .

[0122] The bipartite graph of unlabeled data and anchor points is:

[0123]

[0124] Minimize the following issues:

[0125]

[0126] Among them, F s2 =[F u ; F m ]; L s2 is the Laplace matrix, specifically:

[0127]

[0128] The degree matrix D um_r is a diagonal matrix, and (D um_r ) ii =∑ j (P um ) ij , D um_c is a diagonal matrix, and (D um_c ) jj =∑ i (P um ) ij .

[0129] The solution to the minimization problem is:

[0130]

[0131] Due to D um_r and D um_c is a diagonal matrix, the soft label matrix F of the unlabeled sample u Can be calculated quickly.

[0132] The calculation formula for the label of unlabeled data is:

[0133]

[0134] in is the confidence score of the unlabeled data; and F = [F l ; F u ],y=[y l;y u_pre ].

[0135] S34: Set the confidence score s i Sorting is performed, and the class labels of unlabeled samples with high confidence scores are combined with the class labels of known label samples as known class labels.

[0136] Sort the confidence scores of the unlabeled data and retain the data labels y with the top r% confidence scores ur , combined with the label of the known label sample as the known class label, update t = t + 1; iterate step S31 to step S33 until r = 100; avoid calculating the inverse matrix of the dense matrix in the label prediction process, and achieve fast label prediction. .

[0137] S4: Calculate the classification accuracy on unlabeled samples.

[0138] In this embodiment, the classification accuracy AC of unlabeled samples is calculated as follows:

[0139]

[0140] Among them, y u_pre is the predicted label vector of the unlabeled sample, y u is the actual label vector of the unlabeled sample, and u is the number of unlabeled samples.

[0141] In one embodiment, the present application has been experimentally verified. The classification performance and computational efficiency of the method proposed in the present application are compared with several popular algorithms on incomplete multi-view data, and the results are shown in the following table; Table 1 is a comparison of the classification performance and computational time of the method of the present application and several representative semi-supervised multi-view classification models on ECG data with a missing ratio of 35%.

[0142]

[0143]

[0144] Table 2 compares the classification performance and computation time of the method of the present application with several current representative semi-supervised multi-view classification models on the ALOI100 data with a missing ratio of 30%.

[0145] Missing proportion 30% (average of 10 results) 30% (1 operation time) AC(%) Time(s) SSC_IHIC 58.65 4116.82 MLAN 13.25 330.62 CONMF 73.88 301.01 LACK 31.38 357.65 GPSNMF 67.74 26.04 MVAR 49.34 85.36 SAGIMCSSL 79.08 24.03

[0146] The validation was performed on two multi-view data. The first data is the ECG dataset, which has 162 original electrocardiogram records, namely 96 arrhythmias (ARR), 30 congestive heart failure (CHF), and 36 normal sinus rhythm (NSR). Each data record lasts 512s and the sampling rate is 128Hz. For the ARR category, the first 20 seconds are used as a segmentation. For the CHF and NSR categories, the first 60s are used and evenly divided 3 times. Fourier coefficients (1281 dimensions) and time domain features (2560 dimensions) are used as 2 feature views, resulting in multi-view data containing 294 samples.

[0147] The second data is the ALOI100 dataset, which contains 11,025 images of 100 objects. RGB color histogram features (64 dimensions), HSV color histogram features (64 dimensions), color similarity features (77 dimensions) and Haralick features (13 dimensions) are used as 4 feature views.

[0148] In ECG and ALOI100 data, there are random missing samples in each view. The annotation ratio of both data is 10%. For ECG data, the missing ratio of each view is 35%, and r0 is set to 90; for ALOI100 data, the missing ratio of each view is 30%, and r0 is set to 80. SSC_IHIC, MLAN, CONMF, LACK, GPSNMF and MVAR are the existing methods.

[0149] The above methods are all semi-supervised classification models for complete multi-view data. For the above methods, the missing samples are first filled with the feature average of each view, and then the filled data is input into the model. All methods are tested for computational complexity on the same computer configuration: Ubuntu 18.04.6LTS system, MATLAB R2016b, Intel(R)Xeon(R)CPU and 503GB RAM. The data is randomly missing and fixed, and the random data is repeatedly labeled 10 times, and the mean of the 10 classification results is recorded. It can be concluded that the classification performance of the algorithm of this application is significantly better than the existing algorithms, and the computational efficiency is also the highest.

[0150] In summary, in Example 1 of the present application, an anchor graph matrix is ​​constructed based on all samples of each incomplete view data, and the rows and columns of the anchor graph matrix are expanded to achieve alignment of anchor graphs of different views at the sample level, and an initial fused anchor graph is obtained by averaging; the anchor graph is compressed, label prediction and anchor graph compression are iteratively performed, and a discriminative anchor graph is learned; the prediction of unlabeled samples is split into two steps, and the prediction of unlabeled samples is split into two steps. First, the soft label matrix F of the anchor is calculated. m , and then based on the soft label matrix Fm Calculate the soft label matrix F of the unlabeled samples u And get the predicted label y u_pre , thus avoiding the calculation of the inverse matrix of the dense matrix during the label prediction process and achieving fast label prediction.

[0151] Example 2

[0152] See also Figure 2 , which is a structural diagram of an incomplete multi-view semi-supervised classification system based on a compressed anchor graph according to Example 2 of the present application; the specific contents include:

[0153] Calculation module: Calculate the average of the extended anchor map matrix based on the original multi-view data to obtain the initial fused anchor map;

[0154] Compression module: compressing the initial fusion anchor map based on the class label to obtain a compressed anchor map;

[0155] Judgment module: predicting the class label and confidence score of the unlabeled sample based on the compressed anchor point graph, and combining the class label of the unlabeled sample with a high confidence score with the class label of the known label sample as the known class label;

[0156] Iteration module: learns to discriminate anchor graphs through iterative label prediction and anchor graph compression;

[0157] Classification module: Calculates the classification accuracy on unlabeled samples.

[0158] In summary, in Example 2 of the present application, a calculation module is used to construct respective anchor graph matrices based on incomplete view data, and the anchor graph matrices of each view are expanded to the scale of complete data, and the alignment and fusion of multiple views are achieved by averaging the anchor graphs of multiple views; a compression module is further used to compress the initial fused anchor graph; in addition, a fast label prediction strategy based on a compressed anchor graph is proposed through a judgment module; and an iteration module is used to iterate label prediction and anchor graph compression, so that the learned anchor graph has better discriminability, which can effectively improve the classification accuracy of the model.

[0159] Example 3

[0160] See also Figure 3 , is a schematic diagram of the device structure of Embodiment 3 of the present application. The device 50 includes a processor 51 and a memory 52 coupled to the processor 51 .

[0161] The memory 52 stores program instructions for implementing the above-mentioned incomplete multi-view semi-supervised classification method based on compressed anchor graph.

[0162] The processor 51 is configured to execute program instructions stored in the memory 52 to implement incomplete multi-view semi-supervised classification based on compressed anchor graphs.

[0163] The processor 51 may also be referred to as a CPU (Central Processing Unit).

[0164] The processor 51 may be an integrated circuit chip with signal processing capabilities. The processor 51 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0165] Example 4

[0166] See also Figure 4 , which is a schematic diagram of the structure of the storage medium of Example 4 of the present application. The storage medium of the embodiment of the present application stores a program file 61 that can implement all the above methods, wherein the program file 61 can be stored in the above storage medium in the form of a software product, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or a computer, a server, a mobile phone, a tablet, and other devices.

[0167] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.

[0168] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.

[0169] Although the embodiments of the present application have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the appended claims and their equivalents.

[0170] Of course, the present invention may have many other implementations. Based on this implementation, other implementations obtained by ordinary technicians in this field without any creative work are all within the scope of protection of the present invention.

Claims

1. Incomplete multi-view semi-supervised classification method and system based on compressed anchor graph, It is characterized in that include: The expanded anchor map matrix of the original multi-view data is averaged to obtain the initial fused anchor map; Compressing the initial fused anchor point graph based on the class label to obtain a compressed anchor point graph; Predicting the class labels and confidence scores of the unlabeled samples based on the compressed anchor graph, and combining the class labels of the unlabeled samples with high confidence scores with the class labels of the known label samples as known class labels; Learn the discriminative anchor map by iteratively performing label prediction and anchor map compression; Calculate the classification accuracy on unlabeled samples.

2. The incomplete multi-view semi-supervised classification method based on compressed anchor graph according to claim 1, It is characterized in that The step of calculating the average of the extended anchor map matrix based on the original multi-view data to obtain the initial fused anchor map specifically includes the following steps: Calculate the rows and columns of the expanded anchor map matrix: Among them G v P v G vT To transform the matrix P v The number of rows and columns of is expanded from the number of samples in each view to the size n of the complete data; Based on the average of the extended anchor map matrix, the initial fused anchor map is obtained: Where V is the number of views.

3. The incomplete multi-view semi-supervised classification method based on compressed anchor graph according to claim 1, It is characterized in that The step of compressing the initial fused anchor graph based on the class label to obtain a compressed anchor graph specifically includes the following steps: Based on the class label, the initial fusion anchor map P ave To compress, the shape of the anchor map matrix changes from n v ×n v becomes n v ×C, that is: P=S k (P ave ) Where P = [P 1 ,...,P c ,...,P C ], C is the number of categories, c = 1, 2, ..., C; Specific location, Where N(c) is the sample index set of the cth class; |N(c)| is the set size; (P ave ) ·j is the matrix P ave The jth column of The first iteration is recorded as t = 0, and the initial fusion anchor graph P is ave Compress to obtain the compressed anchor point map P 0 ; If t≠0, update r=min(r0+2×(t-1),100) and update the initial fusion anchor graph P based on the class labels generated by the model. ave Compress to obtain the compressed anchor point map P t .

4. The incomplete multi-view semi-supervised classification method based on compressed anchor graph according to claim 1, It is characterized in that In the step of predicting the class label and confidence score of the unlabeled sample based on the compressed anchor graph, combining the class label of the unlabeled sample with a high confidence score with the class label of the known label sample as the known class label, the following steps are specifically included: The compressed anchor graph is divided into sub-anchor graphs P corresponding to labeled samples lm And the sub-anchor graph P corresponding to the unlabeled sample um ; Based on the sub-anchor graph P lm and label matrix Q l Calculate the anchor soft label matrix F m ; Based on the sub-anchor graph P um and anchor soft label matrix F m Calculate the soft label matrix F for unlabeled data u , and get the corresponding confidence score s i With label y ur ; The confidence score s i Sorting is performed, and the class labels of unlabeled samples with high confidence scores are combined with the class labels of known label samples as known class labels.

5. The incomplete multi-view semi-supervised classification method based on compressed anchor graph according to claim 4, It is characterized in that Based on the sub-anchor graph P lm and label matrix Q l Calculate the anchor soft label matrix F m The steps specifically include the following steps: Assume that the first l samples of the complete data are known label samples, and the last u samples are unknown label samples, n = l + u; the compressed anchor point graph is recorded as P = [P lm ;P um ]; the bipartite graph with labeled data and anchor points is: Minimize the following issues: Among them, F s1 =[F l ; F m ]; L s1 is the Laplace matrix; The solution to the minimization problem is: Where Q l is the label matrix of known data, D lm_r and D lm_c is a diagonal matrix.

6. The incomplete multi-view semi-supervised classification method based on compressed anchor graph according to claim 4, It is characterized in that Based on the sub-anchor graph P um and anchor soft label matrix F m Calculate the soft label matrix F for unlabeled data u , and get the corresponding confidence score s i With label y ur The steps specifically include the following steps: The bipartite graph of unlabeled data and anchor points is: Minimize the following issues: Among them, F s2 =[F u ; F m ]; L s2 is the Laplace matrix; The solution to the minimization problem is: Where D um_r and D um_c is a diagonal matrix; The calculation formula for the label of unlabeled data is: in is the confidence score of the unlabeled data; and F = [F l ; F u ],y=[y l ;y u_pre ].

7. The incomplete multi-view semi-supervised classification method based on compressed anchor graph according to claim 4, It is characterized in that In the confidence score s i The step of sorting and combining the class labels of the unlabeled samples with high confidence scores with the class labels of the known label samples as the known class labels specifically includes the following steps: Sort the confidence scores of the unlabeled data and retain the data labels y with the top r% confidence scores ur , combined with the label of the known labeled sample as the known class label, update t=t+1 until r=100.

8. A system for the incomplete multi-view semi-supervised classification method based on compressed anchor graph according to claim 1, It is characterized in that include: Calculation module: Calculate the average of the extended anchor map matrix based on the original multi-view data to obtain the initial fused anchor map; Compression module: compressing the initial fusion anchor map based on the class label to obtain a compressed anchor map; Judgment module: predicting the class label and confidence score of the unlabeled sample based on the compressed anchor point graph, and combining the class label of the unlabeled sample with a high confidence score with the class label of the known label sample as the known class label; Iteration module: learns to discriminate anchor graphs through iterative label prediction and anchor graph compression; Classification module: Calculates the classification accuracy on unlabeled samples.

9. A device, It is characterized in that The device includes a processor and a memory coupled to the processor, wherein the memory stores program instructions for implementing the incomplete multi-view semi-supervised classification method based on a compressed anchor graph as described in any one of claims 1-7; and the processor is used to execute the program instructions stored in the memory to implement the incomplete multi-view semi-supervised classification based on a compressed anchor graph.

10. A storage medium, It is characterized in that The method stores program instructions executable by a processor, wherein the program instructions are used to execute the incomplete multi-view semi-supervised classification method based on a compressed anchor graph according to any one of claims 1 to 7.