A multi-view multi-manifold classifier with local and global structure preservation

By combining multi-view and multi-manifold learning and utilizing local and global structure to maintain constraints, the problem of unutilized multi-manifold information in multi-view learning is solved, achieving higher classification accuracy and stability, and forming a general multi-view learning framework.

CN116310467BActive Publication Date: 2025-11-11EAST CHINA UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211106294.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-11
Publication Date
2025-11-11
Estimated Expiration
2042-09-11

AI Technical Summary

Technical Problem

Existing kernel-based multi-view learning methods fail to fully consider multi-manifold information, resulting in reduced model comprehensiveness and classification accuracy. They also ignore the global and local structure of training samples, affecting classification performance.

Method used

By combining multi-view learning and multi-manifold learning methods, we preserve constraints through local and global structures, extract features using multi-manifold information, maximize the inter-class Laplacian scattering matrix and minimize the intra-class Laplacian matrix, map the data to a low-dimensional feature space to preserve geometric structure, and use IFSL terms to control the complementarity of different perspectives.

Benefits of technology

It achieves better multi-view data classification performance, improves classification accuracy and stability by preserving the global and local geometric structure of training data, and forms a general multi-view learning framework.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310467B_ABST
    Figure CN116310467B_ABST
Patent Text Reader

Abstract

This invention discloses a multi-view, multi-manifold classifier that preserves both local and global structure. In the feature space obtained through multiple empirical kernel mappings, this invention uses nonlinear feature extraction to preserve the inherent low-dimensional embeddings of the original data from multiple views. Furthermore, this invention utilizes inter-class graphs representing multi-manifold information and intra-class graphs representing sub-manifold information to constrain the training of the discriminative hyperplane. This invention overcomes the deficiency of existing multi-view classifiers that neglect the manifold structure of the original samples, improving the classification performance of multi-view data by utilizing the local and global geometric information surrounding each view's data point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a multi-view, multi-manifold classifier that preserves both local and global structure, belonging to the field of image classification technology. Background Technology

[0002] Multi-view image classification technology has attracted significant attention in the fields of pattern recognition and machine learning. Kernel-based multi-view learning methods, in particular, process multi-view attributes by extracting information from different attributes using different kernel functions. This approach typically integrates information from multiple classification or regression models using different kernel functions. The processing of multi-view data is very common in real-world applications. For example, in images and videos, color and texture information are two distinct features and can be considered as dual-view data. In web page classification, there are usually two perspectives to describe a given web page: the text content of the web page itself and the anchor text of any web pages linking to it. This is because single-view data often cannot comprehensively describe the information of all samples, thus requiring data collection from different measurement methods.

[0003] Despite significant progress in kernel-based multiview image classification, two issues remain to be addressed. First, kernel-based multiview learning feature extraction methods do not adequately consider multi-manifold information. Ignoring this information reduces the model's comprehensiveness and classification accuracy. The manifold hypothesis states that each class has its own individual manifold structure, and each object space is typically a submanifold, with multiple distinct object spaces forming multiple manifolds. Second, kernel-based multiview learning methods do not sufficiently prioritize preserving the global and local structure of training samples. In fact, maintaining the global or local structure among training samples is crucial for kernel-based multiview classification. In machine learning, intuitive geometric methods offer excellent interpretability and generalization.

[0004] This invention addresses the problem of multi-view data classification from an algorithmic perspective, utilizing multi-manifold learning. It combines multi-view and multi-manifold learning methods, applying constraints to feature extraction to preserve local and global structure. The system fully considers multi-manifold and sub-manifold information to find the optimal low-dimensional embedding for the dataset. Furthermore, this invention emphasizes preserving the global and local geometric structure of the training data in the feature space to achieve better classification performance. Local structure is obtained through analysis of sub-manifold information, and global structure is obtained from multi-manifold information by merging all sub-manifold information. This invention first determines the local geometric information of each data point and maps the data to its inherent low-dimensional feature space by maintaining this geometric information. Then, the projection matrix is ​​constrained by simultaneously maximizing the inter-class Laplacian scattering matrix and minimizing the intra-class Laplacian matrix. Ultimately, this ensures that in the low-dimensional subspace, points from the same class remain close together, while samples from different classes are as far apart as possible. Summary of the Invention

[0005] Technical Problem: This invention provides a multi-view, multi-manifold classification algorithm that preserves both local and global structure, utilizing multi-manifold information from multi-kernel multi-view learning to achieve better classification performance. By fully exploring the geometric structure of samples in the feature space, this invention improves classification performance from a geometric perspective.

[0006] Technical Solution: First, the original sample data is divided into training and test sets. Second, different types of kernel functions are used to map training samples from different perspectives to their respective feature spaces. Next, the local structure of the submanifolds of each class of samples in the feature space is comprehensively considered to obtain the global structure of the dataset. Simultaneously, two graphs are used to describe intra-class compactness and inter-class separability, respectively. Then, during classification after feature extraction, the complementarity between different perspectives is controlled by the IFSL term to improve the accuracy of the model in classifying multi-view data. In the testing step, the mapped test samples are substituted into the discriminant function corresponding to the model for identification.

[0007] The technical solution adopted by this invention to solve its technical problem can be further refined. In the second step of the training phase, there are multiple candidate kernel mapping types for each viewpoint, such as linear kernels, RBF kernels, polynomial kernels, etc. In the third step of the training phase, the algorithm originally applied to Euclidean space is transformed using the structure and properties of the manifold, so that it acts on the manifold. In practice, the original data may be in a nonlinear distribution. This invention first determines the local geometric information of each data point, and then maps the data by preserving the local geometric information in the low-dimensional feature space. High-dimensional data is regarded as a set of geometrically related points on a smooth low-dimensional manifold. This invention uses inter-class graphs to carry multi-manifold information and intra-class graphs to carry sub-manifold information.

[0008] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0009] Existing kernel-based multi-view learning methods do not adequately consider multi-manifold information, leading to reduced model comprehensiveness and classification accuracy. This invention, for the first time, utilizes multi-manifold information to achieve better classification performance for multi-view data. Unlike existing multi-view methods, this method does not perform classification directly in the feature space, but rather comprehensively considers multi-manifold and sub-manifold information to find the optimal low-dimensional embedding for the dataset. Furthermore, it preserves the global and local geometric structure of the training data in the low-dimensional subspace, thereby generating more reasonable classification boundaries. This invention not only combines multi-view learning and multi-manifold learning methods but also forms a relatively general multi-view learning framework. Attached Figure Description

[0010] Figure 1This is a flowchart of the training process for the multi-view, multi-manifold classifier of this invention. Detailed Implementation

[0011] To more clearly describe the content of this invention, further explanation is provided below with reference to examples and accompanying drawings. The embodiments described below are not intended to limit the scope of this invention. The multi-view multi-manifold classifier algorithm of this invention, which preserves both local and global structure, includes the following steps:

[0012] Step 1: Input the total sample set Where N is the number of training samples, ψ i ∈{+1,-1} is a sample x i The class labels are given, where v = [1, ..., V] represents the v-th viewpoint, and there are a total of V viewpoints. After splitting the sample set, the training set and test set are output.

[0013] Step 2: Input the training set, and through empirical kernel mapping, map the samples from different perspectives to V feature spaces respectively:

[0014]

[0015] Where, Φ e (x) is the empirical kernel mapping, defined as:

[0016]

[0017] Where Λ is a diagonal matrix composed of the r positive eigenvalues ​​of the kernel matrix K, and P is composed of the corresponding orthogonal eigenvectors, that is: Kernel function ker(x) i ,x j There are several candidate types for kernel functions, such as linear kernel functions:

[0018] k(x,x i )=x·x i (3)

[0019] Polynomial kernel function:

[0020] k(x,x i )=((x·x i )+1) d (4)

[0021] Gaussian kernel function:

[0022]

[0023] Where σ is the bandwidth parameter. In the experiment, the kernel function with the best classification performance was selected for mapping. After empirical kernel mapping, a new training set is output.

[0024] Step 3: Input a new training set, expand the new training set obtained by mapping the empirical kernel, and obtain the sample set Y in the feature space:

[0025]

[0026] Among them, Y v Let v represent the sample set of the v-th viewpoint in the feature space.

[0027] Step 4: Input the sample set Y in the feature space, and calculate the in-class Laplacian matrix L. w Inter-class Laplace matrix L b The weighted class center matrix M. The compactness of the matrix within a class can be determined by the in-class Laplace matrix L. w This matrix is ​​associated with the intra-class graph. The separability of samples between classes can be characterized by the inter-class Laplacian matrix L, which is associated with the inter-class graph. b It is represented by the weighted class center matrix M. The output L is obtained through the following small steps of calculation. w L b M.

[0028] Step 4.1: The intra-class graph represents the submanifold information. Calculate the intra-class sample y using the following formula. i and y j Similarity weights between:

[0029]

[0030] in, It is a weighting factor.

[0031] Step 4.2: Obtain the similarity matrix S representing the dataset from the above formula:

[0032]

[0033] Where c represents the total number of categories.

[0034] Step 4.3: Next, calculate the block diagonal matrix D corresponding to matrix S. w :

[0035]

[0036] The diagonal element D wii =∑ j S ij It is the column sum of S, representing the importance of each sample in its class.

[0037] Step 4.4: Based on the similarity matrix S and its diagonal matrix D calculated above... w Calculate the intraclass Laplacian matrix L w :

[0038] L w =D w -S (10)

[0039] Step 4.5: Calculate the weighted center of each class based on matrix Dw:

[0040]

[0041] in, This represents the weighting center of the k-th class.

[0042] Step 4.6: Obtain the weighted class center matrix of the dataset from the weighted centers of each class:

[0043]

[0044] Where k and k' represent two different classes.

[0045] Step 4.7: Calculate the distance B between the two class centers. ij Regulation center and The distance between them affects:

[0046]

[0047] Where i and j represent different class labels, i,j∈{1,2,...,c}, i≠j. It is a weighting factor.

[0048] Step 4.8: Obtain the class center spacing matrix B of the dataset from the above formula:

[0049]

[0050] Step 4.9: Next, calculate the block diagonal matrix D corresponding to matrix B. b :

[0051]

[0052] Its diagonal element D bii =∑ j B ij It is the column sum of B.

[0053] Step 4.10: Based on the above class center spacing matrix B and its diagonal matrix D... b The inter-class Laplace matrix L is obtained. b :

[0054] L b =D b -B (16)

[0055] Step 5: Utilizing the multi-kernel mapping, multi-manifold feature extraction, local and global structure-preserving constraints, and the MHKS base classifier described above, the objective function R(F) of this invention is obtained. It comprises five parts, represented as follows:

[0056]

[0057] Part 1 R emp (f v ) is the experience-based risk item, which is defined as:

[0058] R emp (f v )=(Y v w v -1 N×1 -b v ) T ×(Y v w v -1 N×1 -b v (18)

[0059] Among them, f v The weighted projection vector w represents the sub-classifier at the v-th viewpoint. v It is a classification hyperplane, b v It is a margin vector and each element has a non-negative value, 1 N×1 It is an N-dimensional column vector in which all components are set to 1.

[0060] Part Two R reg (f v ) is a regularization term that penalizes roughness, and it is defined as:

[0061]

[0062] The classifier f from the vth perspective is penalized. v The roughness or smoothness. The first and second parts constitute the MHKS base classifier.

[0063] Part 3 R w (f v ) is a regularization term that minimizes the Laplace divergence within a class, and is defined as:

[0064]

[0065] Part 4 R b (f v ) is the regularization term that maximizes the inter-class Laplace divergence, and it is defined as:

[0066]

[0067] Part 5 RIFSL (F) plays a role in the consistency penalty of outputs from multiple perspectives. This term is defined as:

[0068]

[0069] Where F is the overall classifier.

[0070] Combining the above five parts, the final objective function is:

[0071]

[0072] Where α, β, γ, and λ are the regularization parameters for the corresponding terms.

[0073] Step 6: Based on the objective function of this invention, solve this optimization problem using the heuristic gradient descent method, and output the w to be solved. v and b v The solution steps are as follows:

[0074] Step 6.1: R(F) versus w v Taking the partial derivative and setting it to 0, we get:

[0075]

[0076] Where I represents a matrix consisting entirely of 1s. In the k-th iteration, we obtain:

[0077]

[0078] Step 6.2: R(F) versus b v Taking the partial derivative, we get:

[0079]

[0080] Where e v Let b be the error vector at the v-th viewpoint. v The components determine the hyperplane w from the sample to the classification hyperplane. v The distance. To prevent it from converging to zero, set b to... v Starting the iteration from greater than or equal to zero, at the k-th iteration, we obtain:

[0081]

[0082] Where, ρ v >0 represents the learning rate.

[0083] Step 6.3: Iterate until the condition ||R(F) is met. k+1 -R(F) k When ||2≤ε, the iteration terminates and the current w is output. v and b v, where ε is the iteration threshold preset by the user.

[0084] Step 7: Input test sample x i The test samples are then substituted into the discriminant function. Using the training results obtained from the above steps, the final output is the category of each test sample. The discriminant function for this example is:

[0085]

[0086] According to the appendix above Figure 1 Specific embodiments of the present invention have been described. However, those skilled in the art will understand that several improvements and equivalent substitutions can be made without departing from the spirit and principles of the present invention. The improved and equivalent substitution techniques and solutions described in the claims of the present invention all fall within the protection scope of the present invention.

[0087] Experimental Design

[0088] Experimental Dataset Selection: This invention uses five real-world datasets for multi-view image classification for experiments. Their details are shown in the table below.

[0089] Table 1: Dataset Description Table

[0090]

[0091]

[0092] The parameter selection for all algorithms employs a 5-round crossover method. First, the dataset is divided into two parts: one for training and one for testing. Then, the training set is further divided into five parts (1, 2, 3, 4, 5). First, parts 1-4 are trained, and part 5 is validated. Then, parts 1-5 are trained, and part 4 is validated, for a total of five rounds. The best parameters are then selected and tested on the test set.

[0093] Comparison Model: The system proposed in this invention is named MML-LGSP. We compare the classification performance of MvCCDA, AdaRaker, CGMKL, MCCD, MEKL-LPC, MultiK-MHKS, and our system MML-LGSP.

[0094] Performance metrics: This experiment uses two evaluation metrics, one of which is accuracy (ACC). ACC represents the percentage of the entire dataset that is correctly judged. The definition of ACC is as follows:

[0095]

[0096] In this context, TP, TN, FP, and FN represent True Positives, True Negatives, False Positives, and False Negatives, respectively.

[0097] The second evaluation metric is the Macro F1-score, defined as:

[0098]

[0099] in,

[0100] Experimental results:

[0101] Table 2: Classification Accuracy ACC [mean±std(%)]

[0102]

[0103] Table 3: Macro F1-score[mean±std(%)]

[0104]

[0105] Tables 2 and 3 show the prediction results and their mean squared errors under the two metrics of ACC and Macro F1-score, respectively. Each row corresponds to a dataset, and each column corresponds to an algorithm. The best results for each dataset are marked in bold.

[0106] As can be seen from the classification results and scores of various algorithms, the present invention achieves relatively good classification results on multi-view image datasets, and its stability on all datasets results in a high average classification performance.

Claims

1. A multi-view, multi-manifold classifier with local and global structure preservation, characterized in that, The training method for this classifier includes the following steps: 1) Divide the original sample data into two parts: a training set and a test set; 2) Through empirical kernel mapping, samples from different perspectives are mapped to their respective feature spaces, and after expansion, the sample set in the kernel feature space is obtained; 3) Sample features are extracted using a multi-manifold information feature extraction method. By constraining the training of the classification hyperplane, the following constraints are applied to maximize inter-class separability and minimize intra-class divergence: It is a regularization term used to minimize the intra-class divergence, where w v It is the classification hyperplane of the v-th view that needs to be trained, Y v It is the dataset on the v-th viewpoint after empirical kernel mapping, L w It is based on dataset Y v The calculated intraclass Laplacian matrix; It is a regularization term used to maximize the separability between classes, where M is the weighted class center matrix, and L is the regularization term used to maximize the separability between classes. b It is the inter-class Laplace matrix; Maximizing inter-class separability while minimizing intra-class divergence is achieved by minimizing the following formula: Where β and γ are the regularization parameters for the corresponding terms; 4) Introducing multi-manifold learning into a kernel-based multi-view framework yields the training results for the training samples. In the training step, the objective function is established using the MHKS base classifier, combined with empirical kernel mapping. The objective function of this classifier consists of five parts: Part 1 R emp (f v ) is the experience risk item, which is defined as: R emp (f v )=(Y v w v -1 N×1 -b v ) T ×(Y v w v -1 N×1 -b v ) Where f v Y represents the subclassifier at the v-th viewpoint. v It is the training set mapped from the empirical kernel to the feature space, and the weighted projection vector w v It is a classification hyperplane, b v It is a margin vector and each element has a non-negative value, 1 N×1 It is an N-dimensional column vector with all components set to 1; Part Two R reg (f v ) is a regularization term that penalizes roughness, and this regularization term is defined as: The regularization term penalizes the classifier f on the v-th view. v The roughness or smoothness of the material, the first part and the second part together form the MHKS base classifier; Part 3 R w (f v ) is the regularization term that minimizes the intra-class Laplace divergence, defined as: Part 4 R b (f v ) is the regularization term that maximizes the between-class Laplace divergence, defined as: Part 5 R IFSL (F) plays a role in the consistency penalty of outputs from multiple perspectives, R IFSL (F) is defined as: Where F is the overall classifier; Combining the above five parts, the final objective function is: Where α, β, γ, and λ are the parameters of the corresponding terms, and V represents the total number of viewpoints; 5) In the testing step, the test samples are substituted into the discriminant function corresponding to the multi-view multi-manifold classifier for identification.

2. A multi-view multi-manifold classifier with local and global structure preservation as described in claim 1, which acquires submanifold information and calculates the intra-class Laplacian matrix L. w The specific process is as follows: First, input the sample set, then use the following formula to calculate and output the within-class samples y. i and y j Similarity weights between: in It is a weighting factor, ψ i and ψ j They represent y respectively i and y j The class label, N is the number of samples; Then, the similarity matrix S representing the dataset is calculated and output using the following formula: Where c represents the number of classes; Next, calculate and output the block diagonal matrix D corresponding to matrix S using the following formula. w : The diagonal element D wii =∑ j S ij It is the column sum of S, representing the importance of each sample in its class; Finally, input the similarity matrix S calculated above and its diagonal matrix D. w Calculate and output the in-class Laplacian matrix L w : L w =D w -S。 3. A multi-view, multi-manifold classifier with local and global structure preservation as described in claim 1, acquiring multi-manifold information and calculating the inter-class Laplacian matrix L. b The specific process for the weighted class center matrix M is as follows: First, input matrix D w Calculate and output the weighted center for each class: in, This represents the weighting center of the k-th class; Secondly, input the weighting center for each class. Output the weighted class center matrix of the dataset: Where k and k' represent two different classes; Furthermore, given input M, calculate and output the distance B between the two class centers. ij Used to regulate the two class centers and The distance between them affects: Where i and j represent different class labels, i,j∈{1,2,...,c}, i≠j. It is a weighting factor; Next, enter B. ij Output the class center spacing matrix B of the dataset: Then, calculate and output the block diagonal matrix D corresponding to matrix B. b : Its diagonal element D bii =∑ j B ij It is the column sum of B; Finally, input the class center spacing matrix B and its diagonal matrix D obtained above. b Calculate and output the inter-class Laplacian matrix L b : L b =D b -B。

Citation Information

Patent Citations

  • Local spline embedding-based orthogonal semi-monitoring subspace image classification method

    CN101916376A

  • Method for establishing Alzheimer's disease layered multi-manifold analysis model

    CN106202916A