Improved multi-view machine learning classification method
By adopting hierarchical clustering and alternating optimization methods in multi-perspective machine learning, the information between different perspectives is effectively integrated, and the computational complexity and information integration problems of traditional SVM in the processing of large multi-perspective data sets is solved, and the accuracy and robustness of the classification model are improved.
Patent Information
- Application Number
- CN202510194740.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional support vector machines (SVMs) have high computational complexity when processing large multi-view data sets, making it difficult to effectively integrate information between different perspectives, resulting in limited classification performance.
The improved multi-view machine learning classification method is adopted to enhance information exchange between perspectives through hierarchical clustering, an initial classification model is constructed, and the optimal weight and hyperplane are alternately optimized to achieve effective integration of information between perspectives.
It significantly improves the accuracy and robustness of the classification model, avoids the computational complexity of traditional SVM in large data set processing, and is suitable for a variety of application scenarios such as image recognition, medical diagnosis and sentiment analysis.
Smart Images

Figure CN120125889A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly to an improved multi-view machine learning classification method. Background Art
[0002] Support Vector Machine (SVM) is a powerful machine learning algorithm mainly used for classification tasks. In the field of object recognition, image classification often adopts the SVM model. For example, in face recognition applications, SVM can effectively distinguish the facial features of different people, especially when dealing with images with a large amount of complex backgrounds or different lighting conditions; in the field of natural language processing, SVM is often applied to tasks such as spam filtering and sentiment analysis. Its powerful feature extraction ability enables SVM to remain efficient and accurate when facing high-dimensional text data; in the field of financial technology, support vector machines are widely used in credit scoring and market trend prediction. Due to its good classification performance, SVM can effectively identify potential high-risk customers; in bioinformatics: in tasks such as gene classification and protein structure prediction, SVM is used as a powerful classification tool, and its processing ability for high-dimensional characteristics is particularly suitable for dealing with complex biological data. These applications can help scientists better understand disease mechanisms and develop new treatment strategies. In order to improve the efficiency of Support Vector Machine (SVM) in some cases that may be complex and computationally intensive, especially when dealing with large datasets, the improved LSSVM classification algorithm uses a least squares cost function to minimize the squared error between the predicted value and the actual output, which results in a simpler system of linear equations that is generally faster to solve than the quadratic programming problem in traditional SVM.
[0003] With the development of artificial intelligence, the technical means of information collection have been continuously enhanced, and data shows a development trend of being massive, diverse, and high-dimensional. People are not limited to understanding or describing a certain sample from a single data source, such as news written in different languages, image descriptors obtained by different feature extractors, etc. For the same object, the feature data obtained from different channels or different levels is called multi-view data. For example, fingerprints, voices, irises, etc. constitute multiple perspectives for identifying a person's identity. The features from these different perspectives describe the same object, but the feature distributions of different perspectives are in different feature spaces. Multi-view data presents characteristics such as polymorphism, multi-source, multi-description, and high-dimensional heterogeneity. There are both internal connections and differences between different perspectives, and a new learning method is needed to process and process these data or features, so as to make full and reasonable use of the information in multi-view data, that is, multi-view learning. Multi-view learning (MVL) is a machine learning method that uses multiple feature sets or perspectives to enhance the learning process. It can play an important role in many fields and applications, such as enhancing classification accuracy: by integrating data from different perspectives (such as images, texts, audios, etc.), the performance of the classifier on multi-modal data sets can be improved; dealing with occlusion situations: in complex scenarios, the target may be partially occluded, and multi-view learning can use the information supplement between different perspectives to better identify the occluded object; group behavior analysis: in social network analysis or group behavior research, different data perspectives (such as social media dynamics, user interactions, etc.) can be used to understand and predict group behavior; medical diagnosis: by combining medical data from multiple perspectives (such as images, genes, and clinical information), more accurate disease diagnosis and personalized treatment can be carried out. Sentiment analysis: In sentiment analysis, multiple perspectives such as text, speech, and vision can be considered to enhance the understanding of the emotional state.
[0004] Due to the diversity of data sources in multi-view learning models, in the innovative field of multi-view learning, the effective integration of two key principles has received high attention, namely the complementarity principle and the consistency principle. The complementarity principle holds that different perspectives should provide unique but complementary insights. The diversity of viewpoints enhances the overall understanding and improves the model performance because each perspective contributes unique information, reducing the risk of overfitting to a single source. The consistency principle emphasizes that the information obtained from different perspectives should be aligned and reinforced with each other. Ensuring the consistency between viewpoints helps to maintain the consistency of the learning process, enabling the model to tend to a more accurate representation of the underlying pattern. The model reconstructed by these principles provides a robust framework for multi-view learning, making the developed model not only comprehensive but also reliable in different environments. Summary of the Invention
[0005] To overcome the deficiencies of the prior art, the object of the present invention is to provide an improved multi-view machine learning classification method, which avoids the computational complexity of traditional support vector machines (SVMs) when dealing with large datasets, and enhances the information exchange between views by using hierarchical clustering, thereby promoting the overall performance improvement of the model in multi-view learning tasks.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] An improved multi-view machine learning classification method, comprising:
[0008] Obtain a multi-view image dataset;
[0009] Extract features from the multi-view image dataset to obtain multi-view feature data; the multi-view feature data includes first-view features and second-view features;
[0010] Construct an initial classification model;
[0011] Input the multi-view feature data into the initial classification model to obtain a trained image classification model;
[0012] Input the image to be tested into the image classification model to obtain a classification result.
[0013] Preferably, the expression of the initial classification model is:
[0014]
[0015] where λ, c, and d are all predetermined parameters, whose values are greater than 0, V is the number of views of the multi-view image dataset, X is the multi-view feature data, y is the label of the multi-view feature data, ξ is the loss, θ [v] is the optimal weight of the v-th view, and w [v] is the hyperplane for multi-view classification.
[0016] Preferably, the method for determining the optimal weight and the hyperplane includes:
[0017] S1: Fix the weight θ [v] , and update w [v] , specifically including:
[0018] Define a matrix H, the expression of which is: H = blockdiag{θ 1 + dΣ [1] ,..., θ V + dΣ [V]}; where in the formula, l is the number of samples, and T is the transpose matrix;
[0019] Construct the Lagrangian function \(L\) according to the matrix \(H\). 2 The Lagrangian function \(L\) 2 has the following expression: where is a non - negative Lagrange multiplier vector, \(w\) is a vector composed of \(w\) from multiple perspectives [v] \(\xi\) is a loss vector composed of \(\xi\) from multiple perspectives [v] \(Y\) is a label diagonal matrix composed of \(y\), \(X\) is the training data set composed of all perspectives, and \(e\) is a column vector of all 1s;
[0020] Take the derivative of the Lagrangian function \(L\) 2 with respect to the variables to obtain the KKT conditions; the expression of the KKT conditions is:
[0021] Solve the KKT conditions to obtain a non - negative Lagrange multiplier vector; the expression of the non - negative Lagrange multiplier vector is:
[0022] Solve according to the non - negative Lagrange multiplier vector to obtain the hyperplane \(w\) [v] ;
[0023] S2: Fix the hyperplane \(w\) [v] , and update the weight through the quadratic programming QPP with respect to the weight \(\theta\) [v] ;
[0024] Alternate steps S1 and S2 until the algorithm converges, so as to obtain the optimal perspective weights and the optimal hyperplane.
[0025] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0026] The present invention provides an improved multi-view machine learning classification method, including: obtaining a multi-view image data set; extracting features from the multi-view image data set to obtain multi-view feature data; the multi-view feature data includes first-view features and second-view features; constructing an initial classification model; inputting the multi-view feature data into the initial classification model to obtain a trained image classification model; inputting the image to be tested into the image classification model to obtain a classification result. By following the principles of complementarity and consistency, the present invention effectively integrates feature data from different views, significantly improving the accuracy and robustness of the classification model. This method avoids the computational complexity of traditional support vector machines (SVMs) when dealing with large data sets, and uses hierarchical clustering to enhance information exchange between views, thereby promoting the overall performance improvement of the model in multi-view learning tasks. Finally, the trained image classification model can more accurately identify and classify the images to be tested, with stronger adaptability, and is applicable to a variety of application scenarios, such as image recognition, medical diagnosis, and sentiment analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0028] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention;
[0029] Figure 2 It is an algorithm flowchart provided by the embodiment of the present invention;
[0030] Figure 3 It is a comparison chart of F1-scores of this algorithm and seven other algorithms on the AWA45 data set provided by the embodiment of the present invention;
[0031] Figure 4 It is a comparison chart of the classification accuracy results of the present invention provided by the embodiment of the present invention;
[0032] Figure 5 It is a schematic diagram of parameter analysis of parameters λ, c, and d on three data sets (AWA_1, Bankruptcy, businessvs.sport) provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] The object of the present invention is to provide an improved multi-view machine learning classification method, which avoids the computational complexity of traditional support vector machines (SVMs) when dealing with large datasets, and uses hierarchical clustering to enhance information exchange between views, thereby promoting the overall performance improvement of the model in multi-view learning tasks.
[0035] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] Figure 1 The flowchart of the method provided for the embodiments of the present invention is as Figure 1 shown. The present invention provides an improved multi-view machine learning classification method, including:
[0037] Step 100: Obtain a multi-view image dataset;
[0038] Step 200: Extract features from the multi-view image dataset to obtain multi-view feature data; the multi-view feature data includes first-view features and second-view features;
[0039] Step 300: Construct an initial classification model;
[0040] Step 400: Input the multi-view feature data into the initial classification model to obtain a trained image classification model;
[0041] Step 500: Input the image to be tested into the image classification model to obtain a classification result.
[0042] Specifically, in this embodiment, the AWA45 multi-view dataset is taken as an example. The AWA45 dataset is a commonly used standard dataset for visual recognition and attribute learning, mainly used for developing and evaluating machine learning and computer vision algorithms. It contains 30,475 images covering 50 animals, and each image has 6 pre-extracted features (such as distinguishable features like "horned", "hairy", "flying", etc.). We selected 10 animal categories, namely: beaver, blue whale, skunk, cow, pig, mouse, walrus, weasel, vole, and mole, and paired them into 45 binary classification datasets in a one-versus-all manner.
[0043] Furthermore, taking the AWA45 multi-view dataset as an example, in each animal category, the accelerated robust features (SURF) normalized by 2000-DL1 are regarded as the first view, while the 252-D histogram of oriented gradients (PHOG) features are the second view. Their dimensions are 546 and 252 respectively, and these data are obtained after extracting 90% of the principal components of the original view.
[0044] Furthermore, bring the above multi-view feature data into the following classification model SMvLSSVC-2C of the present invention, and the definition of the model is:
[0045]
[0046] where λ, c, and d are predetermined parameters, V is the number of views, X is the input dataset, y is the label of the input dataset, and ξ is the loss. Exemplarily, the predetermined parameters only need to be greater than 0, and the values are randomly given. The determined numbers will come out during the final optimization.
[0047] The first term of the objective function is the norm regularization term. Minimizing this term means that it follows the principle of minimizing the structural risk of different views. θ [v] is the weight of the v-th view and it is required that the sum of all weights is 1. Optimizing it can complete the mining of complementary information in different views and achieve complementarity. The second term in the objective function avoids the appearance of trivial solutions. The third term of the objective function, its slack variable can reduce the misclassification degree of the classification task in each view and can simplify the computational complexity. The fourth term of the objective function represents the weighted combination of multiple covariance matrices and captures the structural information from different views. Since the current view can represent the structural information from all views, it can utilize the complementary and consistent information from different angles. The potential information from different views can enhance or correct each other.
[0048] As Figure 2 shown, the hyperplane that satisfies multi-view classification with consistency and complementarity is obtained through the following process and the view-optimal weight θ [v] :
[0049] (1) Fix θ to optimize the following problem to update
[0050] For simplicity, define H = blockdiag{θ 1 +dΣ [1] ,...,θ V +dΣ [V]}, where
[0051] In this embodiment, the Lagrangian function of its algorithm can be obtained:
[0052]
[0053] wherein is a non - negative Lagrange multiplier vector. For the Lagrangian function L 2 Taking the derivative with respect to the variables, the following KKT conditions can be obtained:
[0054]
[0055] Solving gives the non - negative Lagrange multiplier vector
[0056] Then the classification hyperplane can be solved
[0057] (2) By fixing the optimal view weight θ can be obtained by solving through quadratic programming QPP to update θ [v] ;
[0058] (3) Combining the obtained classification hyperplane and the optimal view weight θ [v] , the conclusion can be drawn.
[0059] Figure 3 Different symbols in
[0060] Figure 4 represent the F1 - score results of eight different algorithms. The abscissa represents a total of 45 datasets, and the ordinate is the value of the F1 - score. The larger the value of the F1 - score, the better the effect. Through this result, it is found that the proposed method obtains relatively high F1 - score values on most datasets.
[0061] Figure 5 In
[0062] the x, y, and z axes respectively represent the values of three parameters. Different values of the three parameters will result in different accuracies. The dots in the figure represent the accuracies, and the colors of the dots represent the levels of accuracy. Among them, dark blue represents relatively low accuracy, and bright yellow represents relatively high accuracy. This result shows that the accuracy of the proposed method is highly dependent on the selection of parameters. Therefore, it is very necessary to perform parameter optimization.
[0063] The beneficial effects of the present invention are as follows:
[0063] The present invention provides an improved multi-view machine learning classification method that follows the principles of complementarity and consistency. It can avoid the need to solve large quadratic programming problems (OPP) in traditional SVM classification algorithms and can handle multi-view data. A model based on structural information, called SMvLSSVC-2C, is proposed. This classifier effectively minimizes the square of the difference in decision functions between different views, while integrating information from multiple views through coupling terms. It uses hierarchical clustering to enhance information exchange between views, thereby promoting complementarity and consistency, improving computational efficiency, and enhancing the overall performance in multi-view learning tasks.
[0064] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.
[0065] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only for helping to understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An improved multi-view machine learning classification method, characterized in that: include: Obtain a multi-view image dataset; Extracting features from the multi-view image data set to obtain multi-view feature data; The multi-view feature data includes a first view feature and a second view feature; Build an initial classification model; Inputting the multi-view feature data into the initial classification model to obtain a trained image classification model; The image to be tested is input into the image classification model to obtain a classification result.
2. The improved multi-view machine learning classification method according to claim 1, characterized in that: The expression of the initial classification model is: Wherein, λ, c and d are all predetermined parameters, the values of λ, c and d are greater than 0, V is the number of views of the multi-view image data set, X is the multi-view feature data, y is the label of the multi-view feature data, ξ [v] is the loss, θ [v] is the weight of the vth view, w [v] is the hyperplane for multi-view classification, Σ [v] is the covariance matrix composed of clustered data, and V is the total number of viewpoints.
3. The improved multi-view machine learning classification method according to claim 2, characterized in that: The method for determining the optimal weight and the hyperplane includes: S1: Fixed weight θ [v] , update w [v] , specifically including: Define the matrix H, the expression is: H = blockdiag{θ1+dΣ [1] ,...,θ V +dΣ [V] };in In the formula, l is the number of samples, T is the transposed matrix; The Lagrangian function L2 is constructed according to the matrix H; the expression of the Lagrangian function L2 is: in is a non-negative Lagrange multiplier vector, w is the w from multiple perspectives [v] The vector composed of ξ is composed of multiple perspectives. [v] The loss vector is composed of , Y is the label diagonal matrix composed of y, X is the training data set composed of all viewpoints, and e is the all-1 column vector; The Lagrangian function L2 is differentiated with respect to the variable to obtain the KKT condition; the expression of the KKT condition is: The KKT condition is solved to obtain a non-negative Lagrange multiplier vector; the expression of the non-negative Lagrange multiplier vector is: Solving according to the non-negative Lagrange multiplier vector, we get the hyperplane w [v] ; S2: Fixed hyperplane w [v] , by the weight θ [v] The quadratic programming QPP updates the weights; Steps S1 and S2 are performed alternately until the algorithm converges, thereby obtaining the optimal view weight and the optimal hyperplane.