Parallel Support Vector Machine Optimization Method Based on Relative Entropy and Cosine Similarity
Through data division based on relative entropy and redundant hierarchical detection, combined with the MapReduce framework to optimize the parallel support vector machine algorithm, the computing performance bottleneck in the big data environment is solved, and efficient classification and accuracy are improved.
Patent Information
- Application Number
- CN202210285548.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-03-22
AI Technical Summary
The existing parallel support vector machine algorithms face computing performance bottlenecks in the big data environment, especially the problems of large deviations in subset distribution, low parallel efficiency and inaccurate filtering of non-support vectors during data division.
The data division strategy based on relative entropy is adopted to balance the subset distribution of DPRE, combined with the MapReduce framework, and the local SVM model is trained in parallel. Through redundant hierarchical detection and non-support vector filtering strategies, the parallel support vector machine algorithm is optimized.
It significantly improves classification efficiency and accuracy, reduces subset distribution deviation, identifies and stops redundant levels, accurately filters non-support vectors, and improves parallel efficiency and algorithm performance.
Smart Images

Figure CN114638311B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of big data mining, and particularly relates to an optimization method for a parallel support vector machine based on relative entropy and cosine similarity. Background Art
[0002] As a classification algorithm, the support vector machine algorithm takes structural risk minimization as the principle, training error minimization as the constraint condition, and confidence risk minimization as the optimization goal. By calculating the kernel matrix and solving the quadratic programming problem to find support vectors, it has strong generalization ability and robustness, and is applied to fields such as text recognition, image analysis, face recognition, target detection, and time series prediction, and has received extensive attention.
[0003] Although the support vector machine algorithm has good classification performance, problems such as solving the quadratic programming problem and calculating the kernel matrix lead to a computational performance bottleneck with high space-time complexity for the support vector machine algorithm. Especially with the advent of the big data era, the explosively growing data makes the computational performance bottleneck of the support vector machine algorithm more prominent. And the parallelization idea can decompose complex problems into multiple sub-problems for parallel solution to break through the bottleneck. Therefore, the research on the parallel support vector machine algorithm in the big data environment has become a research hotspot for current classification algorithms.
[0004] Facing the problem of difficult calculation brought by the explosively growing data in the big data era, Google has developed a series of distributed computing frameworks to accept the challenges of the big data era. Among them, the MapReduce parallel programming model is favored by many scholars and enterprises due to its simple operation, automatic fault tolerance, strong scalability, etc. Currently, many important progress has been made in the parallel support vector machine algorithm combined with the distributed computing framework MapReduce. Among them, Graf et al. proposed the Cascade Support Vector Machine (Cascade SVM) algorithm based on the idea of early identification of non-support vectors, parallelly trained the subset support vector machine model to eliminate non-support vectors, and merged the support vectors as the training set to continue training until the global support vector machine model. It can be seen from the experimental results that the running time of the cascade support vector machine algorithm is greatly reduced, which greatly avoids the performance bottleneck of the support vector machine algorithm in the big data set environment. However, this algorithm has the following three deficiencies: the algorithm fails to consider problems such as large deviation in subset distribution, low parallel efficiency, and inaccurate filtering of non-support vectors during data partitioning. Summary of the Invention
[0005] The present invention aims to at least solve the technical problems existing in the prior art, and particularly innovatively proposes an optimization method for a parallel support vector machine based on relative entropy and cosine similarity.
[0006] To achieve the above object of the present invention, the present invention provides an optimization method for a parallel support vector machine based on relative entropy and cosine similarity, comprising the following steps:
[0007] S1, data partitioning: adopt a data partitioning strategy DPRE based on relative entropy for data partitioning, balance the relative entropy between the current subset and the original data set, partition the samples into subsets, and reduce the distribution deviation of the subsets;
[0008] S2, parallel SVM training: combine the MapReduce framework to implement a multi-level cascade structure, filter out non-support vectors layer by layer to streamline the training set, and obtain the trained SVM model;
[0009] S3, input the data to be measured into the trained SVM model to obtain the data classification result.
[0010] Further, the data partitioning strategy DPRE based on relative entropy comprises the following steps:
[0011] S1-1, initial partitioning: partition the original training set into non-overlapping internal validation set, positive example initial subset, negative example initial subset and training set to be partitioned, and obtain positive example initial subset and negative example initial subset with similar distributions;
[0012] S1-2, partitioning of the training set to be partitioned: each time select a sample from the training set to be partitioned, calculate and update the relative entropy of the sample in different initial subsets, then calculate the partitioning similarity and partitioning similarity gain in turn, and partition the sample into the subset with the largest similarity gain to further reduce the distribution deviation.
[0013] Further, the S1-1 comprises:
[0014] First, based on the label y, divide the data set D into a data set D + containing only positive labels and a data set D - containing only negative labels; then extract 5% from D + and D - as the internal validation set innerTest respectively; finally, based on the number of samples n + and n - of the current D + and D - , randomly divide D + and D - into k positive example initial subsets with the number of samples being , k negative example initial subsets with the number of samples being and a training set to be partitioned D with the number of samples being w . If the value is a decimal, round up.
[0015] Further, the S1-2 includes:
[0016] First, use the relative entropy calculation formula to calculate the relative entropy of k initial positive example subsets and D + and the relative entropy of k initial negative example subsets and D - where D is a data set containing only positive labels, and D + is a data set containing only negative labels; -
[0017] Subsequently, calculate the partition similarity dsf of the initial positive example subset and the initial negative example subset based on and The partition similarity formula SSF is as follows: + and dsf - where
[0018]
[0019] is the mean value of , k is the number of subsets for partitioning;
[0020] re
[0021] is the relative entropy of D i and D; i
[0022] D i represents the i-th subset;
[0023] D is the original data set;
[0024] Next, traverse all samples in the training set D w to be partitioned, and calculate the relative entropy after the sample d is partitioned into k initial subsets with the same label where the sample d is a certain sample in D w ; and calculate the partition similarity of the partitioned subsets based on and the subset similarity formula Then, propose a partition formula PF, and partition the sample d into a suitable subset based on the pre-partition and the post-partition The partition formula PF is as follows:
[0025]
[0026] where argmax(·) is a function to find the subscript i of the maximum value in the set;
[0027] ssf is the relative entropy of the original data set D and the subset Subset similarity;
[0028] ssf i * be D i Subset similarity after adding sample d;
[0029] n i be D i The number of vectors;
[0030] D i be the i-th subset of D;
[0031] n min is the minimum value among n1, n2, …, n k in,
[0032] n1 is the number of vectors in D1, n2 is the number of vectors in D2, and n k be D k The number of vectors,
[0033] D1 is the 1st subset of D, D2 is the 2nd subset of D, and D k be the k-th subset of D.
[0034] Furthermore, the S2 includes:
[0035] S2-1, Local SVM model training: Based on the k subsets obtained in the data partitioning stage, locally train SVM models in parallel in the Map stage;
[0036] S2-2, Redundant layer detection: After the local SVM model training is completed, calculate the similarity of adjacent layers in parallel at the Map nodes, detect and stop redundant layers;
[0037] S2-3, Non-support vector filtering: Based on multiple local SVM models trained in the Map stage, filter non-support vector machines in the current training set, streamline the training set and continue training in the next layer;
[0038] That is, by combining the distances from samples to the decision boundaries of multiple local support vector models, calculate the support vector similarity to identify non-support vectors, solving the problem of inaccurate filtering of non-support vectors.
[0039] S2-4, Global SVM model construction: Based on multiple local SVM models with early stopping, construct a global SVM model.
[0040] Furthermore, the S2-1 includes:
[0041] (1) Convert the k subsets in the data partitioning stage into k triples A set, where X and y are the feature vector and label of the sample respectively, α is the corresponding Lagrange multiplier and is initially 0, and n is the number of elements in the set; X i , y i , α i Compared with X j , y j , α j The subscripts i and j represent two different models. The model refers to the local support vector machine model generated during the training process. The support vector is actually a special sample, so α i , y i and X i can also be interpreted as the Lagrange multiplier, label, and feature vector of the support vector respectively.
[0042] (2) In the Map stage, the current k sets are distributed to k Map nodes and the sequential minimal optimization algorithm SMO is used to train with α in the set elements {X, y, α} as the Lagrange initial value, where α is the Lagrange multiplier, and the trained Lagrange multiplier is updated into the set to obtain
[0043] Furthermore, the redundant level detection is implemented by using the redundant level detection strategy based on cosine similarity CS-RLDS. CS-RLDS includes:
[0044] Calculate the cosine similarity of the normal vectors between adjacent layers of local support vector machines, compare the set threshold with the similarity, identify and stop the redundant level, which improves the parallel efficiency.
[0045] Furthermore, the specific steps of the redundant level detection strategy based on cosine similarity CS-RLDS are as follows:
[0046] (1) After the local SVM model training is completed, the normal vector ω of the local SVM models in the k Map nodes can be represented by the sample set in this node as:
[0047]
[0048] where n is the number of local SVM support vectors, that is, the sample quantity n;
[0049] α i , y i and X i are the Lagrange multiplier, label, and feature vector of the support vector respectively;
[0050] Φ(·) is the mapping function;
[0051] (2) Combine the k in the previous layer among the k Map nodes* The normal vector of a local SVM and the normal vector ω of the local SVM of the current i-th node i Parallelly calculate the hierarchical similarity HS. The hierarchical similarity formula HS is as follows:
[0052]
[0053]
[0054] where is the inner product of ω i and ;
[0055] ω i represents the normal vector of the local SVM of the current i-th node;
[0056] represents the normal vector of a local SVM of the j-th node in the previous layer;
[0057] ni and nj are the numbers of support vectors corresponding to ω i and respectively;
[0058] α l and α m are the Lagrange multipliers corresponding to the l-th sample and the m-th sample respectively;
[0059] y l and y m are the label values corresponding to the l-th sample and the m-th sample respectively;
[0060] K(·,·) is the kernel function;
[0061] x l , x m represent the l-th vector and the m-th vector respectively;
[0062] (3) Compare the hierarchical similarity HS with the cosine-set threshold τ. If HS ≥ τ, terminate the parallel training SVM stage in advance and transfer the k local SVMs of the current layer to the global SVM construction stage. Otherwise, retain the normal vectors of the current layer for use in the next layer.
[0063] Furthermore, the non-support vector filtering is implemented using the non-support vector filtering strategy NSVF. The non-support vector filtering strategy NSVF includes:
[0064] (1) Coarse filtering: Identify non-support vectors, singular vectors, and support vectors through the coarse identification formula, filter out non-support vectors, and retain support vectors as the refined training set;
[0065] Rough identification formula RI:
[0066]
[0067] where RI(d) represents the rough identification result of sample d;
[0068]
[0069]
[0070] f i (d) is the predicted value of sample d in SVM i ;
[0071] SVM i represents the i-th local SVM model;
[0072] y is the label of sample d;
[0073] (2) Singular vector filtering: Based on the performance of singular vectors in multiple local SVMs, calculate the support vector similarity of singular vectors, and accurately filter out non-support vectors from the singular vectors according to the similarity, and retain the remaining vectors in the refined training set.
[0074] Furthermore, the singular vector filtering includes:
[0075] 1) Evaluate the prediction accuracy of k local SVM models using a cross-training set That is:
[0076]
[0077] where p is the accuracy;
[0078] ni is the number of support vectors from the training set of another model, and another model refers to any one of the k local SVM models other than the current model;
[0079] nj is the number of support vectors of the model to be evaluated, that is, nj is the corresponding number of local SVM support vectors; here, all k local SVM models need to be evaluated, and the model to be evaluated refers to the current model to be evaluated.
[0080] max(·) is the maximum function;
[0081] sign(·) is the sign function;
[0082] α j 、y j 、X ja and b are respectively the Lagrange multipliers, labels, eigenvectors, and intercepts corresponding to the support vectors of a model;
[0083] K(·,·) is the kernel function;
[0084] X i , X j are respectively the vectors of the i-th and j-th samples;
[0085] X i is the eigenvalue of the support vector of the model;
[0086] 2) For each singular vector d and k local SVM models, calculate The k local SVM models are divided into two groups by |f(d)| ≤ 1 and |f(d)| > 1, that is and where represents a models that identify d as a support vector, represents b models that identify d as a non-support vector, is the set of predicted values of the k models for d, and f(d) represents the predicted value of the model for sample d; The singular vector is a special sample and belongs to the samples.
[0087] 3) Calculate the distances from the sample d to be filtered with label y to and the corresponding decision boundaries and Then the distance from sample d to the corresponding decision boundary of the model can be expressed as:
[0088]
[0089] where n is the number of support vectors of the model, i.e., the number of samples n;
[0090] X is the eigenvector;
[0091] b is the intercept in SVM;
[0092] α i , y i and X i are respectively the Lagrange multiplier, label, and eigenvector of the support vector;
[0093] K(·,·) is the kernel function;
[0094] y is the label of sample d;
[0095] y i is the label of sample d i ;
[0096] |·| represents the absolute value;
[0097] 4) Based on the comprehensive performance of the singular vectors in the a·b pairs calculate the support vector similarity of the singular vectors, where respectively represent the i-th local SVM model that classifies the singular vector as a support vector and the j-th local SVM model that classifies the singular vector as a non-support vector; then filter out the vectors with a similarity less than a pre-set threshold μ as non-support vectors, and merge the remaining vectors into the refined training set for continued training. The support vector similarity formula SVSF is as follows:
[0098] Given that the sample d to be filtered is a support vector in a SVMs s and a non-support vector in b SVMs ns then the support vector similarity of d can be expressed as:
[0099]
[0100]
[0101] where and are the accuracies of the i-th SVM s and SVM ns respectively;
[0102] SVM s represents the model that classifies d as a support vector;
[0103] SVM ns represents the model that classifies d as a non-support vector;
[0104] is the accuracy of the a-th SVM s ;
[0105] is the accuracy of the b-th SVM ns ;
[0106] [·] T represents the transpose of the matrix;
[0107] and are the distances from d to the and corresponding decision boundaries respectively.
[0108] In summary, due to the adoption of the above technical solutions, the present invention can significantly improve both the classification efficiency and the classification accuracy.
[0109] By adopting the relative entropy-based data partitioning strategy DPRE, the relative entropy between the current subset and the original data set can be balanced, samples are partitioned into suitable subsets, and the distribution deviation of the subsets is reduced;
[0110] By adopting the cosine similarity-based redundant level detection strategy CS-RLDS, the cosine similarity of the normal vectors between adjacent layer local support vector machines is calculated, the set threshold is compared with the similarity, redundant levels are identified and stopped, and the parallel efficiency is improved;
[0111] By adopting the non-support vector filtering strategy NSVF, combined with the distances from samples to the decision boundaries of multiple local support vector models, the support vector similarity is calculated to identify non-support vectors, and the problem of inaccurate filtering of non-support vectors is solved.
[0112] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Brief Description of the Drawings
[0113] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, wherein:
[0114] Figure 1 is a schematic structural diagram of the parallel SVM training stage of the present invention.
[0115] Figure 2 is a schematic diagram of the speedup ratios of four algorithms of the present invention in different data sets.
[0116] Figure 2 (a) is the speedup ratio of the four algorithms on the Buzz data set.
[0117] Figure 2 (b) is the speedup ratio of the four algorithms on the Income data set.
[0118] Figure 2 (c) is the speedup ratio of the four algorithms on the Covertype data set.
[0119] Figure 2 (d) is the speedup ratio of the four algorithms on the Poker-Hand data set.
[0120] Figure 3 is the classification accuracy of four algorithms of the present invention on different data sets. Detailed Embodiments
[0121] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0122] The present invention proposes an optimization method for a parallel support vector machine based on relative entropy and cosine similarity. The specific embodiments are as follows and include the following steps:
[0123] S1, data partitioning: Adopt the data partitioning strategy DPRE (Data Partitioning Based on Relative Entropy) based on relative entropy to partition medical image data, balance the relative entropy between the current subset and the original data set, and partition samples into subsets.
[0124] S2, parallel SVM training: Combine the MapReduce framework to implement a multi-level cascade structure, layer by layer filter out non-support vectors to streamline the training set, and obtain the trained SVM model.
[0125] S3, input the medical image data to be measured into the trained SVM model to obtain the classification result of the medical image data.
[0126] Based on the advantages of the MapReduce programming model, the present invention proposes a parallel support vector machine algorithm (RC-PSVM) based on relative entropy and cosine similarity. First, the algorithm first proposes a data partitioning strategy DPRE (Data Partitioning Based on Relative Entropy) based on relative entropy, balances the relative entropy between the current subset and the original data set, partitions samples into suitable subsets, and reduces the subset distribution deviation; then proposes a redundancy level detection strategy CS-RLDS (Redundancy level detection strategy based on cosine similarity) based on cosine similarity, calculates the cosine similarity of the normal vectors between adjacent layer local support vector machines, compares the set threshold with the similarity, identifies and stops redundant levels, and improves the parallel efficiency; finally proposes a non-support vector filtering strategy NSVF (Non-Support Vector Filter), combines the distance from the sample to the decision boundaries of multiple local support vector models, calculates the support vector similarity to identify non-support vectors, and solves the problem of inaccurate filtering of non-support vectors. The algorithm proposed by the present invention has significantly improved both in terms of running efficiency and classification accuracy. In addition, the knowledge mined through this method can provide great help in biology, medicine, astronomy and geography.
[0127] 1. Data partitioning
[0128] At present, there is a problem of large deviation in the subset distribution when the parallel SVM algorithm divides data. To address this problem, this paper proposes a data partitioning strategy DPRE based on relative entropy. This strategy mainly includes two steps: (1) Initial partitioning: Divide the original training set into non-overlapping internal validation set, positive example initial subset, negative example initial subset, and training set to be partitioned, obtaining positive example initial subset and negative example initial subset with roughly similar distributions; (2) Partitioning of the training set to be partitioned: Each time, select a sample from the training set to be partitioned, calculate and update the relative entropy of the sample in different initial subsets, then calculate the partitioning similarity and partitioning similarity gain in turn, and partition the sample into the subset with the largest similarity gain to further reduce the distribution deviation.
[0129] 1.1 Initial partitioning
[0130] For the original training set (where X i represents the feature vector of the i-th sample, y i represents the label of the i-th sample, X is the feature vector, y = ±1 is the label, and n is the number of samples), first divide D into the data set D + containing only positive labels and the data set D - containing only negative labels based on the label y; then extract 5% from D + and D - respectively as the internal validation set innerTest; finally, based on the current number of samples n + and n - of D + and n - , randomly divide D + and D - into k positive example initial subsets with the number of samples , k negative example initial subsets with the number of samples , and a training set to be partitioned D with the number of samples w .
[0131] 1.2 Partitioning of the training set to be partitioned
[0132] After obtaining k positive example initial subsets, k negative example initial subsets, and the training set to be partitioned D w through the initial partitioning, first, use the relative entropy calculation formula (1) to calculate the relative entropy + between the k positive example initial subsets and D and the relative entropy - between the k negative example initial subsets and D Subsequently, propose the subset similarity formula SSF (Subset similarity formula), based on and Calculate the partition similarity dsf of the initial positive example subset and the initial negative example subset + and dsf - , the partition similarity formula SSF is as follows:
[0133] Theorem 1 (Subset Similarity Formula SSF): Given that D is the original dataset, are k non-overlapping subsets of D, then the partition similarity between D and the subset can be expressed as:
[0134]
[0135] where k is the number of subsets in the partition, re i is the relative entropy between D i and D, D i represents the i-th subset; is 's mean.
[0136] Proof: Given that relative entropy can be used to measure the similarity between the distributions of two sets, so the mean of the relative entropies of the k subsets and D can represent the average distribution similarity between the k subsets and D, indicating that the average distribution of the k subsets is more similar to the original dataset D. Since is 's second-order central moment, it can represent the degree of difference between the k subsets, indicating that the distributions between the k subsets are more similar. Therefore, adding the average distribution of the k subsets and the distribution difference between the k subsets can comprehensively represent the partition similarity between D and the k subsets obtained by the partition. Therefore, the subset similarity formula SSF effectively measures the similarity after partitioning. Q.E.D.
[0137] Next, traverse all samples in D w and calculate the relative entropy of the sample d after being partitioned into k initial subsets with the same label where the sample d is a certain sample in the training set D to be partitioned w ; and based on and the subset similarity formula (1), calculate the subset similarity after partitioning. Then, a partition formula PF (Partition formula) is proposed. Based on the situation before partitioning and after partitioning partition the sample d into a suitable subset. The partition formula PF is as follows:
[0138] Theorem 2 (Partition formula PF): Given that D is the original dataset, Let \(D_1, D_2, \cdots, D_k\) be \(k\) non - overlapping subsets of \(D\), \(d\in D\) and Then the partitioning formula for \(d\) can be expressed as:
[0139]
[0140] where \(\arg\max(\cdot)\) is a function to find the index \(i\) of the maximum value in the set, \(ssf\) is the subset similarity between the original data set \(D\) and the subset ; \(ssf'\) i * is the subset similarity after adding \(d\) to \(D\), \(n\) i is the number of vectors in \(D\), \(D_i\) i is the \(i\) - th subset of \(D\); \(n'\) i is the number of vectors in \(D\), and \(n_{\min}\) i is the minimum value among \(n_1,n_2,\cdots,n\). min is \(n_1,n_2,\cdots,n\) k .
[0141] Proof: From Theorem 1, we know that \(ssf\) is the subset similarity between the original data set \(D\) and the subset ; \(ssf'\) is the subset similarity after adding the sample \(d\) to \(D\). Therefore, \(ssf - ssf'\) i can represent the degree of change in the subset similarity after adding the sample \(d\) to \(D\). And \((n\) i d - n' + 1)\) i can reflect the degree of difference in the number between the subset \(D\) i and the subset with the smallest number. Therefore, using \((n\) min - n' + 1)\) 2 as a penalty term can balance the number gap between different subsets. In summary, i i min - n' + 1)\) 2 i i can effectively represent the gain of partitioning the sample \(d\) into the subset \(D_i\). Therefore, using the \(\arg\max()\) function to obtain the index value \(i\) corresponding to the maximum gain and partitioning the sample \(d\) into \(D_i\) can effectively reduce the distribution bias. Q.E.D. i i i can effectively reduce the distribution bias. Q.E.D.
[0142] Finally, pairwise merge the \(k\) positive - example subsets and the \(k\) negative - example subsets to obtain \(k\) subsets with similar distributions.
[0143] 2. Parallel SVM Training
[0144] At present, in the big data environment, the parallel SVM algorithm in the training stage is usually based on a cascade structure. Locally parallel SVMs are trained on nodes, and then the support vectors of adjacent nodes are merged pairwise as the input for the next layer. Multiple iterations are performed until all support vectors are merged into the same node to obtain the global SVM model. The cascade structure realizes the parallel training of SVM in the big data environment by partitioning subsets for parallel training and then merging support vectors. However, the characteristic that the number of layers increases with the number of nodes and the filtering strategy of only retaining the support vectors of local SVMs lead to problems of low parallel efficiency and inaccurate filtering of non-support vectors when dealing with large-scale data. To address the above problems, in the parallel SVM training stage, a redundant level detection strategy RLD is first proposed to detect redundant levels and improve parallel efficiency; then a non-support vector filtering strategy NSVF is proposed to comprehensively identify samples using multiple local SVM models and accurately filter non-support vectors.
[0145] The main tasks in the parallel SVM training stage are as follows: Combine the MapReduce framework to implement a multi-level cascade structure and filter non-support vectors layer by layer to streamline the training set. The structure of parallel SVM training is as Figure 1 shown. This structure mainly includes four parts: (1) Local SVM model training: Based on the k subsets obtained in the data partitioning stage, locally parallel SVM models are trained in the Map stage; (2) Redundant level detection: After the local SVM model training is completed, the similarity between adjacent layers is calculated in parallel at the Map nodes to detect and stop redundant levels; (3) Non-support vector filtering: Based on multiple local SVM models trained in the Map stage, non-support vectors in the current training set are filtered, and the streamlined training set continues to be trained in the next layer corresponding to Figure 1 the i-th MapReduce task part in it; (4) Global SVM model construction: Based on multiple local SVM models with early stopping, a global SVM model is constructed;
[0146] 2.1 Local SVM model training
[0147] (1) Convert the k subsets in the data partitioning stage into k triples (where X and y are the feature vector and label of the sample respectively, α is the corresponding Lagrange multiplier and is initially 0, and n is the number of elements in the set) sets;
[0148] (2) In the Map stage, distribute the current k sets to k Map nodes and use the sequential minimal optimization algorithm SMO to train with α in the set element {X, y, α} as the initial value of the Lagrange multiplier, where α is the Lagrange multiplier, and obtain the trained Lagrange multiplier and update it into the set to get
[0149] 2.2 Redundant level detection
[0150] In the traditional cascade structure, the number of layers increases with the increase of the number of initial nodes, and the proportion of support vectors also gradually increases during the process of filtering non-support vectors layer by layer. As a result, the cascade structure has a large training time overhead at higher levels and makes little contribution to the improvement of the algorithm accuracy or even causes negative optimization, leading to low parallel efficiency. To address this problem, a redundant layer detection strategy based on cosine similarity, CS-RLDS, is proposed. The similarity between the current layer and the previous layer is calculated in parallel during the Map stage of each layer as the detection condition to detect and stop redundant layers. The CS-RLDS strategy is as follows:
[0151] (1) After the local SVM model training is completed, the normal vector ω of the local SVM model in k Map nodes can be obtained from the sample set in this node which is expressed as:
[0152]
[0153] where n is the number of local SVM support vectors, α i , y i and X i are the Lagrange multipliers, labels, and feature vectors of the support vectors respectively, and Φ(X) is the mapping function (when the kernel function is not used in SVM training, then Φ(X) = X; when the kernel function K(X i , X j ) is used, then Φ(X i )Φ(X j ) = K(X i , X j ), and at this time Φ(X) cannot be solved).
[0154] (2) A hierarchical similarity formula HS (Hierarchical similarity) is proposed. In k Map nodes, the normal vectors * of the k local SVMs in the previous layer and the normal vector ω of the local SVM in the current i-th node are used to calculate the hierarchical similarity HS in parallel. The hierarchical similarity formula HS is as follows: i Theorem 3 (Hierarchical similarity formula HS): Given the normal vectors
[0155] of the current k local SVMs and the normal vectors of the k local SVMs in the previous layer * then the hierarchical similarity can be expressed as: The hierarchical similarity can be expressed as:
[0156]
[0157]
[0158] where is ω i and inner product, where ni and nj are the numbers of local SVM support vectors corresponding to ω i and respectively, α l and α m are the Lagrange multipliers corresponding to the l-th sample and the m-th sample respectively, y l and y m are the label values corresponding to the l-th sample and the m-th sample respectively, K(x l , x m ) = Φ(x l ) · Φ(x m ) is the kernel function used for training the local SVM, where K(·, ·) is a function for calculating the inner product of vectors in a high-dimensional feature space, also known as the kernel function. When no kernel function is used, then K(x l , x m ) = x l · x m , x i , x m represent the i-th vector and the k-th vector respectively, and · represents the inner product.
[0159] Proof: As can be seen from formula (3), and then:
[0160] Let Z l = α l y l Φ(x l ), then
[0161]
[0162] It can be obtained that Therefore, when the normal vector ω cannot be solved due to the use of the kernel function in the local SVM, the inner product of ω l , x m ) can be calculated through the kernel function K(x i and , and then the cosine similarity of ω i and can be calculated. Since the normal vector ω is also the normal vector of the local SVM hyperplane, it can represent the similarity degree of the SVM i and hyperplanes, where ω i represents the normal vector of the SVM i , so calculating the k normal vectors of the current layer of local SVMs and the k *The mean of the cosine similarities between pairwise local SVM normal vectors can reflect the similarity degree between the current layer and the previous layer of local SVM. Q.E.D.
[0163] (3) Compare the hierarchical similarity HS and the cosine-set threshold τ. If HS ≥ τ, terminate the parallel training SVM stage in advance and transfer the k local SVMs of the current layer to the global SVM construction stage. Otherwise, retain the normal vectors of the current layer for use in the next layer.
[0164] 2.3 Non-Support Vector Filtering
[0165] The essence of the cascade structure lies in parallel training of local SVM models to identify and filter non-support vectors. However, there are inevitably differences between local SVM models and the global SVM model. Therefore, relying solely on a single local SVM model for non-support vector filtering will result in inaccurate non-support vector filtering, thereby affecting the algorithm performance. For this reason, a non-support vector filtering strategy NSVF is proposed to accurately filter non-support vectors by integrating multiple local SVMs. The NSVF strategy mainly includes two steps: (1) Rough filtering: A rough filtering formula (Rough identification, RI) is proposed to identify non-support vectors, singular vectors, and support vectors, filter non-support vectors, and retain support vectors for the refined training set; (2) Singular vector filtering: A support vector similarity formula (Support vector similarity formula, SVSF) is proposed to comprehensively consider the performance of singular vectors in multiple local SVMs, calculate the support vector similarity of singular vectors, and accurately filter non-support vectors from singular vectors according to the similarity, retaining the remaining vectors in the refined training set;
[0166] (1) Rough filtering:
[0167] For the sample d to be filtered, the positive / negative and the absolute value of the prediction value f(d) of the single local SVM model obtained from formula (3) for d reflect the prediction result of the model for d and the distance between the two. Therefore, combining the prediction result f(d) and the true label y of the sample d can obtain the degree of association between the sample d and the single local SVM model. When facing multiple local SVM models, the degrees of association between the sample d and these models are not exactly equal, and there may even be a complete contradiction. Moreover, the training sets of these models are derived from multiple subsets with similar distributions after dividing the same original training set, resulting in a strong correlation between the models. Therefore, it is very likely that the sample d is the support vector or non-support vector in multiple local SVM models at the same time, and there is also a certain possibility that it is a singular vector with completely opposite prediction results. Based on this, the main task of rough filtering is to identify the sample d to be filtered as a support vector, non-support vector or singular vector, directly filter out the non-support vectors, retain the support vectors for continued training in the refined training set, and continue to identify and filter the singular vectors in the next stage. The specific process is as follows:
[0168] 1) Based on k local SVM models, calculate the prediction values of the sample d to be predicted in the k local SVM models
[0169] 2) Propose a rough identification formula RI, and based on identify the sample to be predicted, retain the identified support vectors, filter out the non-support vectors, and leave the singular vectors for continued identification in the next stage; the rough identification formula RI is as follows:
[0170] Theorem 4 (Rough Identification Formula RI): Given SVM1, SVM2, …, SVM k are k local SVM models, then the rough identification result of the sample d can be expressed as:
[0171]
[0172]
[0173]
[0174] where f i (d) is the prediction value of the sample d in SVM i .
[0175] Proof: Since f i (d) is the prediction value of the sample d in the model SVM i , and the attribute value y = ±1 of d is the true label, so f i (d)·y reflects whether SVM i predicts the sample d correctly and the distance from the sample d to SVM iRegarding the distance of the separating hyperplane, it can be obtained that when f i (d)·y > 1, d is correctly classified by the model and the margin is greater than 1, that is, a support vector. When 0 < f i (d)·y < 1, d is correctly classified by the model and the margin is less than 1, that is, a non-support vector. When f i (d)·y ≤ 0, d is misclassified by the model, that is, a non-support vector. And from formulas (7) and (8), it can be known that f max (d) and f min (d) are respectively the maximum value and the minimum value of , that is, when f max (d) ≤ 1 at this time, d is a non-support vector in all local SVM models. When f min (d) > 1 at this time, d is a support vector in all local SVM models. When at this time, d is both a support vector and a non-support vector in all local SVM models. To sum up, the rough recognition formula RI can identify support vectors, non-support vectors, and singular vectors based on the performance of vectors in multiple local SVM models. Q.E.D.
[0176] (2) Singular vector filtering
[0177] In the rough filtering stage, most of the non-support vectors and support vectors in the dataset are briefly identified, while it is difficult to identify singular vectors. Since the proportion of different judgments on the singular vector d, the distance between d and each model, and the difference in the prediction accuracy of each model in all local SVM models all reflect to a certain extent the possibility of d being a support vector. Based on this, the main task of singular vector filtering is to calculate the similarity of singular vectors by combining the comprehensive representation of singular vectors in multiple local SVM models, accurately identify non-support vectors and filter them, and retain the remaining vectors in the refined training set. The specific process is as follows:
[0178] 1) Evaluate the prediction accuracy of k local SVM models on the cross-training set That is:
[0179]
[0180] where p is the accuracy, K(·,·) is the kernel function, X i , X j are the vectors of the i-th and j-th samples respectively, nj is the number of support vectors of the model to be evaluated, α j , y j , X ja and b are the Lagrange multipliers, labels, feature vectors, and intercepts corresponding to the support vectors of the model, respectively. Subscripts i and j represent two different models. ni (i≠j) is the number of support vectors in the training set from another model, and X i is the eigenvalue of the support vector of this model. max(·) is the maximum function, and sign(·) is the sign function.
[0181] 2) For each singular vector d and k local SVM models, the predicted value of d can be calculated by the model, and we can obtain as the set of predicted values of d by k models. The k local SVM models are divided into two groups based on |f(d)|≤1 and |f(d)|>1, namely: a models that identify d as a support vector and b (b = k - a) models that identify d as a non - support vector where the superscript s represents support and the superscript ns represents non - support, and f(d) represents the predicted value of the model for the sample d.
[0182] 3) For the grouped and Calculate the distances from the sample d with label y to be filtered to the and corresponding decision boundaries and Then the distance from the sample d to the decision boundary corresponding to the model can be expressed as:
[0183]
[0184] where n and b are the number of support vectors of the model and the intercept in SVM, and α i 、y i and X i are the Lagrange multiplier, label, and feature vector of the support vector respectively, y is the label of the sample d, and K(·,·) is the kernel function.
[0185] 4) Propose the support vector similarity formula SVSF. Based on the comprehensive performance of the singular vector in a·b pairs calculate the support vector similarity of the singular vector, where represent the i - th local SVM model that classifies the singular vector as a support vector and the j - th local SVM model that classifies the singular vector as a non - support vector respectively; then vectors with similarity less than the pre - set threshold μ are identified as non - support vectors and filtered, and the remaining vectors are merged into the reduced training set for further training. The support vector similarity formula SVSF is as follows:
[0186] Theorem 5 (Support Vector Similarity Formula SVSF): Given that the sample d to be filtered is a support vector in a SVMs s and a non - support vector in b SVMsns If \(d\) is a non-support vector, the support vector similarity of \(d\) can be expressed as:
[0187]
[0188]
[0189] where and are the accuracies of the \(i\)-th SVM s and SVM ns respectively. SVM s represents the model that classifies \(d\) as a support vector, and SVM ns represents the model that classifies \(d\) as a non-support vector. is the accuracy of the \(a\)-th SVM s . is the accuracy of the \(b\)-th SVM ns . [·] T represents the transpose of a matrix. and are the distances from \(d\) to the and corresponding decision boundaries respectively.
[0190] Proof: Since is the distance from \(d\) to the corresponding decision boundary, is the distance from \(d\) to the corresponding decision boundary, and are models constructed from different subsets of the original dataset. Therefore, by using the local model that identifies \(d\) as a support vector and the local model that identifies \(d\) as a non-support vector as references, and performing normalization using , it can reflect the degree to which \(d\) approaches the corresponding decision of the global model constructed from the original dataset, that is, indicates that the higher the probability of \(d\) becoming a support vector in the global model, indicates that the lower the probability of \(d\) becoming a support vector. And by cross-combining the distances from \(d\) to the corresponding decision boundaries of \(a\) SVMs s and \(b\) SVMs ns to obtain the matrix \([h ij a×b , and then supplementing it with the probability vector s of SVM to perform weighted sum on the \(b\) columns \([h 1j h 2j … h aj T and the reciprocal probability vector of SVM ns Penalize the a-th row [h i1 h 2j … h bj , comprehensively measure the possibility of d becoming a support vector in the global model, and it is best to multiply by The value range can be normalized between [0, 1]. It can be comprehensively obtained that the support vector similarity formula can be used as the support vector similarity formula for measuring singular vectors. Q.E.D.
[0191] 3. Effectiveness of the Parallel Support Vector Machine Algorithm (RC-PSVM) Based on Relative Entropy and Cosine Similarity
[0192] To verify the classification effect of the RC-PSVM algorithm, we applied the RC-PSVM method to four datasets, namely Buzz, Income, Covertype, and Poker-Hand. The specific information is shown in Table 1. The RC-PSVM, L-SVM, Projection-SVM, and AM-SVM algorithms were compared in terms of classification accuracy, etc.
[0193] Table 1 Detailed Information of Datasets
[0194]
[0195] 3.1 Parallel Performance Analysis of the RC-PSVM Method
[0196] To verify the speedup ratio of the RC-PSVM algorithm, comparative experiments were conducted on the RC-PSVM algorithm, L-SVM algorithm, Projection-SVM algorithm, and AM-SVM algorithm on four datasets, namely Buzz, Income, Covertype, and Poker-Hand, respectively. The speedup ratio was used as a measurement index, and the speedup ratios of each algorithm under different numbers of nodes were compared respectively, and then the performance of each algorithm was compared and analyzed. The experimental results are as follows:
[0197] From Figure 2 it can be seen that the speedup ratios of the four algorithms in each dataset increase with the increase in the number of nodes, and reach the maximum when the number of nodes is 8. Among them, the upward trend of the speedup ratio of the RC-PSVM algorithm with the increase in the number of nodes in each dataset is more significant than that of the other three algorithms, and it always has the highest speedup ratio. As Figure 2 (a) shows, compared with the L-SVM, Projection-SVM, and AM-SVM algorithms, the speedup ratios of the RC-PSVM algorithm when the number of nodes is 8 are increased by 1.67, 0.56, and 0.84 respectively; as Figure 2As shown in (b), compared with the other three algorithms, the speedup ratios of the RC-PSVM algorithm when the number of nodes is 8 are increased by 1.11, 0.49, and 0.37 respectively; as Figure 2 As shown in (c), compared with the other three algorithms, the speedup ratios of the RC-PSVM algorithm when the number of nodes is 8 are increased by 2.15, 0.70, and 1.03 respectively; as Figure 2 As shown in (d), compared with the other three algorithms, the speedup ratios of the RC-PSVM algorithm when the number of nodes is 8 are increased by 1.85, 0.77, and 0.45 respectively. It can be seen from the data analysis that compared with the L-SVM, Projection-SVM, and AM-SVM algorithms, the RC-PSVM algorithm has better speedup ratio performance. There are mainly two reasons for this: on the one hand, the RC-PSVM algorithm designs a redundant layer detection strategy CS-RLDS to identify and eliminate redundant layers, avoiding invalid calculations and improving the speedup ratio; on the other hand, the RC-PSVM algorithm designs a non-support vector filtering strategy NSVF to accurately identify and filter non-support vectors, greatly reducing the data scale of the next layer and reducing the training time. Therefore, the RC-PSVM algorithm has a higher speedup ratio and higher parallel efficiency in the case of a larger number of nodes.
[0198] 3.2 Classification accuracy analysis of the RC-PSVM algorithm
[0199] To evaluate the classification performance of the RC-PSVM algorithm, F-Measure is used as the evaluation index. The RC-PSVM algorithm, L-SVM algorithm, Projection-SVM algorithm, and AM-SVM algorithm are respectively compared in four datasets. In the experiment, the classification accuracy F-Measure of the above algorithms is compared respectively. The experimental results are as Figure 3 shown.
[0200] From Figure 3As can be seen, in the four datasets, the classification accuracy of the RC-PSVM algorithm is always higher than that of the L-SVM, Project-SVM, and AM-SVM algorithms. In the Buzz dataset, compared with the L-SVM, Projection-SVM, and AM-SVM algorithms, the classification accuracy of the RC-SVM algorithm is 4.52%, 1.11%, and 2.43% higher respectively; in the Income dataset, compared with the L-SVM, Projection-SVM, and AM-SVM algorithms, the classification accuracy of the RC-SVM algorithm is 3.07%, 0.43%, and 1.61% higher respectively; in the Covertype dataset, compared with the L-SVM, Projection-SVM, and AM-SVM algorithms, the classification accuracy of the RC-SVM algorithm is 5.54%, 0.87%, and 1.46% higher respectively; in the Poker-Hand dataset, compared with the L-SVM, Projection-SVM, and AM-SVM algorithms, the classification accuracy of the RC-SVM algorithm is 3.11%, 1.21%, and 0.5% higher respectively. From the above data, it can be seen that the RC-PSVM has a significant advantage in classification accuracy in the four datasets, and the advantage is more obvious in the Buzz dataset and the Covertype dataset. The main reasons for this result are: (1) The RC-PSVM algorithm designs a data partitioning strategy DPRE based on relative entropy. By balancing the relative entropy between the subset and the original dataset, the distribution deviation of the subset is reduced, making the local support vector machine model closer to the global model and improving the algorithm accuracy; (2) The RC-PSVM algorithm designs a non-support vector filtering strategy NSVF, which accurately identifies and filters non-support vectors, avoids the omission of support vectors, and ensures the algorithm accuracy. Therefore, from the above experimental comparison results, it can be seen that compared with the L-SVM, Projection-SVM, and AM-SVM algorithms, the RC-PSVM algorithm has better classification accuracy on the four datasets.
[0201] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A parallel support vector machine optimization method based on relative entropy and cosine similarity, characterized in that It includes the following steps: S1, data partitioning: Adopt the data partitioning strategy DPRE based on relative entropy to partition medical images, balance the relative entropy between the current subset and the original dataset, and partition the samples into subsets; S2, parallel SVM training: Combine the MapReduce framework to implement a multi-level cascade structure, filter out non-support vectors layer by layer to streamline the training set, and obtain the trained SVM model; S3, input the medical image data to be tested into the trained SVM model to obtain the classification result of the medical image data; The data partitioning strategy DPRE based on relative entropy includes the following steps: S1-1, initial partitioning: Partition the original training set into non-overlapping internal validation sets, positive example initial subsets, negative example initial subsets, and the training set to be partitioned, and obtain positive example initial subsets and negative example initial subsets with similar distributions; S1-2, partitioning of the training set to be partitioned: Each time, select a sample from the training set to be partitioned, calculate and update the relative entropy of the sample in different initial subsets, then calculate the partitioning similarity and partitioning similarity gain in turn, and partition the sample into the subset with the largest similarity gain; S2 includes: S2-1, Local SVM Model Training: Based on the subsets obtained in the data partitioning phase, locally parallel training SVM models in the Map phase; S2-2, redundant layer detection: After the local SVM model training is completed, calculate the similarity between adjacent layers in parallel at the Map node, detect and stop redundant layers; S2-3, non-support vector filtering: Based on multiple local SVM models trained in the Map stage, filter out non-support vector machines in the current training set, and streamline the training set to continue training in the next layer; S2-4, global SVM model construction: Based on multiple local SVM models with early stopping, construct a global SVM model.
2. The parallel support vector machine optimization method based on relative entropy and cosine similarity according to claim 1, wherein, S1-1 includes: First, based on the labels divide the data set into a data set containing only positive labels and a data set containing only negative labels ; then, extract from and respectively as the internal validation set ; finally, based on the current and sample numbers and , divide and randomly into positive example initial subsets with a sample number of , negative example initial subsets with a sample number of and a training set to be divided with a sample number of . 3. A parallel support vector machine optimization method based on relative entropy and cosine similarity according to claim 1, characterized in that S1-2 includes: First, use the relative entropy calculation formula to calculate the relative entropy of initial subsets of positive examples and respectively and initial subsets of negative examples and respectively ; where is the data set containing only positive labels, is the data set containing only negative labels; Subsequently, based on and calculate the partition similarity of the initial positive example subset and the initial negative example subset and , and the partition similarity formula is as follows: , wherein is mean value of; The number of subsets to be partitioned; is and relative entropy of; Indicates the th subset; is the original dataset; Next, traverse all samples in the training set to be partitioned and calculate the relative entropy of the sample after being partitioned into initial subsets with the same label , where the sample is a certain sample in; and calculate the subset similarity after partitioning based on and the subset similarity formula; then, propose the partitioning formula , and based on the before partitioning and the after partitioning, partition the sample into a suitable subset. The partitioning formula is as follows: , Among them is a function to find the subscript of the maximum value in a set ; is the original dataset and the subset similarity of the subset; For the subset similarity after adding a sample For the number of vectors; is the subset; is the minimum value in is the number of vectors of is the number of vectors of is the number of vectors of is the first subset of is the second subset of is the -th subset of 4. A parallel support vector machine optimization method based on relative entropy and cosine similarity according to claim 1, characterized in that, S2-1 includes: (1) Convert the subsets in the data partitioning stage into triple sets, where are the feature vectors and labels of the samples respectively, the corresponding Lagrange multipliers and initially 0, is the number of set elements; (2) During the Map stage, the current sets are distributed to Map nodes and the Sequential Minimal Optimization algorithm SMO is used to train with the in the as the initial Lagrangian value, where is the Lagrange multiplier, and the trained Lagrange multiplier is obtained and updated into the set to get .
5. A parallel support vector machine optimization method based on relative entropy and cosine similarity according to claim 1, characterized in that The redundant layer detection is implemented using the redundant layer detection strategy CS-RLDS based on cosine similarity. CS-RLDS includes: Calculate the cosine similarity of the normal vectors between adjacent layer local support vector machines, compare the set threshold with the similarity, and identify and stop redundant layers.
6. The parallel support vector machine optimization method based on relative entropy and cosine similarity according to claim 5, characterized in that The specific steps of the redundant layer detection strategy CS-RLDS based on cosine similarity are as follows: (1) After the local SVM model is trained, The normal vectors of the local SVM models in the Map nodes can be represented by the sample sets in these nodes as follows: , Among them is the number of local SVM support vectors; , and are the Lagrange multipliers, labels, and feature vectors of the support vectors, respectively; is a mapping function; (2) Combine the normal vectors of the local SVMs of the previous layer in Map nodes with the normal vectors of the local SVMs of the current and the th node to calculate the hierarchical similarity in parallel . The hierarchical similarity formula is as follows: , , wherein is and inner product of; Indicates the normal vector of the local SVM of the current th node; Indicates the normal vector of the local SVM of the th node in the upper layer; and are respectively and the corresponding number of local SVM support vectors; and are respectively the Lagrange multipliers corresponding to the th sample and the th sample; and are the label values corresponding to the th sample and the th sample, respectively; is a kernel function; respectively represent the th vector and the th vector; (3) Compare the hierarchical similarity with the cosine-set threshold , if , then end the parallel training SVM phase in advance and transfer the local SVMs of the current layer to the global SVM construction phase. Otherwise, retain the normal vector of the current layer for use in the next layer.
7. A parallel support vector machine optimization method based on relative entropy and cosine similarity according to claim 1, characterized in that The non-support vector filtering is implemented using the non-support vector filtering strategy NSVF. The non-support vector filtering strategy NSVF includes: (1) Coarse filtering: Identify non-support vectors, singular vectors, and support vectors through the coarse identification formula, filter out non-support vectors and retain support vectors as the streamlined training set; Rough recognition formula : , Among them represents the rough recognition result of the sample ; , , is the sample in is the predicted value; Indicates the th local SVM model; is a sample label; (2) Singular vector filtering: Synthesize the performance of singular vectors in multiple local SVMs, calculate the support vector similarity of singular vectors, and accurately filter non-support vectors from singular vectors according to the similarity, and retain the remaining vectors in the streamlined training set.
8. An optimization method for a parallel support vector machine based on relative entropy and cosine similarity according to claim 7, characterized in that The singular vector filtering includes: 1) Cross-validation set evaluation Prediction accuracy of individual local SVM models That is: , Among them is the accuracy rate; is the number of support vectors for the training set from another model; is the number of support vectors of the model to be evaluated; For the maximum function; is the sign function; , , and are the Lagrange multipliers, labels, feature vectors, and intercepts corresponding to the support vectors of a model, respectively; is a kernel function; They are respectively the and the vectors of the samples; 2) For each singular vector and local SVM models, calculate which is composed of and Divide local SVM models into two groups, namely and ; where represents models that identify as support vectors, represents models that identify as non-support vectors, is the set of predicted values of the models for represents the predicted value of the model for the sample ; 3) Calculate the sample to be filtered with label to and and the distances to the corresponding decision boundaries and If so, the distance from the sample to the corresponding decision boundary of the model can be expressed as: , Among them is the number of model support vectors; is the eigenvector; is the intercept in SVM; , and are the Lagrange multipliers, labels, and feature vectors of the support vectors, respectively; is a kernel function; is a sample label; is the sample label; Represents the absolute value; 4) Based on the comprehensive performance of the singular vectors in pair , calculate the support vector similarity of the singular vectors, where respectively represent the th local SVM model that divides the singular vectors into support vectors and the th local SVM model that divides the singular vectors into non-support vectors; then identify the vectors with similarity less than the pre-set threshold as non-support vectors for filtering, and merge the remaining vectors into the reduced training set for continued training. The support vector similarity formula is as follows: Known samples to be filtered Among ones are support vectors, and among ones are non - support vectors, then The support vector similarity can be expressed as: , , wherein and are respectively the th and accuracy rates; Denotes a model that is divided into support vectors; Indicates that is divided into a model of non-support vectors; For the th accuracy rate; For the th accuracy rate; Denotes the transpose of a matrix; and are respectively to and the distances to the corresponding decision boundaries.