A data classification method with knowledge transfer discrimination ability

By introducing weight coefficients and weighted WMMD into the classification model and combining LDA dimensionality reduction technology, the poor classification results caused by inconsistent feature distribution of training samples and test samples in the existing technology are solved, and higher knowledge transfer capabilities and classification accuracy are achieved.

CN114861813BActive Publication Date: 2025-05-30HENAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210568499.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-24
Publication Date
2025-05-30
Estimated Expiration
2042-05-24

AI Technical Summary

Technical Problem

The classification results of the existing classification model are not satisfactory when the feature distribution of training samples and test samples are inconsistent. The existing transfer learning method ignores the differences in the global metrics of a single sample, which affects the performance of the classifier.

Method used

By designing weight coefficients and weighted WMMD, the system's knowledge transfer capabilities are improved, and LDA is introduced to reduce the dimensionality of the source and target domain samples, so that the classification model can more accurately distinguish different types of data.

Benefits of technology

Effective knowledge migration under different data distributions of source and target domains is realized, and classification accuracy is improved, making the classification model perform better in cross-domain data classification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114861813B_ABST
    Figure CN114861813B_ABST
Patent Text Reader

Abstract

The present invention relates to a data classification method with the ability to discriminate knowledge transfer. The present invention effectively solves the problem that the existing classification models cannot perform classification and have low classification accuracy in the case where the data distributions of the source domain and the target domain are different; the technical solutions adopted include: by designing Weighted Maximum Mean Divergence (WMMD), the cross-domain knowledge transfer ability of the system is improved. In order to distinguish different types of data, Linear Discriminant Analysis (LDA) is introduced to reduce the dimensions of the source domain samples and the target domain samples, so that the data after dimension reduction has the smallest within-class variance and the largest between-class variance, and the classification accuracy is improved compared with the existing classification models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image and text data classification, and particularly to a data classification method with knowledge transfer discrimination ability. Background Art

[0002] Recently, machine learning has achieved great success in fields such as autonomous driving, speech recognition, computer vision, and object detection. However, for some real-world scenarios, it still has some limitations. Building an accurate machine learning model requires a large number of different types of training samples. However, in reality, the data obtained is usually unclassified;

[0003] Existing classification models such as Softmax regression are suitable for multi-classification and have the advantages of simple application, easy training, and intuitive results. Therefore, they have received more attention from researchers. Traditional Softmax uses the cross-entropy loss function for optimization, but it cannot guarantee the optimized feature distribution. The ideal application scenario of the Softmax classifier is that the training samples and test samples have the same feature distribution. However, in many cases, collecting training samples that meet the requirements is usually time-consuming and expensive, which makes the classification results of the Softmax model unsatisfactory;

[0004] For the case where the training samples and test samples do not have the same distribution, training transfer learning is a relatively effective method to solve the above problems. However, existing transfer learning methods only focus on the overall distribution information and global structure information of the source domain (training samples) and target domain (test samples), ignoring the differences in the contributions of individual samples to the global metric, thus affecting the performance of the classifier;

[0005] In view of the above, we provide a data classification method with knowledge transfer discrimination ability to solve the above problems. Summary of the Invention

[0006] In view of the above situation, the present invention provides a data classification method with knowledge transfer discrimination ability. By designing the weight coefficient and weighted WMMD, the system's ability to cross knowledge transfer is improved. To distinguish different types of data, LDA is introduced to reduce the dimensions of the source domain samples and target domain samples, making the data after dimensionality reduction have the smallest within-class variance and the largest between-class variance. Compared with existing classification models, the classification accuracy is improved.

[0007] A data classification method with knowledge transfer discrimination ability, characterized by including a source domain, a target domain, a feature space, a Softmax classifier, and a prediction result. The process includes the following steps:

[0008] S1: After the source domain samples and target domain samples are feature-mapped, their approximate same distribution can be found in the feature space;

[0009] S2: Perform WMMD and Linear Discriminant Analysis (LDA) operations in the feature space. The WMMD unit calculates the differences in the marginal distribution and conditional distribution between source domain samples and target domain samples, and tries to eliminate the differences between the source domain and the target domain. The Linear Discriminant Analysis (LDA) unit aims to minimize the distance between samples of the same class and maximize the distance between different samples;

[0010] S3: The Softmax classifier calculates the probabilities for each label corresponding to each sample in the target domain respectively, and selects the label corresponding to the maximum probability value as the label of this sample, thus completing the sample classification;

[0011] S4: The prediction result saves the classification result of the target domain samples, that is, the label corresponding to each sample.

[0012] The beneficial effects of the above technical solutions are as follows:

[0013] (1) In the case where the data distributions of the source domain (training samples) and the target domain (test samples) are different, apply the classification model learned in the source domain to the target domain classification task by using the knowledge transfer method;

[0014] (2) By designing weighted MMD (referred to as WMMD, MMD is a metric method that measures the distance between different types of data of source domain samples and target domain samples. WMMD takes into account the different contributions of different types of samples during the measurement and realizes it by adding a weight coefficient W during the measurement), the system's cross - knowledge transfer ability is improved, thereby realizing the application of the classification model learned in the source domain to the target domain classification task by using the knowledge transfer method. Compared with the existing classification models, the classification accuracy is improved;

[0015] (3) In order to distinguish different types of data, LDA is introduced to reduce the dimensions of the source domain samples and target domain samples, so that the data after dimensionality reduction has the smallest within - class variance and the largest between - class variance, further improving the classification performance of the classification model. Brief Description of the Drawings

[0016] Figure 1 It is the block diagram of the Softmax - LWMMD system of the present invention;

[0017] Figure 2 It is the internal structure block diagram of WMMD of the present invention;

[0018] Figure 3 It is the schematic diagram of the calculation principle of Linear Discriminant Analysis (LDA) of the present invention;

[0019] Figure 4 It is a schematic diagram showing the relationship between the values of parameters α and β of the present invention and the accuracy rate;

[0020] Figure 5 It is a schematic diagram showing the relationship between the values of parameters λ and μ of the present invention and the accuracy rate;

[0021] Figure 6 It is a schematic diagram showing the relationship between the number of iterations of the present invention and the accuracy rate. Detailed implementation manners

[0022] Regarding the foregoing and other technical contents, features and effects of the present invention, they can be clearly presented in the following detailed description of the embodiments in conjunction with the attached Figures 1 to 6 drawings. The structural contents mentioned in the following embodiments are all referenced to the drawings of the specification.

[0023] Embodiment 1. This embodiment provides a data classification method with knowledge transfer discrimination ability, which is characterized in that it includes a source domain (the source domain consists of two parts, sample x and sample label y x , and the source domain is represented by ), a target domain (there is only a sample in the target domain, represented by ), a feature space, a Softmax classifier, and a prediction result;

[0024] Its classification process includes the following steps:

[0025] S1: After the source domain samples and the target domain samples are feature-mapped, their approximate same distribution can be found in the feature space;

[0026] S2: Perform WMMD and Linear Discriminant Analysis (LDA) operations in the feature space. The WMMD unit calculates the differences in the marginal distribution and conditional distribution between the source domain samples and the target domain samples, and tries to eliminate the differences between the source domain and the target domain;

[0027] As shown in the attached Figure 3 drawings, it is a schematic diagram of the LDA calculation principle. The purpose of the Linear Discriminant Analysis (LDA) unit is to minimize the distance between samples of the same class (make the data of the same class as close as possible) and maximize the distance between different samples (make the data of different classes as far away as possible);

[0028] In the attached Figure 1 drawings, the pseudo-label means the label assigned to the target sample during the training process, which is not necessarily the true label, and the pseudo-label will be updated after each training. The condition for the end of the training is: there is a pseudo-label that no longer updates or the maximum number of training times (Tmax) is reached;

[0029] S3: The Softmax classifier calculates the probabilities corresponding to each label for each sample in the target domain respectively, and selects the label corresponding to the maximum probability value as the label of this sample, thus completing the sample classification;

[0030] S4: The prediction result saves the classification result of the target domain samples, that is, the label corresponding to each sample.

[0031] I. To implement the above classification steps process, first construct the system objective function: (1)

[0032] Among them,

[0033]

[0034]

[0035]

[0036]

[0037]

[0038] In the above function, θ is all the system item parameters, that is, the parameters that the system needs to determine through training;

[0039] J 1 (θ) is the cost function of Softmax regression;

[0040] J 2 (θ) is based on the WMMD calculation result;

[0041] J 3 (θ) is used to reduce the distribution difference between domains in the probability output layer of Softmax regression through the conditional distribution and the marginal distribution;

[0042] J 4 (θ) is the sparse control term of the model parameters to prevent the model from overfitting;

[0043] J 5 (θ) is the LDA calculation item in the system;

[0044] α, β, λ, μ are the balance parameters of J 2 (θ), J 3 (θ), J 4 (θ), J 5 (θ) respectively;

[0045] L 0 and W 0 are the marginal distribution matrix and the marginal distribution weight coefficient matrix respectively;

[0046] L c and W c are the conditional distribution matrix and the conditional distribution weight coefficient matrix, respectively;

[0047] G b is the between-class matrix, G u is the within-class matrix, G v is the total divergence matrix, G v =G b +G u 。

[0048] L 0 The matrix is calculated by the following formula:

[0049]

[0050] where, is the source domain, is the target domain, x is the sample in the corresponding domain, and n is the number of samples in the corresponding domain;

[0051] The weight coefficient W 0 is defined as follows:

[0052]

[0053] where, is the weight coefficient of the i-th sample, is the mean of the source domain or target domain samples;

[0054] L c The matrix is calculated according to the following formula:

[0055]

[0056] where, is the set of source domain samples belonging to class c, is the correct label of the source domain sample x i correspondingly, is the set of target domain samples belonging to class c;

[0057] The class weight coefficient W c is defined as follows:

[0058]

[0059] where, is the weight coefficient of the i-th sample belonging to class c, is the mean of the source domain or target domain samples belonging to class c.

[0060] The derivation process of the above formula (3) is as follows:

[0061] The calculation process of the WMMD marginal distribution and the conditional distribution calculation formula is as follows:

[0062] The calculation of the marginal distribution distance is as follows:

[0063]

[0064] Among them, the matrix L 0 , and the weight coefficient W 0 The calculation formulas are formulas (7) and (8). Our innovation is to add ω i x i and ω j x j terms in formula (10-1), which represent the source domain and the target domain respectively;

[0065] After adding the weight coefficient to the conditional distribution distance objective function expression, it is as follows:

[0066]

[0067] Among them, the matrix L c , and the weight coefficient W c The calculation formulas are formulas (9) and (10). Our innovation is to add ω i x ic and ω j x jc terms in formula (10-2), which represent different types of data in the source domain and the target domain respectively;

[0068] Based on formulas (10-1) and (10-2), the formula based on the WMMD calculation structure is finally obtained as follows:

[0069]

[0070] As shown in the appendix Figure 2 is the internal structure diagram of WMMD. The traditional MMD (Maximum Mean Divergence) only considers the overall information of the samples when calculating the marginal distribution and the conditional distribution, and does not consider the influence of individual sample information on the data distribution (does not consider the contribution of individual samples in migration). MMD only considers the L matrix, while WMMD not only considers the overall information of the samples, but also considers the contribution of individual samples. A weight coefficient W is added to each sample (to distinguish the contribution degree of each sample, a weight coefficient is added to each sample. Samples with more contributions have larger weights, and vice versa, samples with less contributions have smaller weights), so as to make the classification accuracy higher.

[0071] The calculation of the between-class matrix is as follows:

[0072]

[0073] Here represents the mean of all samples;

[0074] The within-class matrix is calculated as follows:

[0075]

[0076] Here represents the mean of the samples in the \(i\)-th class, represents the data subset of the \(i\)-th class samples in the total samples;

[0077] From equations (11) and (12), the total scatter matrix is obtained:

[0078]

[0079] II. Optimize the above system objective function

[0080] Minimizing the objective function \(J(\theta)\) gives the optimal solution of the parameter \(\theta\). Since equation (1) has no closed-form solution, its solution can be obtained using the gradient descent method. The partial derivative of the objective function \(J(\theta)\) with respect to the parameter \(\theta\) belonging to the \(j\)-th class is:

[0081]

[0082] where,

[0083]

[0084]

[0085]

[0086]

[0087]

[0088] The model parameters are updated by the following formula,

[0089]

[0090] After obtaining the model parameter \(\theta\), classify the target domain The probability that each test sample \(x\) belongs to class \(j\) is:

[0091]

[0092] Calculate the probabilities that \(x\) belongs to all classes through equation (21), and the class corresponding to the maximum probability is the classification class of \(x\).

[0093] III. The following demonstrates the operation process of the above system:

[0094] First, input: source domain samples Target domain samples Parameters α, β, λ, μ, and the maximum number of iterations T max .

[0095] 1. Initialize the model parameter θ.

[0096] 2. Construct the marginal distribution difference matrix L according to formula (7) 0 , let L 0 = 0. At the same time, construct the marginal distribution weight matrix W according to formula (8) 0 , let W 0 = 0.

[0097] 3. Calculate the G b , G u and G v matrices respectively according to formulas (11), (12), and (13).

[0098] 4. Calculate the objective function J(θ) according to formula (1), and then find the partial derivative value according to formula (14).

[0099] 5. Optimize the parameter θ value using the gradient descent method.

[0100] 6. Substitute the parameter θ value into the Softmax regression model to classify the target samples and predict their labels

[0101] 7. Calculate the conditional distribution difference matrix L c according to formula (9), and at the same time, calculate the conditional distribution weight matrix W c .

[0102] 8. Jump to step 3 and continue to execute until the predicted label no longer changes, or reaches the maximum number of iterations T max .

[0103] 9. Terminate the program (Output: model parameter θ and the classification prediction labels of the target domain samples ).

[0104] IV. The following verifies the classification performance of this solution through specific experimental data

[0105] 1. Description of the dataset

[0106] The experiment was completed on five datasets. The datasets were divided into three groups, namely two groups of image datasets: Office-31 + Caltech-256 and MSRC + VOC2007, and one group of text dataset: Reuters-21578. Table 2 shows the basic information of the image datasets.

[0107]

[0108] Table 2 Image Dataset Information

[0109] Office-31+Caltech-256: Office-31 is a commonly used cross-domain image recognition dataset, consisting of 4,652 images and 31 categories. Its images come from three different domains, namely Amazon (abbreviated as A, pictures downloaded from the Amazon website), Webcam (abbreviated as W, low-resolution images taken from a webcam), and DSLR (abbreviated as D, high-resolution images taken from a digital single-lens reflex camera);

[0110] Caltech-256 is also a standard image recognition dataset, containing 30,607 images in 256 categories. The Office-31+Caltech-256 combined dataset contains 4 domains, A, W, D, and C (Caltech-256). When conducting experiments, any two different domains are randomly selected to form the source domain and the target domain, and 12 different cross-domain tasks can be constructed, namely (1) A→C, (2) A→D, (3) A→W, …, (11) W→C, (12) W→D;

[0111] MSRC+VOC2007: The MSRC dataset contains 18 categories and a total of 4,323 images. The VOC2007 dataset contains 20 categories and a total of 5,011 images. 1,296 images are selected from MSRC and 1,530 images are selected from VOC2007 to form two groups of experimental data. One group uses MSRC as the source data and VOC2007 as the target data, denoted as MSRC→VOC. In the other group, the source domain and the target domain are swapped, denoted as VOC→MSRC. The selected images are converted into black-and-white images with 256 pixels, and the feature dimension is 240 dimensions;

[0112] Reuters-21578: It is a commonly used dataset for text classification, containing 5 top-level categories, namely Orgs, People, Places, Topics, and Exchanges. We select 3 major categories, Orgs, People, and Places, to construct 6 cross-domain text classification tasks, namely Orgs vs People, People vs Orgs, Orgs vs Places, Places vs Orgs, People vs Places, and Places vs People. In each group of data, the first category is the source data and the second category is the target data;

[0113] 2. Selection of Classification Models

[0114] We selected representative classification models to conduct comparative experiments on the above three groups of datasets to evaluate their classification performance;

[0115] The selected models are divided into two categories. One is the standard classifier, including SVM and Softmax regression, and the other is the transfer learning classifier, including TCA+SVM (TCA1), TCA+Softmax regression (TCA2), JDA+SVM (JDA1), JDA+Softmax regression (JDA2), UTSR, and Softmax-WMMD and Softmax-LWMMD proposed in this scheme;

[0116] In the comparative experiment, the parameters of all algorithms are set to the default values or the recommended values mentioned in the original text. For typical hyperparameter settings, the penalty term parameter C of SVM takes values {0.1, 0.5, 1, 5, 10, 50, 100}, and the balance parameter λ of the sparse constraint term in Softmax regression takes values [10 -4 , 10 -3 , in the feature transfer algorithm, the dimension k of the feature subspace takes values [20, 100]. Refer to Appendix Figure 4 、 5 、6, which is a schematic diagram of the relationship between the parameters α, β, λ, μ and the maximum number of iterations and the accuracy of the Softmax-LWMMD model;

[0117] First, we conducted sensitivity experiments and analyses to verify that the Softmax-LWMMD model can achieve the best performance under the balance parameters α, β, λ, μ and the maximum number of iterations. Four classification tasks were selected from the three groups of datasets, namely (5)C→D, (7)D→A, MSRC→VOC, and People vs Place. The results are as shown in Appendix Figure 4 、 5 、6:

[0118] (1) Figure 4 In it, for the task People vs Place, when the value of α is less than or equal to 1, the classification accuracy of the model always remains the best. When the value of α is less than 1, the proportion of the WMMD term increases and the classification accuracy of the model decreases. For the other three tasks, when the value of α is in a specific interval, the classification performance of the model reaches the optimal.

[0119] (2) When the value of β is relatively small, the change in the classification accuracy of the model on the four tasks is relatively small, but it does not reach the best. When β takes a suitable value, the classification accuracy of the model reaches the best. As β increases, when reaching the critical point, the performance of the model decreases rapidly, indicating that the proportion of the introduced JDA term needs to be within a reasonable range, as Figure 4 shown.

[0120] (3) Figure 5 Among them, the parameters λ, μ are similar to β. Whether the proportions of the Softmax regularization term and the LDA term are too large or too small will affect the classification accuracy of the model.

[0121] (4) As the number of iterations increases, in the classification task, the model accuracy will continuously improve. When the number of iterations exceeds 14 times, the model performance reaches the best and tends to be stable, indicating that it has good convergence. As Figure 6 shown.

[0122] To sum up, select the maximum number of iterations T max = 50. For other parameters α, β, λ, μ, the selected values are different according to different classification data sets, and the values of parameters α, β, λ, μ are selected accordingly according to different classification data sets during the experiment.

[0123] 3. Analysis of Experimental Results

[0124]

[0125]

[0126] Table 3 Accuracy of Different Classification Algorithms on Image Data Sets

[0127]

[0128] Table 4 Accuracy of Different Classification Algorithms on Text Data Sets According to the experimental results, we can obtain the following three conclusions:

[0129] First of all, we observe that on the three experimental data sets, the Softmax-LWMMD model achieves better results than other classification models. The average classification accuracy of Softmax-LWMMD is 53.24% and 51.11% on the two image data sets respectively, and 76.11% on the text data set, which is 1.7%, 1.42% and 2.28% higher than the reference model UTSR respectively. This verifies that Softmax-LWMMD can build a more effective performance for cross-domain classification tasks;

[0130] Secondly, the classification effect of the Softmax-LWMMD model is better than that of the Softmax-WMMD model. This is because the Softmax-LWMMD model adds the LDA algorithm. After the data is reduced in dimension by LDA, it has the smallest within-class variance and the largest between-class variance, further improving the classification performance of the Softmax-LWMMD model;

[0131] This solution is applied to the classification of image data and text data in image classification tasks and text classification tasks:

[0132] In the image classification task, select the dataset W→D, and set the parameters as α = 10 -2 , β = 10 -3 , λ = 10 -3 , μ = 10 -6 , when the number of iterations is greater than or equal to 50, the system classification accuracy is greater than or equal to 88.55;

[0133] In the text classification task, select the dataset people vs orgs, and set the parameters as α = 2, β = 3, λ = 3, μ = 1. When the number of iterations is greater than or equal to 50, the system classification accuracy is greater than or equal to 85.42;

[0134] So that this solution has a higher classification accuracy compared with other methods when applied to the image classification task.

[0135] The above is only for the purpose of illustrating the present invention. It should be understood that the present invention is not limited to the above embodiments, and various equivalent forms that conform to the idea of the present invention are within the protection scope of the present invention.

Claims

1. A data classification method with the ability to discriminate knowledge transfer, characterized in that, it includes a source domain, a target domain, a feature space, a Softmax classifier, and a prediction result. The source domain and the target domain are image data sets or text data sets. The classification process includes the following steps: S1: After the source domain samples and the target domain samples are feature-mapped, their approximate same distribution can be found in the feature space; S2: Perform WMMD and LDA operations in the feature space to distinguish samples with the same features from samples with different features, and different sample labels correspond to different sample features; Construct a system objective function: J(θ) = J 1 (θ) + α·J 2 (θ) + β·J 3 (θ) + λ·J 4 (θ) + μ·J 5 (θ)(1) Where, the θ is all the system item parameters, that is, the parameters that the system needs to determine through training; J 1 (θ) is the cost function of Softmax regression; J 2 (θ) is based on the WMMD calculation result; J 3 (θ) is used to reduce the distribution difference between domains in the Softmax regression probability output layer through conditional distribution and marginal distribution; J 4 (θ) is a sparse control term for model parameters, preventing the model from overfitting; J 5 (θ) is the LDA calculation term in the system; α, β, λ, μ are the balance parameters of J 2 (θ), J 3 (θ), J 4 (θ), J 5 (θ); L 0 and W 0 are the edge distribution matrix and the edge distribution weight coefficient matrix, respectively; L c and W c are the conditional distribution matrix and the conditional distribution weight coefficient matrix, respectively; G b is the between-class matrix, G u is the within-class matrix, G v is the total scatter matrix, G v = G b + G u . S3: The Softmax classifier calculates the probability of each sample in the target domain under each label respectively, and selects the label corresponding to the maximum probability value as the label of the sample to complete the sample classification; S4: The prediction result saves the classification result of the target domain samples, that is, the label corresponding to each sample.

2. A data classification method with the ability to discriminate knowledge transfer according to claim 1, characterized in that, In the above S1, the source domain consists of two parts, the sample x and the sample label y x , the source domain is denoted by . In the target domain, there are only samples, denoted by . The feature space includes WMMD and LDA.

3. A data classification method with the ability to discriminate knowledge transfer according to claim 1, characterized in that, The said L 0 matrix is calculated by the following formula: Among them, is the source domain, is the target domain, x is the sample in the corresponding domain, and n is the number of samples in the corresponding domain; the weight coefficient W 0 is defined as follows: Among them, is the weight coefficient of the i-th sample, is the mean of the source domain or target domain samples; L c matrix is calculated according to the following formula: Among them, is the source domain sample set belonging to class c, is the correct label of the source domain sample x i , correspondingly, is the target domain sample set belonging to class c; the class weight coefficient W c is defined as follows: wherein, is the weight coefficient of the i-th sample belonging to class c, is the mean value of the samples belonging to class c in the source domain or the target domain; The between-class matrix is calculated as follows: Here represents the mean of all samples; The within-class matrix is calculated as follows: Here represents the mean of the i-th class of samples, represents the data subset of the i-th class of samples in the total samples; From formula (11) and formula (12), the total divergence matrix is obtained:

4. A data classification method with the ability to discriminate knowledge transfer according to claim 3, characterized in that, Minimize the objective function J(θ) to obtain the optimal solution of the parameter θ. The objective function J(θ) is partially derived with respect to the parameter θ belonging to the j-th class, and the result is: Where, The model parameters are updated by the following formula, After obtaining the model parameter θ, for the target domain is classified, and the probability that each test sample x belongs to class j is: Calculate the probability that x belongs to all classes through formula (21), and the class corresponding to the maximum probability is the classification class of x.