Deep dictionary learning method based on intra-class weighting and class sharing
By introducing a mechanism of in-class weighting and class sharing in deep dictionary learning, combining specific class and class sharing dictionary learning, the problems of insufficient feature discrimination ability and insufficient class sharing feature mining in deep dictionary learning technology are solved, and higher image recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510194328.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing deep dictionary learning technology has problems in image recognition tasks such as insufficient feature identification capabilities, insufficient class sharing feature mining, and insufficient optimization strategies, resulting in insufficient classification accuracy and efficiency.
A deep dictionary learning method based on in-class weighting and class sharing is adopted. By designing an in-class weighting and inter-class weighting mechanism, combining specific class and class sharing dictionary learning, specific class features and common pattern features of the sample are extracted, and the discrimination ability of feature encoding is enhanced.
It improves the accuracy and robustness of image recognition, enhances the expression ability and classification efficiency of the dictionary, and solves the problems of insufficient feature identification ability and insufficient mining of class-sharing feature.
Smart Images

Figure CN120088561A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly relates to a deep dictionary learning method based on intra-class weighting and class sharing. Background Art
[0002] With the rapid development of artificial intelligence technology, computer vision has been increasingly widely applied in fields such as intelligent monitoring, autonomous driving, and medical diagnosis. In recent years, deep dictionary learning technology has shown remarkable effects in computer vision fields such as image representation, image restoration, denoising, and classification. As a branch of machine learning, deep dictionary learning combines the advantages of deep learning and dictionary learning (DL), aiming to extract data features by constructing a deeper dictionary learning architecture. The key to deep dictionary learning is that it expands the single-layer dictionary learning problem into a multi-layer problem, that is, using the representation of the previous layer as the input for the next layer's learning, so as to explore the deep structure and features in the data. This method is not only a new deep learning network structure, but also can be regarded as a unified framework combining dictionary learning and deep learning, with the effect of enhancing image recognition ability.
[0003] Regarding the deep dictionary learning technology, the patent document with the publication number CN115082727A discloses a scene classification method and system based on multi-layer local perception deep dictionary learning. This technology improves the deep dictionary learning method. In the first-layer dictionary learning, the sample features extracted by the PCA method are used as the input; in the dictionary learning from the second layer to the last layer, the features after passing through the ReLU activation function in the previous layer are used as the input for this layer's dictionary learning layer. An activation function is added to the deep dictionary learning to further ensure the sparsity and effectiveness of the obtained feature encoding, thereby obtaining a deeper dictionary with stronger discriminative ability and improving the model's performance in the scene classification task; through this method, it is possible to effectively extract the more essential features of the samples and obtain better scene classification ability. However, after careful analysis, it is found that this technology still has the following technical problems:
[0004] 1. Insufficient feature discrimination ability: In visual tasks, in addition to specific class features, different objects usually also contain the same common pattern features, which has been proven in the low-rank shared dictionary learning of shallow dictionary learning methods. However, this technology does not consider the highly discriminative characteristics of intra-class features and inter-class features, and has limitations in feature extraction, resulting in difficulty in further improving the classification accuracy of the model when facing complex image recognition tasks and being unable to better distinguish images of different categories.
[0005] 2. Insufficient mining of class - shared features: This technology fails to effectively mine the common pattern features among different images. As a result, when the model processes multiple categories with similar features, it cannot fully utilize these common pattern features to improve the classification efficiency and accuracy, restricting the model's expressive ability and robustness.
[0006] 3. Insufficient optimization strategy: This technology uses the PCA method to extract sample features during model learning. Specifically, it needs to first extract PCA features and then solve the feature encoding and dictionary for each input data one by one. For a dataset with a large amount of data, there will be technical problems of low training efficiency and slow convergence speed.
[0007] Therefore, it is necessary to provide a new technology to solve the above - mentioned technical problems. Summary of the Invention
[0008] The purpose of the present invention is to overcome the above - mentioned problems existing in the prior art and provide a deep dictionary learning method based on intra - class weighting and class sharing. By designing an intra - class weighting and inter - class weighting mechanism and combining specific - class and class - shared dictionary learning, this method can better extract the specific - class features and common pattern features of samples, thereby enhancing the discriminative ability of feature encoding. It can not only improve the classification efficiency, image recognition accuracy and robustness, but also solve the technical problems that the classification accuracy and classification efficiency are still not high enough due to insufficient feature discriminative ability and insufficient mining of class - shared features in the prior art.
[0009] To achieve the above purpose, the technical scheme adopted by the present invention is as follows:
[0010] A deep dictionary learning method based on intra - class weighting and class sharing, which includes the following steps:
[0011] Step S1, pre - processing and feature extraction of sample pictures
[0012] Obtain the sample data of the sample pictures, perform cropping, scaling, grayscale conversion, normalization and data augmentation on the sample data, divide the processed sample data into training samples and test samples, and extract spatial pyramid features from the training samples and test samples respectively.
[0013] Step S2, constructing a deep intra - class weighted and class - shared dictionary learning model
[0014] Introduce a discriminative reconstruction fidelity constraint term based on the reconstruction error of the input data on the specific - class dictionary and the class - shared dictionary;
[0015] Introduce a sparse coefficient regularization constraint term based on the sparse coefficients in the l 2 - norm;
[0016] Introduce the intra-class weighting factor and the inter-class weighting factor to establish the weighted Fisher discriminant constraint term;
[0017] Introduce nuclear norm regularization to establish the low-rank shared dictionary constraint term;
[0018] Construct a deep intra-class weighted and class-shared dictionary learning model according to the discriminant reconstruction fidelity constraint term, the sparse coefficient regularization constraint term, the weighted Fisher discriminant constraint term, and the low-rank shared dictionary constraint term.
[0019] Step S3: Train the model
[0020] Input the spatial pyramid features extracted from the training samples into the deep intra-class weighted and class-shared dictionary learning model for learning, and for the specific class dictionaries Specific class feature encoding Class-shared dictionary and class-shared feature encoding perform optimization and update to obtain the learned specific class dictionaries and class-shared dictionaries, as well as the feature encodings of the training samples learned at each dictionary learning layer.
[0021] Step S4: Test the model
[0022] Establish a deep dictionary learning test model based on the learned specific class dictionaries and class-shared dictionaries in step S3, input the spatial pyramid features extracted from the test samples into the deep dictionary learning test model for learning, and obtain the feature encodings of the test samples on the learned specific class dictionaries and class-shared dictionaries at each dictionary learning layer.
[0023] Step S5: Train the classifier and output the results
[0024] Connect the feature encodings obtained in step S3 and input them into the classifier to complete the training of the classifier; connect the feature encodings obtained in step S4 and input them into the trained classifier, output the test sample labels and compare them with the true labels of the test samples. When the loss function or the number of iterations of the deep intra-class weighted and class-shared dictionary learning model reaches the preset value, it indicates that the learning is completed; otherwise, repeat steps S3 - S5 until the loss function reaches the preset value or the number of iterations reaches the preset value.
[0025] In step S2, the discriminant reconstruction fidelity constraint term is used to ensure that the input data at the l-th layer can be well reconstructed and represented using the specific class dictionaries, the class-shared dictionaries, and their corresponding feature encodings. The established discriminant reconstruction fidelity constraint term is as follows:
[0026]
[0027] In the formula, Res(D l ,Zl ) represents the discriminative reconstruction fidelity constraint term; C represents the number of input sample categories; represents the input data of the l-th layer, which is obtained from the output of the feature encoding learned by the (l - 1)-th layer after passing through the activation function, that is represents the activation function layer (such as ReLU); is the input data of the k-th class. When l = 1, represents n training samples of a total of C categories with dimension d; represents the training sample of the c-th class, where The constraint term means that it is expected that the input data of the k-th class is mainly represented by the specific class dictionary of the k-th class, the feature encoding obtained from the input data on it, the class-shared dictionary, and the feature encoding obtained from the input data on it, rather than the specific class dictionary of the j-th (j ≠ k) class; represents the feature encoding of the input data of the k-th class on the specific class dictionary of the j-th class; represents the feature encoding matrix learned by the input data of the k-th class on the class-shared dictionary.
[0028] In step S2, the weighted Fisher discriminant constraint term is expected to aggregate the within-class sample representations and disperse the between-class sample representations, avoiding the inclusion of specific class features in the shared features. The weighted Fisher discriminant constraint term established according to the introduced within-class weighting factor and between-class weighting factor is as follows:
[0029]
[0030] Among them,
[0031]
[0032] In the formula, Dis(Z l ) represents the weighted Fisher discriminant constraint term; represents the within-class mean matrix of the feature encoding learned by the input data of the c-th class on the specific class dictionary, and each column vector of it is represents the between-class mean matrix of the specific class feature encoding, and each column vector of it is represents the mean matrix of the feature encoding of the input data on the class-shared dictionary, and each column of it is w k represents the within-class weighting factor; b k represents the between-class weighting factor; σ is a positive constant; represents the within-class mean matrix of the k-th class training sample X k of which each column vector is represents the between-class mean matrix of the training sample X, and each column vector of it is
[0033] In step S2, the established deep intra-class weighted and class-shared dictionary learning model is specifically as follows:
[0034]
[0035] In the formula, F(D l , Z l ) represents the deep intra-class weighted and class-shared dictionary learning model, and α, β, and γ are all hyperparameters; represents the total dictionary of the l-th layer, which is composed of a group of specific class dictionaries and the class-shared dictionary ; represents the feature encoding matrix of the input data on the total dictionary D l ; represents the specific class feature encoding learned in the l-th layer; represents the feature encoding of the input data of the c-th class on the specific class dictionary;
[0036] represents the feature encoding of the input data on the class-shared dictionary; is the nuclear norm, which represents the sum of the singular values of the matrix , and is used to ensure that the class-shared dictionary satisfies the low-rank space constraint.
[0037] In step S3, the spatial pyramid features extracted from the training samples are used as the input data for the first-layer dictionary learning. In the dictionary learning from the second layer to the last layer, the feature encoding after passing through the ReLU activation function in the previous-layer dictionary learning is used as the input data for this layer of dictionary learning.
[0038] In step S3, the ADMM method is used to optimize and update the specific class dictionary the specific class feature encoding the class-shared dictionary and the class-shared feature encoding .
[0039] In step S3, the deep intra-class weighted and class-shared dictionary learning model adopts a three-layer deep dictionary learning structure, which specifically includes the following steps:
[0040] Step S3.1, perform the first-layer dictionary learning
[0041] Input the spatial pyramid features extracted from the training samples into the following dictionary learning model to update the specific class dictionary, the class-shared dictionary, and the feature encoding:
[0042]
[0043]
[0044] In the formula, D 1 represents the dictionary learned in the first layer, and Z 1 represents the feature encoding of the input data on the dictionary D 1 .
[0045] Step 3.1.1: Initialize the specific-class dictionary Specific-class feature encoding Class-shared dictionary and class-shared feature encoding
[0046] For the specific-class dictionary Randomly select data from each class of input data as the initial dictionary atoms;
[0047] For the specific-class feature encoding Use the following model for initialization:
[0048]
[0049] For the class-shared dictionary According to the model Randomly select from the matrix input data as the initial dictionary atoms;
[0050] For the class-shared feature encoding Adopt the same initialization model as the specific-class feature encoding to complete the parameter initialization, replacing the input data X with The specific initialization model is as follows:
[0051]
[0052] After completing the initialization of all variables, alternately optimize the specific-class dictionary Specific-class feature encoding Class-shared dictionary and class-shared feature encoding
[0053] Step 3.1.2, update the specific-class dictionary At this time, the remaining variables in the deep intra-class weighted and class-shared dictionary learning model are all in a fixed state, and replacing can obtain Use the LRSDL algorithm to solve the specific-class dictionary Finally obtained The calculation is as follows:
[0054]
[0055] Among them, the matrix is shown separately as follows:
[0056]
[0057] Step 3.1.3, update the specific class feature encoding For each specific class feature encoding introduce different within-class weighting factors w k , w k is calculated based on the original input data of different classes; therefore, to calculate the gradient solution of the specific class feature encoding , first, the deep within-class weighting and class-shared dictionary learning model needs to be rewritten in the form of a single class, that is The remaining variables that have nothing to do with the solution of the current gradient variable are defined as constants constant; the rewritten model is as follows:
[0058]
[0059] Calculate with respect to , and set its gradient to 0, that is, set to obtain the optimal solution of the k-th specific class feature encoding in the deep within-class weighting and class-shared dictionary learning model
[0060] With respect to The specific solution process is shown as follows:
[0061]
[0062] Among them is shown in the following formula. After calculating the gradient solution of each class, it is finally expressed as:
[0063]
[0064]
[0065] In the formula, k - 1 and C - k represent the number of 0 elements;
[0066] Step 3.1.4, update the class-shared dictionary Update the class-shared dictionary When are all in a fixed state, the deep within-class weighting and class-shared dictionary learning model can be rewritten as:
[0067]
[0068] Among them, G 1 and H 1 are defined as follows:
[0069]
[0070] Step 3.1.5, update the class-shared feature encoding When updating the class-shared feature encoding at that time,
[0071] are all in a fixed state, and the deep in-class weighted and class-shared dictionary learning model can be rewritten as:
[0072]
[0073] Among them, Regarding the gradient of is calculated as follows:
[0074]
[0075] Finally, let can obtain as shown below:
[0076]
[0077] Step S3.2, perform the second-layer dictionary learning
[0078] Take the output of the first-layer dictionary learning as the input data for the second-layer dictionary learning, and use it for the optimization and update of the second-layer dictionary and feature encoding. The specific process is the same as that of the first-layer dictionary learning layer;
[0079] Step S3.3, perform the third-layer dictionary learning
[0080] Take the output of the second-layer dictionary learning as the input data for the third-layer dictionary learning, and use it for the optimization and update of the third-layer dictionary and feature encoding. The specific process is the same as that of the second-layer dictionary learning layer.
[0081] In step S4, the established deep dictionary learning test model is:
[0082]
[0083] In the formula, represents the deep dictionary learning test model, and β represents the hyperparameter; represents the test input data of the l-th layer. When l = 1, Otherwise represents the output after the feature encoding obtained from the upper-layer dictionary learning of the test input data passes through the activation function, that is represents the feature encoding learned by the test data on the overall dictionary; D l is represented by a specific class dictionary and a class-shared dictionary.
[0084] In step S1, the sample data of the sample pictures is the Caltech-101 image dataset. This image dataset contains 101 object categories and one background category image. The number of images in each category ranges from 31 to 800, and there are a total of 9,144 color images; the images include multiple categories such as animals, vehicles, and flowers. 30 samples are randomly selected from the images in each category as training samples, and the remaining samples are used as test samples.
[0085] The advantages of adopting the present invention are as follows:
[0086] 1. The present invention constructs a deep intra-class weighted and class-shared dictionary learning model through the established discriminative reconstruction fidelity constraint term, sparse coefficient regularization constraint term, weighted Fisher discriminant constraint term, and low-rank shared dictionary constraint term. Among them, the discriminative reconstruction fidelity constraint term introduces the reconstruction error based on the input data and the input data on the specific class dictionary and class-shared dictionary, which is beneficial to ensuring that the learned specific class dictionary and shared dictionary can optimally represent the input data. The sparse coefficient regularization constraint term introduces the sparse coefficient based on the l 2 -norm, which is beneficial to preventing the model from overfitting and improving the optimization efficiency. The weighted Fisher discriminant constraint term introduces the intra-class weighting and inter-class weighting mechanisms, which can better reflect the geometric distribution of samples in the same class in the original space and the data distribution difference between different class samples, facilitating better extraction of the specific class features and shared features of the samples. This weighting mechanism helps to optimize the sparse coding, making the model pay more attention to the clustering of intra-class samples and the discreteness of inter-class samples during the learning process, thus being beneficial to enhancing the discriminative ability of the feature encoding. The shared dictionary needs to ensure low rankness. The low-rank shared dictionary constraint term is beneficial to preventing the shared dictionary from absorbing discriminative dictionary atoms by introducing nuclear norm regularization.
[0087] Generally speaking, based on deep dictionary learning, the present invention combines the learning of specific class dictionaries and class-shared dictionaries, not only learning the specific class dictionaries of each category, but also learning a class-shared dictionary to represent the common pattern features between different objects. This combination method enables the model to simultaneously extract the specific class features and common pattern features of the samples, and make full use of these common pattern features to improve the classification efficiency and accuracy, enhance the expression ability and robustness of the dictionary, make the feature representation more comprehensive and abstract, and contribute to improving the accuracy and robustness of image recognition.
[0088] 2. The present invention uses the ADMM method to optimize and update the dictionary and feature encoding, which can efficiently solve the parameters in the model and is beneficial to improving the training efficiency and convergence speed of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 is a flowchart during the model training of the present invention;
[0090] Figure 2 is a flowchart during the model testing of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0091] As Figure 1 , 2 shown, the present invention provides a deep dictionary learning method based on intra-class weighting and class sharing, which includes the following steps:
[0092] Step S1, preprocessing and feature extraction of sample images
[0093] Obtain the sample data of the sample images, perform cropping, scaling, grayscaling, normalization, and data augmentation on the sample data, divide the processed sample data into training samples and test samples, and extract spatial pyramid (SPF) features from the training samples and test samples respectively.
[0094] The sample data of the above sample images is the Caltech-101 image dataset, which contains 101 object categories and one background category image. The number of images in each category ranges from 31 to 800, with a total of 9144 color images; the images include multiple categories such as animals, vehicles, and flowers. Randomly select 30 samples from the images in each category as training samples, and the remaining samples as test samples.
[0095] Step S2, constructing a deep intra-class weighted and class sharing dictionary learning model
[0096] To ensure that the learned specific class dictionary and class sharing dictionary can optimally represent the input data, the present invention introduces a discriminative reconstruction fidelity constraint term based on the reconstruction error of the input data on the specific class dictionary and class sharing dictionary. This discriminative reconstruction fidelity constraint term is used to ensure that the input data at the l-th layer can be well reconstructed and represented using the specific class dictionary, class sharing dictionary, and their corresponding feature encodings. The established discriminative reconstruction fidelity constraint term is as follows:
[0097]
[0098] In the formula, Res(D l , Z l ) represents the discriminative reconstruction fidelity constraint term; C represents the number of input sample categories; represents the input data of the l-th layer, which is obtained from the output of the feature encoding learned by the (l - 1)-th layer after passing through the activation function, that is represents the activation function layer (such as ReLU); is the input data of the k-th class. When l = 1, represents n training samples of a total of C classes with a dimension of d; represents the training sample of the c-th class, where The constraint term indicates that it is expected that the input data of the k-th class is mainly represented by the specific class dictionary of the k-th class, the feature encoding obtained from the input data on it, the class-shared dictionary, and the feature encoding obtained from the input data on it, rather than the specific class dictionary of the j-th (j ≠ k) class; represents the feature encoding of the input data of the k-th class on the specific class dictionary of the j-th class; represents the feature encoding matrix learned by the input data of the k-th class on the class-shared dictionary.
[0099] To prevent the model from overfitting and improve the optimization efficiency, the present invention establishes a sparse coefficient regularization constraint term by introducing a sparse coefficient based on the l 2 -norm.
[0100] To better reflect the geometric distribution of samples of the same class in the original space and the data distribution difference between samples of different classes, and to facilitate better extraction of specific class features and shared features of samples, the present invention establishes a weighted Fisher discriminant constraint term by introducing an intra-class weighting factor and an inter-class weighting factor. The weighted Fisher discriminant constraint term is expected to aggregate the representation of intra-class samples and discrete the representation of inter-class samples, and avoid including specific class features in the shared features. The weighted Fisher discriminant constraint term established according to the introduced intra-class weighting factor and inter-class weighting factor is as follows:
[0101]
[0102] where,
[0103]
[0104] In the formula, Dis(Z l ) represents the weighted Fisher discriminant constraint term; represents the intra-class mean matrix of the feature encoding learned by the input data of the c-th class on the specific class dictionary, and each column vector of it is represents the inter-class mean matrix of the specific class feature encoding, and each column vector of it is represents the mean matrix of the feature encoding of the input data on the class-shared dictionary, and each column of it is w kdenotes the within-class weighting factor; b k denotes the between-class weighting factor; σ is a positive constant; denotes the within-class mean matrix of the training samples X k of the k-th class, and each column vector of it is denotes the between-class mean matrix of the training samples X, and each column vector of it is
[0105] The shared dictionary needs to ensure low rank. To prevent the shared dictionary from absorbing discriminative dictionary atoms, the present invention establishes a low-rank shared dictionary constraint term by introducing nuclear norm regularization.
[0106] Construct a deep within-class weighted and class-shared dictionary learning model according to the above discriminative reconstruction fidelity constraint term, sparse coefficient regularization constraint term, weighted Fisher discriminative constraint term, and low-rank shared dictionary constraint term. The established deep within-class weighted and class-shared dictionary learning model is specifically as follows:
[0107]
[0108] In the formula, F(D l , Z l ) represents the deep within-class weighted and class-shared dictionary learning model, and α, β, and γ are all hyperparameters; denotes the total dictionary of the l-th layer, which consists of a group of specific class dictionaries and the class-shared dictionary ; denotes the feature encoding matrix of the input data on the total dictionary D l ; denotes the specific class feature encoding learned in the l-th layer; denotes the feature encoding of the input data of the c-th class on the specific class dictionary;
[0109] denotes the feature encoding of the input data on the class-shared dictionary; is the nuclear norm, which represents the sum of the singular values of the matrix , and is used to ensure that the class-shared dictionary satisfies the low-rank space constraint.
[0110] Step S3, Train the model
[0111] Input the spatial pyramid features extracted from the training samples into the deep within-class weighted and class-shared dictionary learning model for learning, and for the specific class dictionaries in the deep within-class weighted and class-shared dictionary learning model specific class feature encoding class-shared dictionary and class-shared feature encoding Optimize and update to obtain the learned specific class dictionary, class - shared dictionary, and the feature encodings of the training samples learned at each dictionary learning layer.
[0112] Specifically, the first - layer dictionary learning uses the spatial pyramid features extracted from the training samples as the input data. From the second - layer dictionary learning to the last - layer dictionary learning, the feature encodings after passing through the ReLU activation function in the previous - layer dictionary learning are used as the input data for this layer's dictionary learning layer. Additionally, in each layer of dictionary learning, the ADMM method is used to optimize and update the specific class dictionary Specific class feature encoding Class - shared dictionary and class - shared feature encoding for optimization and update.
[0113] In a specific embodiment, the deep intra - class weighted and class - shared dictionary learning model of the present invention adopts a three - layer deep dictionary learning structure, specifically including the following steps:
[0114] Step S3.1, perform the first - layer dictionary learning
[0115] Input the spatial pyramid features extracted from the training samples into the following dictionary learning model for updating the specific class dictionary, class - shared dictionary, and feature encoding:
[0116]
[0117] In the formula, D 1 represents the dictionary learned in the first layer, and Z 1 represents the feature encoding of the input data on the dictionary D 1 .
[0118] Step 3.1.1: Initialize the specific class dictionary Specific class feature encoding Class - shared dictionary and class - shared feature encoding
[0119] For the specific class dictionary Randomly select data from each class of input data as the initial dictionary atoms.
[0120] For the specific class feature encoding Use the following model for initialization:
[0121]
[0122]
[0123] For the class - shared dictionary According to the model Randomly select from the matrix several input data as the initial dictionary atoms.
[0124] For class-shared feature encoding Adopt the same initialization model as the specific-class feature encoding to complete the parameter initialization, and replace the input data X with The specific initialization model is as follows:
[0125]
[0126] After completing the initialization of all variables, the specific-class dictionary specific-class feature encoding class-shared dictionary and class-shared feature encoding
[0127] Step 3.1.2, update the specific-class dictionary At this time, the remaining variables in the deep within-class weighted and class-shared dictionary learning model are all in a fixed state, and replace to obtain Use the LRSDL algorithm to solve the specific-class dictionary The finally obtained The calculation is as follows:
[0128]
[0129] Among them, the matrix is shown respectively as follows:
[0130]
[0131] Step 3.1.3, update the specific-class feature encoding Introduce different within-class weighted factors w for each specific-class feature encoding k , w k is calculated according to the original input data of different classes; therefore, to calculate the gradient solution of the specific-class feature encoding , first rewrite the deep within-class weighted and class-shared dictionary learning model into the form of a single class, that is The remaining variables irrelevant to the current gradient variable are all defined as constants constant; the rewritten model is as follows:
[0132]
[0133]
[0134] Calculate the gradient with respect to , and set its gradient to 0, i.e., set to obtain the optimal solution of the k-th class-specific feature encoding in the deep within-class weighted and class-shared dictionary learning model
[0135] With respect to The specific solution process is shown as follows:
[0136]
[0137] where is shown in the following formula. After calculating the gradient solution for each class, it finally represents:
[0138]
[0139] In the formula, k - 1 and C - k represent the number of 0 elements.
[0140] Step 3.1.4, update the class-shared dictionary Update the class-shared dictionary When are all in a fixed state, the deep within-class weighted and class-shared dictionary learning model can be rewritten as:
[0141]
[0142] where G 1 and H 1 are defined as follows:
[0143]
[0144] Step 3.1.5, update the class-shared feature encoding When updating the class-shared feature encoding When
[0145] are all in a fixed state, the deep within-class weighted and class-shared dictionary learning model can be rewritten as:
[0146]
[0147] where The gradient with respect to is calculated as follows:
[0148]
[0149] Finally, set to obtain As shown below:
[0150]
[0151] Step S3.2, perform the second-layer dictionary learning
[0152] The learning process of the second-layer dictionary is the same as that of the first-layer dictionary learning layer. The specific process is as follows:
[0153] Take the output of the first-layer dictionary learning as the input data for the second-layer dictionary learning, and input it into the following dictionary learning model for the optimization and update of the second-layer dictionary and feature encoding:
[0154]
[0155] In the formula, D 2 represents the dictionary learned in the second layer, and Z 2 represents the feature encoding of the input data on the dictionary D 2
[0156] coding.
[0157] Step 3.2.1: Initialization For Randomly select data from each class of input data as the initial dictionary atoms.
[0158] For Use the following model for initialization:
[0159]
[0160] For According to the model Randomly select from the matrix input data as the initial dictionary atoms.
[0161] For Adopt the same initialization model as to complete the parameter initialization, and replace the input data with The specific initialization model is as follows:
[0162]
[0163] After completing the initialization of all variables, the alternating optimization can be performed
[0164] Step 3.2.2, update the specific-class dictionary At this time, the remaining variables in the depth within-class weighted and class-shared dictionary learning model are all in a fixed state and are replaced to obtain The efficient algorithm proposed in LRSDL can be used to solve it, and the finally obtained is calculated as follows:
[0165]
[0166] where the matrix is shown separately as follows:
[0167]
[0168] Step 3.2.3, update the specific class feature encoding In the present invention, for each specific class feature encoding different within-class weighting factors w k , w k are calculated according to the original input data of different classes; therefore, to calculate the gradient solution of the specific class feature encoding , first, the depth within-class weighted and class-shared dictionary learning model needs to be rewritten in the form of a single class, that is The remaining variables irrelevant to the solution of the current gradient variable are defined as constants constant; the rewritten model is as follows:
[0169]
[0170] Calculate with respect to and set its gradient to 0, that is, set to obtain the optimal solution of the specific class encoding coefficient of the k-th class in the depth within-class weighted and class-shared dictionary learning model
[0171] with respect to The specific solution process is shown as follows:
[0172]
[0173] where is shown in the following formula. After calculating the gradient solution of each class, it is finally expressed as:
[0174]
[0175]
[0176]
[0177] In the formula, k - 1 and C - k represent the number of 0 elements.
[0178] Step 3.2.4, update the class - shared dictionary During the update when are all in a fixed state, the deep intra - class weighted and class - shared dictionary learning model can be rewritten as:
[0179]
[0180] where G 2 and h 2 are defined as follows:
[0181]
[0182]
[0183] Step 3.2.5, update the class - shared feature encoding During the update when are all in a fixed state, the deep intra - class weighted and class - shared dictionary learning model can be rewritten as:
[0184]
[0185] where, The gradient with respect to is calculated as follows:
[0186]
[0187] Finally, let We can obtain as follows:
[0188]
[0189] Step S3.3, perform the third - layer dictionary learning
[0190] Take the output of the second - layer dictionary learning as the input data for the third - layer dictionary learning, which is used to optimize and update the third - layer dictionary and feature encoding. The learning process of the third - layer dictionary is the same as that of the second - layer dictionary learning layer and will not be elaborated here.
[0191] Step S4, test the model
[0192] Establish a deep dictionary learning test model according to the specific - class dictionary and class - shared dictionary learned in Step S3. The established deep dictionary learning test model is:
[0193]
[0194] In the formula, represents the deep dictionary learning test model, and β represents the hyperparameter; represents the l-th layer of test input data. When l = 1, otherwise represents the output after the feature encoding obtained from the previous layer of dictionary learning of the test input data passes through the activation function, that is, represents the feature encoding learned by the test data on the overall dictionary; D l represents the composition of the specific class dictionary and the class-shared dictionary.
[0195] It should be noted that the deep dictionary learning test model is a strictly convex function and can be simply obtained by the gradient descent algorithm The specific display is as follows:
[0196]
[0197] After obtaining the deep dictionary learning test model, the spatial pyramid features extracted from the test samples are input into the deep dictionary learning test model for learning, and the feature encodings of the test samples on the specific class dictionary and the class-shared dictionary that have been learned in each dictionary learning layer are obtained.
[0198] Step S5, Training the classifier and outputting the results
[0199] The feature encodings obtained in step S3 are concatenated to form the final feature encoding for training the classifier, that is, the training sample feature encoding The training sample feature encoding Z ‰‘ is input into the classifier to complete the training of the classifier.
[0200] The feature encodings obtained in step S4 are concatenated to form the final feature encoding for testing the classifier, that is, the test sample feature encoding The test sample feature encoding Z te is input into the trained classifier, and the test sample label is output.
[0201] The test sample label is compared with the true label of the test sample. When the loss function or the number of iterations of the deep within-class weighted and class-shared dictionary learning model reaches the preset value, it indicates that the learning is completed; otherwise, steps S3 - S5 are repeated until the loss function reaches the preset value or the number of iterations reaches the preset value.
[0202] The patent document with the publication number CN115082727A is used as Comparative Example 1 below. The traditional shallow shared dictionary learning (LRSDL) and deep dictionary learning (DDL) methods are used as Comparative Examples 2 and 3 respectively. The Caltech-101 image dataset in step S1 is used as the experimental data to test the present invention and Comparative Examples 1-3 respectively. It is found through experimental verification that:
[0203] 79.54%, 80.29% and 78.98% of Comparative Examples 1-3 on the Caltech-101 dataset.
[0204] However, due to the combination of the specific class dictionary and class-shared dictionary learning in the present invention, not only the specific class dictionary of each class is learned, but also a class-shared dictionary is learned to represent the common pattern features between different objects. This combination method enables the model to simultaneously extract the specific class features and common pattern features of the samples, and make full use of these common pattern features to improve the classification efficiency and accuracy. Therefore, the accuracy of the present invention on this dataset can reach 81.21%. Compared with the prior art, the present invention significantly improves the recognition accuracy, enhances the expressive ability and robustness of the dictionary, makes the feature representation more comprehensive and abstract, helps to improve the accuracy and robustness of image recognition, and verifies the effectiveness of the present invention.
[0205] The above is only the specific implementation manner of the present invention. Any feature disclosed in this specification, unless specifically described, can be replaced by other equivalent or similar-purpose alternative features; all the features disclosed, or all the steps in any method or process, except for mutually exclusive features and / or steps, can be combined in any way.
Claims
1. A deep dictionary learning method based on intra-class weighting and class sharing, characterized in that The following steps are involved: Step S1: Sample image preprocessing and feature extraction Obtain sample data of the sample image, perform cropping, scaling, grayscale, normalization and data enhancement processing on the sample data, divide the processed sample data into training samples and test samples, and extract spatial pyramid features from the training samples and the test samples respectively; Step S2: Constructing a deep intra-class weighted and class-shared dictionary learning model Introduce a discrimination reconstruction fidelity constraint item based on input data and reconstruction error of input data on a specific class dictionary and a class shared dictionary; Introduce sparse coefficients based on l2 norm to establish sparse coefficient regularization constraints; Introduce intra-class weighting factors and inter-class weighting factors to establish weighted Fisher discriminant constraints; Nuclear norm regularization is introduced to establish low-rank shared dictionary constraints; A deep intra-class weighted and class shared dictionary learning model is constructed based on the discrimination reconstruction fidelity constraint, sparse coefficient regularization constraint, weighted Fisher discrimination constraint and low-rank shared dictionary constraint. Step S3: training model The spatial pyramid features extracted from the training samples are input into the deep intra-class weighted and class-shared dictionary learning model for learning. Class-specific feature encoding Class Shared Dictionary Shared feature encoding with class Perform optimization and update to obtain the learned specific class dictionary and class shared dictionary, as well as the feature encodings of the training samples learned in each dictionary learning layer; Step S4: Test model A deep dictionary learning test model is established based on the specific class dictionary and the class shared dictionary that have been learned in step S3, and the spatial pyramid features extracted from the test sample are input into the deep dictionary learning test model for learning, so as to obtain the feature encoding of the test sample on the specific class dictionary and the class shared dictionary that have been learned in each dictionary learning layer; Step S5: training the classifier and outputting the results The feature codes obtained in step S3 are connected and input into the classifier to complete the training of the classifier; the feature codes obtained in step S4 are connected and input into the trained classifier, the test sample label is output and compared with the true label of the test sample. When the loss function or the number of iterations of the deep intra-class weighted and class shared dictionary learning model reaches the preset value, it means that the learning is completed; otherwise, repeat steps S3-S5 until the loss function reaches the preset value or the number of iterations reaches the preset value.
2. The deep dictionary learning method based on intra-class weighting and class sharing according to claim 1, characterized in that: In step S2, the identification and reconstruction fidelity constraint items are used to ensure that the input data of the first layer can be well reconstructed and represented using the specific class dictionary and the class shared dictionary and the corresponding feature encodings. The identification and reconstruction fidelity constraint items established are as follows: In the formula, Res(D l ,Z l ) represents the identification and reconstruction fidelity constraint; C represents the number of input sample categories; represents the input data of the lth layer, which is obtained by the output of the feature encoding learned in the l-1th layer after the activation function, that is, represents the activation function layer; is the kth type of input data, when l = 1, represents a total of C categories with n training samples of dimension d; represents the c-th class training sample, where The constraint term indicates that the k-th class input data is expected to be mainly represented by the k-th class-specific dictionary and the feature encoding of the input data obtained thereon and the class-shared dictionary and the feature encoding of the input data obtained thereon, rather than the j-th (j≠k) class-specific dictionary; Represents the feature encoding of the k-th input data on the j-th class-specific dictionary; Represents the feature encoding matrix learned by the k-th class input data on the class shared dictionary.
3. The deep dictionary learning method based on intra-class weighting and class sharing according to claim 2, characterized in that: In step S2, the weighted Fisher discriminant constraint item is expected to aggregate the intra-class sample representation and discretize the inter-class sample representation to avoid the inclusion of specific class features in the shared features. The weighted Fisher discriminant constraint item established according to the introduced intra-class weighting factor and inter-class weighting factor is as follows: in, In the formula, Dis(Z l ) represents the weighted Fisher discriminant constraint; Represents the intra-class mean matrix of the feature encoding learned on the specific class dictionary for the c-th class input data, and each column vector is Represents the inter-class mean matrix of the feature encoding of a specific class, and each column vector is Represents the mean matrix of the feature encoding of the input data on the class shared dictionary, each column of which is w k represents the intra-class weighting factor; b k represents the inter-class weighting factor; σ is a positive constant; Represents the k-th class training sample X k The intra-class mean matrix, each column vector of which is Represents the inter-class mean matrix of training sample X, each column vector of which is 4. The deep dictionary learning method based on intra-class weighting and class sharing according to claim 3 is characterized by: In step S2, the deep intra-class weighted and class-shared dictionary learning model established is as follows: In the formula, F(D l ,Z l ) represents the deep intra-class weighted and class-shared dictionary learning model, α, β, γ are all hyperparameters; Represents the total dictionary of the lth layer, which consists of a set of specific class dictionaries Sharing a dictionary with a class composition; Indicates that the input data is in the total dictionary D l The feature encoding matrix on ; Represents the class-specific feature encoding learned by the lth layer; Represents the feature encoding of the c-th type of input data on a specific class dictionary; Represents the feature encoding of the input data on the class shared dictionary; is the nuclear norm, which means the matrix The sum of the singular values, used to ensure that the class shares the dictionary Satisfy the low-rank spatial constraint.
5. The deep dictionary learning method based on intra-class weighting and class sharing according to claim 4, characterized in that: In step S3, the first layer of dictionary learning uses the spatial pyramid features extracted from the training samples as input data, and in the second layer of dictionary learning to the last layer of dictionary learning, the feature encoding after the ReLU activation function in the previous layer of dictionary learning is used as the input data of the dictionary learning layer of this layer.
6. The deep dictionary learning method based on intra-class weighting and class sharing according to claim 1, characterized in that: In step S3, the ADMM method is used to classify the specific dictionary Class-specific feature encoding Class Shared Dictionary Shared feature encoding with class Perform optimization updates.
7. The deep dictionary learning method based on intra-class weighting and class sharing according to claim 5, characterized in that: In step S3, the deep intra-class weighted and class-shared dictionary learning model adopts a three-layer deep dictionary learning structure, which specifically includes the following steps: Step S3.1, perform first-layer dictionary learning The spatial pyramid features extracted from the training samples Input to the following dictionary learning model to update the class-specific dictionary, class-shared dictionary, and feature encoding: Where D 1 Represents the dictionary learned by the first layer, Z 1 Indicates that the input data is in the dictionary D 1 Feature encoding on ; Step 3.1.1: Initialize a class-specific dictionary Class-specific feature encoding Class Shared Dictionary Shared feature encoding with class For a specific class dictionary Randomly select from each class of input data data as the initial dictionary atom; For specific class feature encoding Initialize with the following model: For class shared dictionary According to the model From the matrix Random selection Input data as initial dictionary atoms; For class-shared feature encoding Using class-specific feature encoding The same initialization model is used to complete parameter initialization, replacing the input data X with The specific initialization model is as follows: After initializing all variables, you can optimize the class-specific dictionary alternately. Class-specific feature encoding Class Shared Dictionary Shared feature encoding with class Step 3.1.2, Update the class-specific dictionary At this time, the remaining variables in the deep intra-class weighted and class shared dictionary learning model All are fixed and replaced Be able to get Solve the class-specific dictionary using the LRSDL algorithm The final result The calculation is as follows: Among them, the matrix and They are shown as follows: Step 3.1.3, Update the specific class feature encoding Encode each specific class feature Introducing different intra-class weighting factors w k , w k It is calculated based on the original input data of different classes; therefore, to calculate the feature encoding of a specific class To solve the gradient problem of , we first need to rewrite the deep intra-class weighted and class-shared dictionary learning model into a single-class form, that is, The rest of the variables are the same as the current gradient All variables that are irrelevant to the solution are defined as constants; the rewritten model is as follows: calculate about The gradient of , and set its gradient to 0, that is, The optimal solution for encoding the k-th class-specific feature in the deep intra-class weighted and class-shared dictionary learning model can be obtained. about The specific solution process is shown as follows: in and Shown in the following formula, after calculating the gradient solution for each category, it is finally expressed as: In the formula, k-1 and Ck represent the number of 0 elements; Step 3.1.4, Update the class shared dictionary Update class shared dictionary hour, All are in a fixed state, and the deep intra-class weighted and class-shared dictionary learning model can be rewritten as: Among them G 1 With H 1 The definition is as follows: Step 3.1.5, Update class shared feature encoding Shared feature encoding in update class hour, All are in a fixed state, and the deep intra-class weighted and class-shared dictionary learning model can be rewritten as: in, about Gradient The calculation is as follows: Finally, let Be able to get As shown below: Step S3.2, perform second-layer dictionary learning The output of the first layer of dictionary learning As the input data of the second-layer dictionary learning, it is used for optimizing and updating the second-layer dictionary and feature encoding. The specific process is consistent with the first-layer dictionary learning layer. Step S3.3, perform third-layer dictionary learning The output of the second layer dictionary learning As the input data of the third-layer dictionary learning, it is used for optimizing and updating the third-layer dictionary and feature encoding. The specific process is consistent with the second-layer dictionary learning layer.
8. A deep dictionary learning method based on intra-class weighting and class sharing according to any one of claims 1-5, characterized in that: In step S4, the deep dictionary learning test model established is: In the formula, represents the deep dictionary learning test model, β represents the hyperparameter; Represents the test input data of the lth layer. When l=1, otherwise It represents the output of the feature encoding of the test input data obtained in the previous layer of dictionary learning after the activation function, that is, represents the feature encoding learned by the test data on the overall dictionary; D l The representation consists of a class-specific dictionary and a class-shared dictionary.
9. The deep dictionary learning method based on intra-class weighting and class sharing according to claim 1, characterized in that: In step S1, the sample data of the sample images is the Caltech-101 image dataset, which contains 101 object categories and one background category image. The number of images in each category ranges from 31 to 800, with a total of 9144 color images; the images include multiple categories such as animals, vehicles, and flowers. 30 samples are randomly selected from the images of each category as training samples, and the remaining samples are used as test samples.
Citation Information
Patent Citations
Scene classification method and system based on multilayer local perception depth dictionary learning
CN115082727A