Image classification model training method and device, equipment and storage medium
By using sparse dictionary learning and classifier to extract image features when the data set only has category proportional labels, the problem that the prior art is difficult to achieve fine-grained image instance-level classification is solved, and effective instance-level classification of images is realized.
Patent Information
- Application Number
- CN202510023917.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-07
AI Technical Summary
The prior art is difficult to effectively mine discriminant features in fine-grained images without having only category proportional labels, so as to realize instance-level classification of fine-grained images.
By obtaining the basic features and category proportional labels of the image bag, the target sparse features are extracted using the sparse dictionary learning sub-model and inputting them into the classifier to obtain the predicted category labels. The loss function value is determined based on these labels, and the model parameters are adjusted for training.
In the case where only the category proportional label is available, the image classification model is effectively mined for the discriminant features of each image to realize instance-level classification of the image.
Smart Images

Figure CN120107647A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to an image classification model training method, device, equipment and storage medium. Background Art
[0002] In real life, a lot of useful information is often released in the form of group-level statistics rather than specific individual data, such as the influenza infection rate in a certain area, the support rate of different candidates in a certain area, the distribution of education levels in a certain city, etc. This type of information plays an important role in data analysis tasks, but traditional learning methods based on individual instances are difficult to directly use such data. In addition, the success of existing algorithms is largely due to large-scale annotated data, and obtaining this data often requires very high time and labor costs; at the same time, some instance-level label information is sometimes difficult or even impossible to obtain directly due to privacy issues, and only provides a proportion of fine-grained labels.
[0003] Therefore, it is necessary to effectively mine discriminative features in fine-grained images with only category ratio labels and achieve instance-level classification of fine-grained images. Summary of the invention
[0004] The present application proposes an image classification model training method, apparatus, device and storage medium, which can effectively mine the discriminative features of the image classification model in the image when the data set only has category ratio labels, so as to realize instance-level classification of the image through the image classification model.
[0005] The first embodiment of the present application proposes an image classification model training method, comprising:
[0006] Obtain basic features of image bags in a sample set and category ratio labels of the image bags;
[0007] Inputting the basic features into a sparse dictionary learning sub-model of an image classification model to obtain target sparse features of the image bag;
[0008] Inputting the sparse features into a classifier of the image classification model to obtain a predicted category label for each image in the bag of images;
[0009] Determining a predicted category ratio label for the bag of images based on the predicted category label for each image in the bag of images;
[0010] Determining a first loss function value based on the predicted category ratio label of the bag of images and the category ratio label of the bag of images;
[0011] Based on the first loss function value, the parameters of the image classification model are adjusted, and training is continued until a preset first training completion condition is met to obtain a trained image classification model.
[0012] The embodiment of the second aspect of the present application provides an image classification model training device, including:
[0013] An acquisition module, used to acquire basic features of image bags in a sample set and category ratio labels of the image bags;
[0014] An input module, used for inputting the basic features into a sparse dictionary learning sub-model of an image classification model to obtain target sparse features of the image bag;
[0015] The input module is further used to input the sparse features into the classifier of the image classification model to obtain a predicted category label for each image in the image bag;
[0016] A determination module, configured to determine a predicted category ratio label of the image bag based on the predicted category label of each image in the image bag;
[0017] The determination module is used to determine a first loss function value based on the predicted category ratio label of the image bag and the category ratio label of the image bag;
[0018] An adjustment module is used to adjust the parameters of the image classification model based on the first loss function value, and continue training until a preset first training completion condition is met to obtain a trained image classification model.
[0019] An embodiment of the third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect above.
[0020] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium on which a computer program is stored, and the program is executed by a processor to implement the method described in the first aspect above.
[0021] The technical solution provided in the embodiments of the present application has at least the following technical effects or advantages:
[0022] The present application proposes an image classification model training method, apparatus, device and storage medium, the method comprising: obtaining basic features of image bags in a sample set and category ratio labels of image bags; inputting the basic features into a sparse dictionary learning sub-model to obtain target sparse features of the image bag; inputting the target sparse features into a classifier to obtain predicted category labels for each image in the image bag; determining predicted category ratio labels of the image bag based on the predicted category labels of each image in the image bag; determining a first loss function value based on the predicted category ratio labels and the category ratio labels; training the image classification model with the first loss function value to obtain a trained image classification model. The embodiments of the present application effectively mine the discriminative features of the image classification model for each image when the data set only has category ratio labels, so as to achieve instance-level classification of images through bag-level category ratio labels.
[0023] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] By reading the detailed description of the preferred embodiment below, various other advantages and benefits will become clear to those of ordinary skill in the art. The accompanying drawings are only used for the purpose of illustrating the preferred embodiment and are not considered to be limitations of the present application. In addition, the same reference symbols are used to represent the same components throughout the accompanying drawings.
[0025] In the attached picture:
[0026] Figure 1 A schematic diagram of a process flow of an image classification model training method provided by an embodiment of the present application is shown;
[0027] Figure 2 A schematic diagram of the structure of a sparse dictionary learning sub-model provided in an embodiment of the present application is shown;
[0028] Figure 3 A schematic diagram of the structure of an image classification model training device provided by an embodiment of the present application is shown;
[0029] Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application is shown;
[0030] Figure 5 A schematic diagram of a storage medium provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0031] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0032] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this application should have the common meanings understood by technicians in the field to which this application belongs.
[0033] The design method of the synchronization signal block of the present application can be executed by a computing device, and the computing device applies cloud computing and virtualization technology. The computing device can be a server, such as one server, multiple servers, a server cluster, a cloud computing platform, etc. Optionally, the computing device can also be a terminal device, such as a mobile phone, a tablet computer, a game console, a portable computer, a desktop computer, an advertising machine, an all-in-one machine, etc. The present application does not limit the device type and number of the computing device.
[0034] Based on the above background technology, fine-grained images and their labels usually contain a lot of details and privacy-sensitive information, while most of the existing computer vision tasks for fine-grained images are still based on instance-level fine-grained images, and the corresponding privacy protection strategies are limited to masking or encoding instance-level labels. No relevant research has considered whether the privacy protection of fine-grained images can be achieved through label ratio learning. In addition, applications with high privacy requirements may only provide the ratio of fine-grained labels instead of the specific category labels of each instance. In this case, directly using label ratio learning can not only effectively utilize the existing annotation information, but also achieve privacy protection of fine-grained information. The challenge of label ratio learning for fine-grained images is mainly how to effectively mine the discriminative features in fine-grained images with only category ratio labels and realize instance-level classification of fine-grained images. A key solution is to make full use of the bag-level structure under the label ratio learning paradigm, extract and mine features of fine-grained images by constructing a sparse dictionary, and train a fine-grained visual classification model based on label ratio learning.
[0035] In order to solve the above problems, the present application proposes an image classification model training method, device, equipment and storage medium, the method comprising: obtaining basic features of image bags in a sample set and category ratio labels of image bags; inputting the basic features into a sparse dictionary learning sub-model to obtain target sparse features of the image bag; inputting the target sparse features into a classifier to obtain predicted category labels for each image in the image bag; determining predicted category ratio labels of the image bag based on the predicted category labels of each image in the image bag; determining a first loss function value based on the predicted category ratio labels and the category ratio labels; training the image classification model with the first loss function value to obtain a trained image classification model. The embodiment of the present application effectively mines the discriminative features of the image classification model for each image when the data set only has category ratio labels, so as to achieve instance-level classification of images through bag-level category ratio labels.
[0036] The following describes an image classification model training method, device, equipment and storage medium proposed according to an embodiment of the present application in conjunction with the accompanying drawings.
[0037] See also Figure 1 , the method specifically comprises the following steps:
[0038] S101, obtaining basic features of image bags in a sample set and category ratio labels of image bags.
[0039] The sample set includes multiple image bags, and the image bag includes multiple sample images. Since applications with higher privacy requirements may only provide the proportion of fine-grained labels rather than the specific category label of each instance, the sample data set does not provide the label of each image, but provides the category ratio label of the image bag. For example, an image bag includes 10 images, and the category ratio label of the image bag is 20% for the dog category, 50% for the cat category, and 30% for the mouse category. Then the category ratio label of the image bag indicates that 2 of the 10 images in the bag belong to the dog category, 5 images belong to the cat category, and 3 images belong to the mouse category.
[0040] The image bags in the sample set can be input into a pre-trained neural network model to extract basic features of the image bags, where the neural network model can be a convolutional neural network, a VGG-19 network, a ResNet-50 network, and the like.
[0041] The bag of images can be represented as The basic features can be expressed as
[0042] Among them, B i represents the i-th input image bag, which contains |B i | instance images, represents the jth image instance in the i-th input image bag, W represents the width of the input image, H represents the height of the input image, and the corresponding basic features are expressed as
[0043] S102: Input the basic features into the sparse dictionary learning sub-model of the image classification model to obtain the target sparse features of the image bag.
[0044] Among them, the sparse dictionary learning sub-model includes sparse features composed of a dictionary and a set of sparse coefficient matrices. Generally, basic features can be approximately represented by the product of the dictionary and the sparse features. In the process of representing basic features by the product of the dictionary and the sparse features, the dictionary is optimized to better adapt to the sample set.
[0045] After the dictionary optimization is completed, the target sparse features corresponding to the basic features can be determined through the dictionary and the basic features.
[0046] Since sparse features effectively reduce the amount of data storage and transmission by representing a large amount of information in an image with as few elements as possible, this representation method can remove redundant information in an image while retaining key information. Therefore, sparse features are used in image classification to improve the efficiency of image classification.
[0047] S103: Input the sparse features into the classifier of the image classification model to obtain the predicted category label of each image in the image bag.
[0048] S104: Determine a predicted category ratio label of the image bag based on the predicted category label of each image in the image bag.
[0049] The classifier can output a predicted image category for each image in the image bag based on the sparse features of the image bag, and further can determine a predicted category ratio label of the image bag based on the image category of each image.
[0050] For example, the image bag includes 10 images, and the classifier output may be that the first image is a dog, the second to fifth images are cats, and the remaining images are mice. Therefore, based on the output results, the image bag predicts that the category ratio labels are 10% for dog, 40% for cat, and 50% for dog.
[0051] S105 . Determine a first loss function value based on the predicted category ratio label of the image bag and the category ratio label of the image bag.
[0052] S106. Based on the first loss function value, adjust the parameters of the image classification model, and continue training until a preset first training completion condition is met to obtain a trained image classification model.
[0053] Among them, the first loss function value can be determined based on the difference value between the predicted category ratio label of the image bag and the category ratio label of the image bag, and the parameters of the image classification model can be further adjusted based on the first loss function value, and training is continued until the preset first training completion condition is met to obtain a trained image classification model.
[0054] The first training completion condition may be that the number of training times reaches a first preset training times threshold or the first loss function value is less than the first loss function threshold, wherein the first preset training times threshold and the first loss function threshold can be flexibly set based on actual conditions.
[0055] The present application proposes an image classification model training method, apparatus, device and storage medium, the method comprising: obtaining basic features of image bags in a sample set and category ratio labels of image bags; inputting the basic features into a sparse dictionary learning sub-model to obtain target sparse features of the image bag; inputting the target sparse features into a classifier to obtain predicted category labels for each image in the image bag; determining predicted category ratio labels of the image bag based on the predicted category labels of each image in the image bag; determining a first loss function value based on the predicted category ratio labels and the category ratio labels; training the image classification model with the first loss function value to obtain a trained image classification model. The embodiments of the present application effectively mine the discriminative features of the image classification model for each image when the data set only has category ratio labels, so as to achieve instance-level classification of images through bag-level category ratio labels.
[0056] In some embodiments, basic deep features are input into a sparse dictionary learning submodel of an image classification model to obtain target sparse features of an image bag, including: determining a first target dictionary of the sparse dictionary learning submodel; determining first sparse features of the image bag based on basic features of the image bag and the first target dictionary; determining second sparse features of the image bag based on basic features of the image bag and a target sub-dictionary related to the label ratio of the image bag in the first target dictionary; constraining the first target dictionary based on the first sparse features of the image bag and the second sparse features of the image bag to obtain a second target dictionary; and determining target sparse features of the image bag based on the second target dictionary and the basic features of the image bag.
[0057] After optimizing the dictionary, a first target dictionary is obtained, which can better adapt to the sample set.
[0058] The dictionary can be iteratively optimized by setting an optimization target until the dictionary meets the optimization target, thereby obtaining a first target dictionary.
[0059] Multiple network layers may also be set, each network layer is provided with an update strategy, and the dictionary is continuously updated through the multiple network layers until the last network layer outputs the first target dictionary.
[0060] It can be understood that the first target dictionary includes basis functions corresponding to all categories in the sample set, while the categories corresponding to the image bag only include a part of all categories. In order to allow the sparse dictionary learning sub-model to better capture and distinguish the differences between different categories, the basis functions related to the categories corresponding to the image bag in the first target dictionary can be used to represent the basic features of the image bag.
[0061] In some embodiments, a first sparse feature of the image bag can be determined based on the basic features of the image bag and a first target dictionary; a second sparse feature of the image bag can be determined based on the basic features of the image bag and a target sub-dictionary related to the image bag label ratio in the first target dictionary; the first target dictionary is constrained based on the first sparse features of the image bag and the second sparse features of the image bag to obtain a second target dictionary.
[0062] In some embodiments, determining a first target dictionary of a sparse dictionary learning submodel includes: determining sparse features corresponding to basic features of an image bag under a current dictionary; determining reconstructed features of the basic features based on the sparse features and the dictionary; and optimizing the dictionary with minimizing the difference between the reconstructed features and the basic features as an optimization goal to obtain a first target dictionary.
[0063] Among them, the basic features are expressed as The dictionary is represented as D, and the corresponding sparse features are represented as The corresponding optimization goal is:
[0064]
[0065] in, is the reconstruction error of the data. The purpose of this part is to make the dictionary and sparse features reconstruct the basic features as accurately as possible. By minimizing this error, we ensure that the learned dictionary can effectively represent the training data. i || 1 The λ in is the regularization parameter, which controls the strength of the regularization term, ||Z i || 1 Represents the L1 norm of the coefficient feature, that is, Z i The sum of the absolute values of all elements in . L1 regularization minimizes Z i The L1 norm of Z promotes sparsity, making many Z i The elements of tend to zero, thus achieving sparse representation. By adjusting the value of λ, a trade-off can be made between reconstruction error and sparsity. A larger λ value will make the model more sparse, but may increase the reconstruction error; while a smaller λ value will reduce sparsity, but may result in a more accurate reconstruction.
[0066] In short, the purpose of this optimization goal is to find a dictionary that can accurately reconstruct the training data while making the representation of each data as sparse as possible, so as to improve the performance and generalization ability of the model.
[0067] Correspondingly, the update strategy corresponding to the above optimization objective is:
[0068]
[0069] Where μ is the step size, t represents the tth update iteration, D is the difference between the basic feature and the dictionary multiplied by the current sparse feature at the tth iteration. T is the transpose of D, It represents the gradient of sparse features, S λ (v j (=sign(v j )·max{|v j |-λ,0} is the element-by-element soft threshold operator, and max{·} is used to calculate the element-by-element maximum value.
[0070] Corresponding to the above update strategy, the sparse features obtained in each iteration corresponding to the update strategy can be expressed as:
[0071]
[0072] in,
[0073] Furthermore, the model structure diagram of the sparse dictionary learning sub-model is as follows Figure 2 As shown, the network layers are stacked in cascade L times, i.e., t = 1, 2, ..., L, to construct an L-layer unfolded network for sparse dictionary learning, and the dictionary D and the sparse feature representation Z are trained in a trainable manner. i Perform iterative learning. The update strategy corresponding to each single expansion layer is Thus, the dictionary D and the sparse feature representation Z can be expanded through multiple single layers. i Perform iterative learning to obtain the first target dictionary.
[0074] In some embodiments, a first target dictionary is constrained based on a first sparse feature of an image bag and a second sparse feature of an image bag to obtain a second target dictionary, including: obtaining a first reconstructed deep feature based on the first sparse feature of the image bag and the first target dictionary; obtaining a second reconstructed deep feature based on the first sparse feature of the image bag and a target sub-dictionary; calculating a second loss function value based on the first reconstructed deep feature and a basic deep feature; calculating a third loss function value based on the second reconstructed deep feature and the basic deep feature; adjusting the first target dictionary based on the second loss function value and the third loss function until a preset iteration completion condition is met to obtain a second target dictionary.
[0075] In some embodiments, by constraining the first target dictionary using a partial dictionary related to the dataset category and reconstructed basic features obtained by corresponding sparse features, it can be ensured that the sparse features are more focused on capturing information related to a specific category, thereby improving the category relevance of the features.
[0076] Among them, the second loss function value can be obtained by the following formula:
[0077]
[0078] Among them, F i,all =DZ i is the feature representation obtained by feature reconstruction using the entire category-related dictionary D, and d(·,·) is used to measure the reconstruction error.
[0079] The third loss function value can be obtained by the following formula:
[0080]
[0081] in, To use only the sub-dictionary D associated with each bag label sub The feature representation is obtained by feature reconstruction.
[0082] The two reconstruction constraints are weighted and summed to jointly constrain the learning of the dictionary and sparse features to make them discriminative:
[0083]
[0084] Among them, α is used to control the weight between the two reconstruction losses.
[0085] In some embodiments, based on the first loss function value, adjusting the parameters of the image classification model, continuing training until a preset first training completion condition is met, and obtaining a trained image classification model, includes:
[0086] Based on the first loss function value, the second loss function value and the third loss function value, the parameters of the image classification model are adjusted, and training is continued until the preset second training completion condition is met to obtain a trained image classification model.
[0087] In some embodiments, the total loss function value of the image classification model is determined based on the first loss function value, the second loss function value and the third loss function value, and the parameters of the image classification model are adjusted based on the total loss function value. Training continues until a preset second training completion condition is met to obtain a trained image classification model.
[0088] Among them, the second training completion condition can be that the number of training times reaches the second preset training times threshold, or it can be that the total loss function value reaches the second loss function threshold, wherein the second preset training times threshold and the second loss function threshold can be flexibly set based on actual conditions.
[0089] The total loss function value can be determined by the following formula:
[0090] Loss = L h_prop +βL bagReconst
[0091] Among them, Loss is the total loss function, L h_prop is the first loss function, L bagReconst It is determined based on the second loss function and the third loss function, and β is used to control the proportional weight between the two loss functions.
[0092] In some embodiments, the classifier includes sub-classifiers of different granularities, and sparse features are input into the classifier of the image classification model to obtain the predicted category label for each image in the image bag, including: inputting the sparse features into multiple sub-classifiers of different granularities to output the category prediction results for each image in the image bag at different granularities.
[0093] In some embodiments, the category ratio label of the image bag includes sub-category ratio labels of different granularities, and the first loss function value is determined based on the predicted category ratio label of the image bag and the category ratio label of the image bag, including: determining multiple sub-loss function values for the image bag based on the sub-predicted category ratio labels corresponding to the image bag at different granularities and the sub-category ratio labels corresponding to the image bag at different granularities; determining the first loss function value based on the multiple sub-loss function values.
[0094] It can be understood that the category ratio labels of the image bag include sub-category ratio labels of different granularities. For example, an image bag includes four images, and the image bag corresponds to sub-category ratio labels of two granularities, one of which is 50% cats and 50% dogs, and the other is 25% Persian cats, 25% Maine Coons, 25% Corgis, and 25% Dachshunds.
[0095] In order to correspond to sub-category proportion labels of different granularities, the classifier includes sub-classifiers of different granularities to output category prediction results of each image in the image bag at different granularities; based on the category prediction results of each image in the image bag at different granularities, the sub-prediction category proportion labels corresponding to the image bag at different granularities are determined.
[0096] The class ratio label can be expressed as
[0097] in, represents the proportion of images with category label c in the i-th input image bag, that is, the total number of image instances with label c in the i-th input image bag Zhang, c∈{1,2,…,C}, C is the total number of fine-grained image categories.
[0098] Further, multiple sub-loss function values for the image bag are determined based on the sub-prediction category ratio labels corresponding to the image bag at different granularities and the sub-category ratio labels corresponding to the image bag at different granularities.
[0099] Among them, the sub-loss function value can be determined by the following formula:
[0100]
[0101] in, is the category ratio loss at the lth granularity, with a total of H granularities, represents the label ratio of category c in the i-th bag at the l-th granularity, It represents the prediction value of the corresponding classifier, which is calculated by the classifier f(·) of the corresponding granularity l.
[0102] The first loss function value can be determined by the following formula:
[0103]
[0104] Where H is the number of particle sizes.
[0105] The present application also provides an image classification model training device, which is used to execute the image classification model training method provided in any of the above embodiments. Figure 3 As shown, the device comprises:
[0106] An acquisition module 301 is used to acquire basic features of image bags in a sample set and category ratio labels of the image bags;
[0107] An input module 302, configured to input the basic features into a sparse dictionary learning sub-model of an image classification model to obtain target sparse features of the image bag;
[0108] The input module 302 is further used to input the sparse features into the classifier of the image classification model to obtain a predicted category label for each image in the image bag;
[0109] A determination module 303 is used to determine the predicted category ratio label of the image bag based on the predicted category label of each image in the image bag.
[0110] The determination module 303 is further configured to determine a first loss function value based on the predicted category ratio label of the image bag and the category ratio label of the image bag;
[0111] The adjustment module 304 is used to adjust the parameters of the image classification model based on the first loss function value, and continue training until a preset first training completion condition is met to obtain a trained image classification model.
[0112] The present application proposes an image classification model training method, apparatus, device and storage medium, the method comprising: obtaining basic features of image bags in a sample set and category ratio labels of image bags; inputting the basic features into a sparse dictionary learning sub-model to obtain target sparse features of the image bag; inputting the target sparse features into a classifier to obtain predicted category labels for each image in the image bag; determining predicted category ratio labels of the image bag based on the predicted category labels of each image in the image bag; determining a first loss function value based on the predicted category ratio labels and the category ratio labels; training the image classification model with the first loss function value to obtain a trained image classification model. The embodiments of the present application effectively mine the discriminative features of the image classification model for each image when the data set only has category ratio labels, so as to achieve instance-level classification of images through bag-level category ratio labels.
[0113] In some embodiments, the input module 302 is specifically used to:
[0114] Determine a first target dictionary of a sparse dictionary learning sub-model;
[0115] Determine a first sparse feature of the image bag based on the basic feature of the image bag and the first target dictionary;
[0116] Determine a second sparse feature of the bag of images based on the basic feature of the bag of images and a target sub-dictionary in the first target dictionary that is related to the label ratio of the bag of images;
[0117] Constraining the first target dictionary based on the first sparse feature of the image bag and the second sparse feature of the image bag to obtain a second target dictionary;
[0118] The target sparse features of the image bag are determined based on the second target dictionary and the basic features of the image bag.
[0119] In some embodiments, the input module 302 is further specifically configured to:
[0120] Determine the sparse features corresponding to the basic features of the image bag under the current dictionary;
[0121] Determine a reconstruction feature of the basic feature based on the sparse feature and the dictionary;
[0122] The dictionary is optimized with minimizing the difference between the reconstructed feature and the basic feature as an optimization goal to obtain a first target dictionary.
[0123] In some embodiments, the input module 302 is further specifically configured to:
[0124] Obtaining a first reconstructed deep feature based on the first sparse feature of the image bag and the first target dictionary;
[0125] Obtaining a second reconstructed deep feature based on the first sparse feature of the image bag and the target sub-dictionary;
[0126] Calculate a second loss function value based on the first reconstructed deep feature and the basic deep feature;
[0127] Calculate a third loss function value based on the second reconstructed deep feature and the basic deep feature;
[0128] The first target dictionary is adjusted based on the second loss function value and the third loss function until a preset iteration completion condition is met to obtain a second target dictionary.
[0129] In some embodiments, the adjustment module 304 is specifically configured to:
[0130] Based on the first loss function value, the second loss function value and the third loss function value, adjust the parameters of the image classification model, continue training, and continue training until the preset second training completion condition is met to obtain a trained image classification model.
[0131] In some embodiments, the classifier includes sub-classifiers of different granularities, and the input module 302 is further specifically used for:
[0132] The sparse features are input into the multiple sub-classifiers of different granularities to output category prediction results of each image in the image bag at different granularities.
[0133] In some embodiments, the category ratio labels of the image bag include subcategory ratio labels of different granularities, and the determination module 303 is specifically configured to:
[0134] Determine a plurality of sub-loss function values for the image bag based on the sub-prediction category ratio labels corresponding to the image bag at different granularities and the sub-category ratio labels corresponding to the image bag at different granularities;
[0135] A first loss function value is determined based on the plurality of sub-loss function values.
[0136] The image classification model training device provided in the embodiment of the present application and the image classification model training method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented therein.
[0137] The present application also provides an electronic device to perform the above-mentioned image classification model training method. Figure 4 It shows a schematic diagram of an electronic device provided by some embodiments of the present application. Figure 4 As shown, the electronic device 8 includes: a processor 800, a memory 801, a bus 802 and a communication interface 803, and the processor 800, the communication interface 803 and the memory 801 are connected via the bus 802; the memory 801 stores a computer program that can be run on the processor 800, and when the processor 800 runs the computer program, it executes the image classification model training method provided in any of the aforementioned embodiments of the present application.
[0138] The memory 801 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the device network element and at least one other network element is realized through at least one communication interface 803 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.
[0139] The bus 802 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 801 is used to store programs, and the processor 800 executes the programs after receiving the execution instruction. The image classification model training method disclosed in any implementation of the aforementioned embodiment of the present application may be applied to the processor 800, or implemented by the processor 800.
[0140] The processor 800 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 800. The above processor 800 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a readily available programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to be executed, or the hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 801, and the processor 800 reads the information in the memory 801 and completes the steps of the above method in combination with its hardware.
[0141] The electronic device provided in the embodiment of the present application and the image classification model training method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented therein.
[0142] The present application also provides a computer-readable storage medium corresponding to the image classification model training method provided in the above embodiment. Figure 5 The computer-readable storage medium shown is a CD 30 on which a computer program (ie, a program product) is stored. When the computer program is run by a processor, the image classification model training method provided by any of the aforementioned embodiments will be executed.
[0143] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.
[0144] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the image classification model training method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0145] It should be noted that:
[0146] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies are not shown in detail so as not to obscure the understanding of this description.
[0147] Similarly, it should be understood that in order to streamline the present application and help understand one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be interpreted as reflecting the following schematic diagram: the claimed application requires more features than the features clearly stated in each claim. More specifically, as reflected in the claims below, the inventive aspects are less than all the features of the single embodiment disclosed above. Therefore, the claims following the specific embodiment are hereby expressly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present application.
[0148] In addition, those skilled in the art will appreciate that, although some embodiments described herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present application and form different embodiments. For example, in the claims below, any one of the claimed embodiments may be used in any combination.
[0149] The above is only a preferred specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A method for training an image classification model, characterized in that: include: Obtain basic features of image bags in a sample set and category ratio labels of the image bags; Inputting the basic features into a sparse dictionary learning sub-model of an image classification model to obtain target sparse features of the image bag; Inputting the sparse features into a classifier of the image classification model to obtain a predicted category label for each image in the bag of images; Determining a predicted category ratio label for the bag of images based on the predicted category label for each image in the bag of images; Determining a first loss function value based on the predicted category ratio label of the bag of images and the category ratio label of the bag of images; Based on the first loss function value, the parameters of the image classification model are adjusted, and training is continued until a preset first training completion condition is met to obtain a trained image classification model.
2. The method according to claim 1, characterized in that The step of inputting the basic deep features into a sparse dictionary learning sub-model of an image classification model to obtain target sparse features of the image bag includes: Determine a first target dictionary of a sparse dictionary learning sub-model; Determine a first sparse feature of the image bag based on the basic feature of the image bag and the first target dictionary; Determine a second sparse feature of the bag of images based on the basic feature of the bag of images and a target sub-dictionary in the first target dictionary that is related to the label ratio of the bag of images; Constraining the first target dictionary based on the first sparse feature of the image bag and the second sparse feature of the image bag to obtain a second target dictionary; The target sparse features of the image bag are determined based on the second target dictionary and the basic features of the image bag.
3. The method according to claim 2, characterized in that The determining of the first target dictionary of the sparse dictionary learning sub-model comprises: Determine the sparse features corresponding to the basic features of the image bag under the current dictionary; Determine a reconstruction feature of the basic feature based on the sparse feature and the dictionary; The dictionary is optimized with minimizing the difference between the reconstructed feature and the basic feature as an optimization goal to obtain a first target dictionary.
4. The method according to claim 2, characterized in that: The constraining the first target dictionary based on the first sparse feature of the image bag and the second sparse feature of the image bag to obtain a second target dictionary includes: Obtaining a first reconstructed deep feature based on the first sparse feature of the image bag and the first target dictionary; Obtaining a second reconstructed deep feature based on the first sparse feature of the image bag and the target sub-dictionary; Calculate a second loss function value based on the first reconstructed deep feature and the basic deep feature; Calculate a third loss function value based on the second reconstructed deep feature and the basic deep feature; The first target dictionary is adjusted based on the second loss function value and the third loss function until a preset iteration completion condition is met to obtain a second target dictionary.
5. The method according to claim 4, characterized in that The adjusting the parameters of the image classification model based on the first loss function value, continuing the training until a preset first training completion condition is met, and obtaining a trained image classification model includes: Based on the first loss function value, the second loss function value and the third loss function value, adjust the parameters of the image classification model, continue training, and continue training until the preset second training completion condition is met to obtain a trained image classification model.
6. The method according to claim 1, characterized in that The classifier includes sub-classifiers of different granularities, and the inputting the sparse features into the classifier of the image classification model to obtain the predicted category label of each image in the image bag includes: The sparse features are input into the multiple sub-classifiers of different granularities to output category prediction results of each image in the image bag at different granularities.
7. The method according to claim 6, characterized in that The category ratio label of the image bag includes subcategory ratio labels of different granularities, and determining the first loss function value based on the predicted category ratio label of the image bag and the category ratio label of the image bag includes: Determine a plurality of sub-loss function values for the image bag based on the sub-prediction category ratio labels corresponding to the image bag at different granularities and the sub-category ratio labels corresponding to the image bag at different granularities; A first loss function value is determined based on the plurality of sub-loss function values.
8. An image classification model training device, characterized in that: include: An acquisition module, used to acquire basic features of image bags in a sample set and category ratio labels of the image bags; An input module, used for inputting the basic features into a sparse dictionary learning sub-model of an image classification model to obtain target sparse features of the image bag; The input module is further used to input the sparse features into the classifier of the image classification model to obtain a predicted category label for each image in the image bag; A determination module, configured to determine a predicted category ratio label of the image bag based on the predicted category label of each image in the image bag; The determination module is used to determine a first loss function value based on the predicted category ratio label of the image bag and the category ratio label of the image bag; An adjustment module is used to adjust the parameters of the image classification model based on the first loss function value, and continue training until a preset first training completion condition is met to obtain a trained image classification model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The processor runs the computer program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Image super-resolution reconstruction method based on multitask KSVD (K singular value decomposition) dictionary learning
CN102156875A
Multi-label classification method, system and device and storage medium
CN109948735A
Training method of multi-task classification model, data classification method and related equipment
CN114359612A
Image classification model training method and device, computer equipment and medium
CN116071613A
Image classification model training method, image processing method and device
CN116824194A