A method, device, equipment and storage medium for training an image classification model
Patent Information
- Application Number
- CN202211140067.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-20
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-09-20
AI Technical Summary
[0054]本申请实施例中,采用已训练的辅助分类模型,对图像分类模型进行训练,使得图像分类模型可以学习到辅助分类模型的图像分类能力,在辅助分类模型的模型参数远多于图像分类模型的模型参数时,训练出的目标图像分类模型可以以较少的模型参数,提供较为准备地图像分类服务。
Smart Images

Figure CN117011567B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for training an image classification model. Background Technology
[0002] With the continuous development of technology, more and more devices can provide image classification services through trained target image classification models. Image classification services can be used to determine the category to which a target in an image belongs.
[0003] For example, a device can use a trained target image classification model to classify face images in order to identify the object corresponding to the face in the face image.
[0004] In related technologies, to balance the image classification time and accuracy of a trained target image classification model, a common method for obtaining a trained target image classification model involves first training a large auxiliary classification model with numerous parameters on a sample image set for multiple rounds, thus obtaining a trained auxiliary classification model. Then, based on the image categories output by the trained auxiliary classification model from the sample image set, a small image classification model with fewer parameters is trained for multiple rounds. This ensures that the image categories output by the obtained trained target image classification model are similar to those output by the trained auxiliary classification model. Therefore, by using a trained target image classification model, image classification time can be reduced without sacrificing too much image classification accuracy.
[0005] However, due to the significant difference in the number of parameters between the auxiliary classification model and the image classification model, training a small image classification model with a limited number of parameters multiple times based solely on the image categories of the sample image set output by the trained auxiliary classification model results in relatively simple constraints. This makes it impossible for the trained target image classification model to accurately learn the image classification capabilities of the auxiliary classification model, thus leading to low image classification accuracy of the trained target image classification model.
[0006] It is evident that the training methods employed under the relevant technologies cannot guarantee the classification accuracy and reliability of the trained target image classification model without increasing the image classification time. Summary of the Invention
[0007] This application provides a method, apparatus, computer device, and storage medium for training an image classification model, which addresses the problem of low classification accuracy and reliability of the trained target image classification model.
[0008] Firstly, a method for training an image classification model is provided, including:
[0009] Obtain a set of sample images associated with multiple sample categories, with each sample category associated with multiple sample images;
[0010] The trained auxiliary classification model is used to extract features from each sample image to obtain the corresponding auxiliary image features. For each sample category, the following operations are performed: the auxiliary image features of multiple sample images associated with a sample category are fused to obtain the corresponding auxiliary category features.
[0011] Based on the sample image set, the image classification model to be trained is subjected to multiple rounds of iterative training, and the trained target image classification model is output. Each round of iteration includes:
[0012] Using the image classification model, feature extraction is performed on multiple sample images associated with a sample category to obtain the corresponding training image features;
[0013] The model parameters of the image classification model are adjusted based on the training image features and auxiliary image features corresponding to each of the multiple sample images, as well as the auxiliary category features of each of the multiple sample categories.
[0014] Secondly, an apparatus for training an image classification model is provided, comprising:
[0015] Acquisition module: Used to acquire sample image sets associated with multiple sample categories, with each sample category associated with multiple sample images;
[0016] Processing module: Used to extract features from each sample image using a trained auxiliary classification model to obtain corresponding auxiliary image features, and to perform the following operations for each sample category: perform feature fusion on the auxiliary image features of multiple sample images associated with a sample category to obtain corresponding auxiliary category features;
[0017] The processing module is further configured to: perform multiple rounds of iterative training on the image classification model to be trained based on the sample image set, and output the trained target image classification model, wherein each round of iteration includes:
[0018] The processing module is further configured to: use the image classification model to extract features from multiple sample images associated with a sample category to obtain corresponding training image features;
[0019] The model parameters of the image classification model are adjusted based on the training image features and auxiliary image features corresponding to each of the multiple sample images, as well as the auxiliary category features of each of the multiple sample categories.
[0020] Optionally, the processing module is specifically used for:
[0021] Based on the training image features and auxiliary image features corresponding to each of the multiple sample images, the feature consistency loss of the image classification model is determined, wherein the feature consistency loss characterizes the consistency of feature extraction using the auxiliary classification model and the image classification model;
[0022] Based on the training image features and auxiliary image features corresponding to each of the multiple sample images, and the auxiliary category features of each of the multiple sample categories, the relation consistency loss of the image classification model is determined, wherein the relation consistency loss represents the consistency of the inter-class relations between the multiple sample categories determined by the auxiliary classification model and the image classification model;
[0023] Based on the obtained feature consistency loss and relation consistency loss, the model parameters of the image classification model are adjusted.
[0024] Optionally, the processing module is specifically used for:
[0025] For each of the multiple sample images, perform the following operations:
[0026] The various sample categories are used as target categories, and the similarity between the training image features corresponding to a sample image and the auxiliary category features of each of the various target categories is determined to obtain the corresponding training similarity.
[0027] The similarity between the auxiliary image features corresponding to each sample image and the auxiliary category features of each of the multiple target categories is determined to obtain the corresponding auxiliary similarity.
[0028] Based on the obtained multiple training similarities and multiple auxiliary similarities, the relation consistency loss of the image classification model is determined.
[0029] Optionally, the processing module is specifically used for:
[0030] For each of the multiple sample images, perform the following operations:
[0031] The similarity between the auxiliary image features corresponding to the sample image and the auxiliary category features of the various sample categories is determined. Among the obtained similarities, the multiple similarities that are greater than a preset similarity threshold are respectively used as auxiliary similarities, and the sample categories corresponding to the multiple obtained auxiliary similarities are respectively used as target categories.
[0032] The similarity between the training image features corresponding to each sample image and the auxiliary category features of each of the multiple target categories is determined to obtain the corresponding training similarity.
[0033] Based on the obtained multiple training similarities and multiple auxiliary similarities, the relation consistency loss of the image classification model is determined.
[0034] Optionally, the processing module is specifically used for:
[0035] For each of the aforementioned target categories, perform the following operations respectively:
[0036] The average training similarity between the multiple sample images and the auxiliary category features of a target category is used as the training inter-class relationship between the sample category and the target category.
[0037] The average value of the auxiliary similarities between the plurality of sample images and the auxiliary category features of the target category is taken as the auxiliary inter-class relationship between the sample category and the target category.
[0038] Based on the errors between multiple training class relationships and multiple auxiliary class relationships obtained, the relationship consistency loss of the image classification model is determined.
[0039] Optionally, the processing module is specifically used for:
[0040] The errors between the training image features and auxiliary image features corresponding to each of the multiple sample images are determined to obtain the corresponding feature consistency errors;
[0041] The feature consistency loss of the image classification model is determined based on the weighted average of the obtained feature consistency errors.
[0042] Optionally, the processing module is specifically used for:
[0043] The weighted sum of the obtained feature consistency loss and relation consistency loss is used as the training loss of the image classification model.
[0044] If the obtained training loss fails to achieve the training objective, the model parameters of the image classification model are adjusted.
[0045] Optionally, the processing module is further configured to:
[0046] Obtain the image to be classified;
[0047] The target image classification model is used to extract features from the image to be classified to obtain target image features;
[0048] Using the target image classification model, the target category of the image to be classified is determined from the multiple sample categories based on the target image features.
[0049] Thirdly, a computer program product is provided, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0050] Fourthly, a computer device is provided, comprising:
[0051] Memory, used to store program instructions;
[0052] A processor is configured to invoke program instructions stored in the memory and execute the method described in the first aspect according to the obtained program instructions.
[0053] Fifthly, a computer-readable storage medium is provided, the computer-readable storage medium storing computer-executable instructions for causing a computer to perform the method as described in the first aspect.
[0054] In this embodiment, a trained auxiliary classification model is used to train the image classification model, so that the image classification model can learn the image classification ability of the auxiliary classification model. When the model parameters of the auxiliary classification model are much more than the model parameters of the image classification model, the trained target image classification model can provide a more accurate image classification service with fewer model parameters.
[0055] Furthermore, during the training of the image classification model, in each training round, the model parameters of the image classification model are adjusted based on the auxiliary image features of multiple sample images extracted by the auxiliary classification model and the training image features of multiple sample images extracted by the image classification model. From the perspective of extracting image features, the image classification model learns the feature extraction capability of the auxiliary classification model, so that the trained target image classification model can provide more accurate and reliable image classification services.
[0056] Furthermore, an auxiliary classification model is used to determine the auxiliary category features of various sample categories. Based on the auxiliary image features of multiple sample images extracted by the auxiliary classification model and the training image features of multiple sample images extracted by the image classification model, the similarity relationship between features can be used to learn the sample category in this round of training and the inter-class relationship between it and other sample categories. Thus, from the perspective of inter-class relationships, the image classification model learns the ability of the auxiliary classification model to selectively acquire features in the image. As a result, the trained target image classification model can provide more accurate and reliable image classification services.
[0057] By training the image classification model from multiple perspectives, including feature relationships and inter-class relationships, the detection accuracy and reliability of the trained target image classification model are improved. Attached Figure Description
[0058] Figure 1A This is a schematic diagram illustrating the application field of the image classification model provided in the embodiments of this application;
[0059] Figure 1B This is one application scenario of the method for training an image classification model provided in the embodiments of this application;
[0060] Figure 2 A flowchart illustrating a method for training an image classification model provided in an embodiment of this application;
[0061] Figure 3 A schematic diagram illustrating the principle of a method for training an image classification model provided in this application embodiment;
[0062] Figure 4 A schematic diagram of the principle of the method for training an image classification model provided in the embodiments of this application. Figure 2 ;
[0063] Figure 5 A schematic diagram of the principle of the method for training an image classification model provided in the embodiments of this application. Figure 3 ;
[0064] Figure 6 A schematic diagram of the principle of the method for training an image classification model provided in the embodiments of this application. Figure 4 ;
[0065] Figure 7A A schematic diagram of the principle of the method for training an image classification model provided in the embodiments of this application. Figure 5 ;
[0066] Figure 7B A schematic diagram of the principle of the method for training an image classification model provided in the embodiments of this application. Figure 6 ;
[0067] Figure 8 A schematic diagram of a device for training an image classification model provided in an embodiment of this application;
[0068] Figure 9 A schematic diagram of the structure of an apparatus for training an image classification model provided in the embodiments of this application. Figure 2 . Detailed Implementation
[0069] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0070] The following explanations of some terms used in the embodiments of this application are provided to facilitate understanding by those skilled in the art.
[0071] (1) Intra-class relations and inter-class relations:
[0072] Intra-class relations represent the relationships between images belonging to the same category, such as the relationship between different face images of the same object.
[0073] Interclass relationships represent the relationships between images belonging to different categories, such as the relationships between facial images of different objects.
[0074] This application relates to the field of Artificial Intelligence (AI), and is designed based on Computer Vision (CV) and Machine Learning (ML) technologies. It can be applied to fields such as cloud computing, smart transportation, smart agriculture, smart healthcare, and mapping.
[0075] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that studies the design principles and implementation methods of various machines, attempting to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence, enabling machines to have perception, reasoning, and decision-making functions.
[0076] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, interactive operating systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, machine learning / deep learning, autonomous driving, and intelligent transportation. With the development and progress of AI, it has been researched and applied in numerous fields, such as smart homes, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, wearable devices, autonomous driving, drones, robots, smart healthcare, vehicle networking, and intelligent transportation. It is believed that with further technological advancements, AI will be applied in even more fields, playing an increasingly important role. The solutions provided in this application's embodiments relate to deep learning and augmented reality technologies in AI, which are further illustrated by the following examples.
[0077] Computer vision is the science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in identifying, tracking, and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), autonomous driving, intelligent transportation, and other technologies, as well as common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0078] Machine learning is a multidisciplinary field that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and many other disciplines. It specifically studies how computers acquire new knowledge or skills by simulating human learning behavior, reorganize existing knowledge structures, and continuously improve their performance.
[0079] Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence. Its applications span all areas of artificial intelligence. Deep learning, the core of machine learning, is a technology that enables machine learning. Machine learning typically includes techniques such as deep learning, reinforcement learning, transfer learning, inductive learning, artificial neural networks, and instructional learning. Deep learning includes techniques such as convolutional neural networks (CNNs), deep belief networks, recurrent neural networks, autoencoders, and generative adversarial networks.
[0080] It should be noted that in the embodiments of this application, data such as sample image sets or images to be classified are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0081] The following is a brief introduction to the application areas of the method for training image classification models provided in the embodiments of this application.
[0082] With the continuous development of technology, more and more devices can provide image classification services through trained target image classification models. Image classification services can be used to determine the category to which a target in an image belongs.
[0083] For example, in practical applications such as security, payment, or access control, a facial recognition system in a mobile terminal can be used to capture a facial image of the target object. Then, a trained target image classification model can be used to classify the facial image to determine the identity of the object corresponding to the face in the image. Please refer to [reference needed]. Figure 1A The trained target image classification model is used to classify the face image A and determine the object corresponding to the face in the face image A as object A.
[0084] Therefore, since it's a mobile terminal, there are high requirements for both the image classification time and accuracy of the target image classification model. To reduce the image classification time, it's common practice to train an image classification model with fewer parameters to obtain a pre-trained target image classification model. If a large number of sample images are used directly to train the image classification model, the model's fitting ability is poor. Therefore, during training, it will get stuck in a local minimum of the loss function and begin to oscillate, making further optimization of the loss function impossible. This results in low image classification accuracy for the trained target image classification model.
[0085] In related technologies, to balance the image classification time and accuracy of a trained target image classification model, a common method for obtaining a trained target image classification model involves first training a large auxiliary classification model with numerous parameters on a sample image set for multiple rounds, thus obtaining a trained auxiliary classification model. Then, based on the image categories output by the trained auxiliary classification model from the sample image set, a small image classification model with fewer parameters is trained for multiple rounds. This ensures that the image categories output by the obtained trained target image classification model are similar to those output by the trained auxiliary classification model. Therefore, by using a trained target image classification model, image classification time can be reduced without sacrificing too much image classification accuracy.
[0086] However, due to the significant difference in the number of parameters between the auxiliary classification model and the image classification model, training a small image classification model with a limited number of parameters multiple times based solely on the image categories of the sample image set output by the trained auxiliary classification model results in relatively simple constraints. This makes it impossible for the trained target image classification model to accurately learn the image classification capabilities of the auxiliary classification model, thus leading to low image classification accuracy of the trained target image classification model.
[0087] It is evident that the training methods employed under the relevant technologies cannot guarantee the classification accuracy and reliability of the trained target image classification model without increasing the image classification time.
[0088] To address the issue of low classification accuracy and reliability in trained target image classification models, this application proposes a method for training image classification models. This method first acquires a set of sample images associated with multiple sample categories, with each category associated with multiple sample images. Using a pre-trained auxiliary classification model, features are extracted from each sample image to obtain corresponding auxiliary image features. Furthermore, for each sample category, the auxiliary image features of multiple sample images associated with a particular category are fused to obtain corresponding auxiliary category features. Based on the sample image set, the image classification model to be trained undergoes multiple rounds of iterative training, outputting a trained target image classification model.
[0089] Each iteration includes: using an image classification model, extracting features from multiple sample images associated with a single sample category to obtain corresponding training image features. Based on the training image features and auxiliary image features corresponding to each of the multiple sample images, as well as the auxiliary category features for each of the various sample categories, the model parameters of the image classification model are adjusted.
[0090] In this embodiment, a trained auxiliary classification model is used to train the image classification model, so that the image classification model can learn the image classification ability of the auxiliary classification model. When the model parameters of the auxiliary classification model are much more than the model parameters of the image classification model, the trained target image classification model can provide a more accurate image classification service with fewer model parameters.
[0091] Furthermore, during the training of the image classification model, in each training round, the model parameters of the image classification model are adjusted based on the auxiliary image features of multiple sample images extracted by the auxiliary classification model and the training image features of multiple sample images extracted by the image classification model. From the perspective of extracting image features, the image classification model learns the feature extraction capability of the auxiliary classification model, so that the trained target image classification model can provide more accurate and reliable image classification services.
[0092] Furthermore, an auxiliary classification model is used to determine the auxiliary category features of various sample categories. Based on the auxiliary image features of multiple sample images extracted by the auxiliary classification model and the training image features of multiple sample images extracted by the image classification model, the similarity relationship between features can be used to learn the sample category in this round of training and the inter-class relationship between it and other sample categories. Thus, from the perspective of inter-class relationships, the image classification model learns the ability of the auxiliary classification model to selectively acquire features in the image. As a result, the trained target image classification model can provide more accurate and reliable image classification services.
[0093] By training the image classification model from multiple perspectives, including feature relationships and inter-class relationships, the detection accuracy and reliability of the trained target image classification model are improved.
[0094] The following describes the application scenarios of the method for training the image classification model provided in this application.
[0095] Please refer to Figure 1B This diagram illustrates an application scenario of the image classification model training method provided in this application. The application scenario includes a client 101 and a server 102. The client 101 and the server 102 can communicate with each other. The communication method can be wired, such as through a network cable or serial cable; or wireless, such as through Bluetooth or Wi-Fi. No specific limitation is imposed.
[0096] Client 101 generally refers to a device that can provide sample image sets to server 102 or use a trained target image classification model, such as a terminal device, a third-party application accessible to the terminal device, or a webpage accessible to the terminal device. Terminal devices include, but are not limited to, mobile phones, computers, smart medical devices, smart home appliances, vehicle terminals, or aircraft. Server 102 generally refers to a device that can train or use a target image classification model, such as a terminal device or a server. Servers include, but are not limited to, cloud servers, local servers, or associated third-party servers. Both client 101 and server 102 can use cloud computing to reduce the consumption of local computing resources; similarly, they can also use cloud storage to reduce the consumption of local storage resources.
[0097] As one embodiment, the client 101 and the server 102 can be the same device, and there is no specific limitation. In this embodiment, the client 101 and the server 102 are described as different devices.
[0098] The following is based on Figure 1B Using server 102 as the main component, this paper provides a detailed description of the method for training an image classification model provided in this embodiment. Please refer to [link / reference]. Figure 2 This is a flowchart illustrating a method for training an image classification model provided in an embodiment of this application.
[0099] S201, Obtain a set of sample images that are associated with multiple sample categories.
[0100] The sample image set contains a large number of sample images, each of which is associated with a corresponding sample category. Each sample category is associated with multiple sample images. Taking face images as an example, each object can be associated with multiple face images. These multiple face images can include frontal face images, side face images, top-view face images, or top-view face images of the corresponding object, etc., without any specific restrictions.
[0101] By acquiring a set of sample images that are associated with multiple sample categories, the image classification model can be trained on the same sample category based on multiple sample images during training. This enriches the prior knowledge of the image classification model, thereby improving the classification accuracy and reliability of the trained target image classification model.
[0102] S202 uses a trained auxiliary classification model to extract features from each sample image to obtain corresponding auxiliary image features, and performs the following operations for each sample category: fuse the auxiliary image features of multiple sample images associated with a sample category to obtain corresponding auxiliary category features.
[0103] Before extracting features from each sample image using a trained auxiliary classification model, the auxiliary classification model can be iteratively trained multiple times based on the sample image set to output a trained auxiliary classification model. An auxiliary classification model contains a large number of parameters, far exceeding the number of parameters in an image classification model. Auxiliary classification models are used for image classification.
[0104] The following example illustrates the process of training an auxiliary classification model using a single round of iterative training. The process for each round of iterative training is similar and will not be repeated here.
[0105] Using an auxiliary classification model to be trained, multiple sample images belonging to each sample category can be read from the sample image set.
[0106] After reading multiple sample images belonging to a single sample category, the auxiliary classification model to be trained sequentially extracts features from these multiple sample images to obtain their individual image features. For example, spatial feature extraction can be performed on each of the multiple sample images to obtain their individual spatial features. The auxiliary classification model to be trained can employ convolutional neural networks to provide feature extraction capabilities. Convolutional neural networks can include convolutional layers, non-linear activation function layers, pooling layers, etc., without specific limitations.
[0107] After obtaining the image features of multiple sample images, an auxiliary classification model to be trained can be used to determine the category features of that sample category based on the image features of the multiple sample images belonging to that category. Thus, based on the multiple sample images read each time, the category features of each sample category can be obtained. The category features of all sample categories can be stored in matrix form, for example, with a matrix shape of (d×m), where d is the dimension of the category features of each sample category, and m is the number of sample categories.
[0108] Based on the image features of multiple sample images belonging to a sample category, the category features of that sample category are determined. This can be done by taking the average of the image features of the multiple sample images belonging to a sample category as the category features of that sample category; or by calculating cluster centers based on the image features of the multiple sample images belonging to a sample category and taking the calculated cluster centers as the category features of that sample category, etc. There are no specific restrictions.
[0109] For each read sample image belonging to one sample category, the auxiliary classification model to be trained is iterated and trained in this round.
[0110] Using the auxiliary classification model to be trained, for the image features of multiple sample images belonging to one sample category, the matrix product between the image features and the category features of each sample category is calculated. An activation function, such as the softmax function, is then used to obtain the probability value of each sample image belonging to each sample category. The probability values of each sample image belonging to each sample category can form a probability vector for that sample image. The sample categories are associated with a fixed arrangement order; arranging the corresponding probability values according to this order yields the probability vector. Based on the error between the probability vector of the sample image and its sample category, the loss function of the auxiliary classification model to be trained is determined. The sample category of a sample image can be viewed as a probability vector, where the probability value of the sample category is 1, and the rest are 0.
[0111] The loss function can be in the form of an activation function or various activation functions with margins, etc., and there are no specific restrictions.
[0112] When the loss function does not meet the training objective, gradient descent algorithms, such as stochastic gradient descent, stochastic gradient descent with a moving average, Adam, or Adamard algorithms, can be used to adjust the parameters of the auxiliary classification model to be trained and enter the next round of iterative training until the obtained loss function meets the training objective, at which point the trained auxiliary classification model is output. The training objective can be reaching a preset threshold for the number of iterations, or the loss function being less than a preset value, etc., and there are no specific restrictions.
[0113] After obtaining a trained auxiliary classification model, which contains a large number of model parameters, and after sufficient training with a large number of sample images, the trained auxiliary classification model has high classification accuracy and reliability. Therefore, the trained auxiliary classification model can be used to extract features from each sample image to obtain the corresponding auxiliary image features.
[0114] For each sample category, the corresponding auxiliary category features are calculated. The following example uses one sample category, and the situation is similar for each sample category, so it will not be repeated here.
[0115] An auxiliary classification model is used to fuse the auxiliary image features of multiple sample images associated with a sample category to obtain the corresponding auxiliary category features. Feature fusion can be performed by calculating the average feature of each auxiliary image feature and using the average feature as the auxiliary category feature; or it can be performed by calculating the cluster centers of each auxiliary image feature and using the cluster centers as the auxiliary category feature, etc., and there is no specific limitation.
[0116] By calculating auxiliary category features for each sample category, for a given sample category, the auxiliary image features of multiple sample images associated with that category show high similarity to the auxiliary category features of that sample category. For different sample categories, the auxiliary image features of sample images associated with one category show low similarity to the auxiliary category features of other sample categories. Therefore, by using auxiliary image features and auxiliary category features, the image classification model to be trained can learn the intra-class and inter-class properties of the features extracted by the auxiliary classification model. This allows for comprehensive learning of the auxiliary classification model from multiple perspectives, thereby improving the classification accuracy and reliability of the trained target image classification model.
[0117] Please refer to Figure 3 An auxiliary classification model may include a data preparation module, a large-scale recognition module, a category feature storage module, a training module, and an optimization module.
[0118] The data preparation module can read multiple sample images belonging to each sample category. The large-scale recognition module contains a large number of model parameters from the auxiliary classification model. This module sequentially extracts features from the read sample images to obtain the individual image features of each sample image. The category feature storage module determines the category features of a single sample category based on the individual image features of the multiple sample images belonging to that category, thus storing the category features for each sample category. The training module performs multiple rounds of iterative training on the auxiliary classification model to be trained. During each round of training, the loss function of the auxiliary classification model to be trained is determined, and it is determined whether the loss function meets the training objective. The optimization module adjusts the parameters of the auxiliary classification model to be trained when the loss function does not meet the training objective; when the loss function meets the training objective, the trained auxiliary classification model is output.
[0119] S203, based on the sample image set, performs multiple rounds of iterative training on the image classification model to be trained, and outputs the trained target image classification model.
[0120] After obtaining the sample image set, the auxiliary image features of each sample image in the sample image set, and the auxiliary category features of each sample category, the image classification model to be trained can be trained in multiple rounds of iteration to output the trained target image classification model.
[0121] The following description uses one round of iterative training as an example. The process of each round of iterative training is similar and will not be repeated here. Please refer to S204 to S205.
[0122] S204 uses an image classification model to extract features from multiple sample images associated with a sample category to obtain the corresponding training image features.
[0123] In each iteration of training, multiple sample images associated with a single sample category can be processed. An image classification model is used to extract features from these multiple sample images associated with a single sample category, obtaining the corresponding training image features. These training image features characterize the current feature extraction capability and image classification ability of the image classification model.
[0124] To make the image classification ability of the image classification model closer to that of the auxiliary classification model, the similarity between the training image features extracted by the image classification model and the auxiliary image features extracted by the auxiliary classification model can be improved during training. There are various methods to improve the similarity between the training image features extracted by the image classification model and the auxiliary image features extracted by the auxiliary classification model. However, if only the error between the training image features and the auxiliary image features is reduced, the constraints are too simplistic, making it difficult to ensure that the training image features of each sample image are highly similar to the auxiliary image features, and the training time is also long. Therefore, in this embodiment, in addition to the training image features and auxiliary image features, constraints are further added based on the auxiliary category features, giving the training process a clearer objective. This shortens the training time while making the image classification ability of the image classification model closer to that of the auxiliary classification model. Please refer to S205.
[0125] S205, based on the training image features and auxiliary image features corresponding to multiple sample images, as well as the auxiliary category features of multiple sample categories, the model parameters of the image classification model are adjusted.
[0126] After obtaining the training image features and auxiliary image features corresponding to multiple sample images, as well as the auxiliary category features for each of the various sample categories, the model parameters of the image classification model can be adjusted based on these features. Compared to methods that only adjust the model parameters based on the training image features and auxiliary image features corresponding to multiple sample images, this approach provides richer prior knowledge, making the training process more efficient and resulting in a more accurate and reliable classification model for the target image.
[0127] As one embodiment, the model parameters of the image classification model are adjusted based on the training image features and auxiliary image features corresponding to multiple sample images, as well as the auxiliary category features of multiple sample categories. This can be achieved by determining the feature consistency loss of the image classification model based on the training image features and auxiliary image features corresponding to multiple sample images, where the feature consistency loss represents the consistency of feature extraction using the auxiliary classification model and the image classification model. Similarly, the relation consistency loss of the image classification model is determined based on the training image features and auxiliary image features corresponding to multiple sample images, as well as the auxiliary category features of multiple sample categories, where the relation consistency loss represents the consistency of inter-class relationships determined using the auxiliary classification model and the image classification model. Based on the obtained feature consistency loss and relation consistency loss, the model parameters of the image classification model are then adjusted.
[0128] Feature consistency loss ensures the similarity between training image features extracted by the image classification model and auxiliary image features extracted by the auxiliary classification model. Relation consistency loss reduces intra-class similarity between sample images belonging to the same sample category and increases inter-class similarity between sample images belonging to different sample categories. Thus, the training process of the image classification model is constrained through multiple dimensions.
[0129] As one embodiment, when determining the feature consistency loss of an image classification model based on the training image features and auxiliary image features corresponding to each of multiple sample images, the errors between the training image features and auxiliary image features corresponding to each of the multiple sample images can be determined separately to obtain the corresponding feature consistency errors. Based on the weighted average of the obtained feature consistency errors, the feature consistency loss L of the image classification model is determined. f Please refer to formula (1).
[0130] L f =||F x -F y ‖twenty one)
[0131] Among them, F x F represents the training image features corresponding to each of the multiple sample images. y The auxiliary image features corresponding to each of the multiple sample images are represented, and the L2 norm is calculated using the ||·||2 characterization.
[0132] As an example, when determining the relationship consistency loss of an image classification model based on the training image features and auxiliary image features corresponding to multiple sample images, as well as the auxiliary category features of multiple sample categories, we will take one sample image as an example for introduction. The situation for each sample image is similar and will not be repeated here.
[0133] Multiple sample categories are used as target categories. For each sample image, the similarity between the training image features and the auxiliary category features of each of the multiple target categories is determined to obtain the corresponding training similarity. Similarly, the similarity between the auxiliary image features and the auxiliary category features of each sample image and the auxiliary category features of each of the multiple target categories is determined to obtain the corresponding auxiliary similarity. Based on the obtained training and auxiliary similarities, the relationship consistency loss of the image classification model is determined.
[0134] Auxiliary similarity can be pre-calculated by the auxiliary relation calculation model after the auxiliary classification model extracts the auxiliary image features of each sample image. Therefore, when it is necessary to use auxiliary similarity to calculate the relation consistency loss of the image classification model, the pre-calculated auxiliary similarity can be directly read from the auxiliary relation calculation model, which improves the efficiency of determining the relation consistency loss and thus improves the efficiency of training the image classification model.
[0135] Please refer to Figure 4 The auxiliary relation calculation model can include a data preparation module, an inter-class relation calculation module, and a relation storage module.
[0136] The data preparation module reads multiple sample images belonging to each sample category. It can share a data preparation module with the auxiliary classification model, and there are no specific restrictions. The inter-class relationship calculation module obtains the auxiliary image features corresponding to each sample image from the auxiliary classification model, as well as the auxiliary category features for each of the multiple target categories. For each sample image, it determines the similarity between the auxiliary image features corresponding to one sample image and the auxiliary category features for each of the multiple target categories, thus obtaining the corresponding auxiliary similarity. The relationship storage module stores the multiple auxiliary similarities corresponding to a single sample image.
[0137] As an example, if there are many sample categories, a subset of these categories can be selected to determine the relationship consistency loss of the image classification model. This will be illustrated using a single sample image as an example; the situation is similar for each sample image and will not be repeated here.
[0138] The similarity between the auxiliary image features corresponding to a sample image and the auxiliary category features of each of the multiple sample categories is determined. Among the obtained similarities, the multiple similarities that are greater than the preset similarity threshold are used as auxiliary similarities.
[0139] For example, after obtaining each similarity score, the similarities can be sorted in descending order. The similarity score ranked (n+1)th can be used as a preset similarity threshold. Then, the first n similarities are the multiple similarities greater than the preset similarity threshold, thus obtaining n auxiliary similarities.
[0140] The sample categories corresponding to the multiple obtained auxiliary similarities are each used as target categories. The similarity between the training image features corresponding to a sample image and the auxiliary category features of each of the multiple target categories is determined to obtain the corresponding training similarity. Based on the obtained training similarities and auxiliary similarities, the relation consistency loss of the image classification model is determined.
[0141] As an example, when determining the relation consistency loss of an image classification model based on multiple training similarities and multiple auxiliary similarities, we will take a sample image as an example. The situation is similar for each sample image, and will not be repeated here.
[0142] The average training similarity between multiple sample images and auxiliary category features of a target category is used as a training inter-class relationship between a sample category and a target category. The average auxiliary similarity between multiple sample images and auxiliary category features of a target category is used as an auxiliary inter-class relationship between a sample category and a target category. Based on the error between the obtained training inter-class relationships and auxiliary inter-class relationships, the relationship consistency loss L of the image classification model is determined. r Please refer to formula (2).
[0143]
[0144] in, This characterizes the auxiliary inter-class relationship between a sample category and n target categories calculated based on the i-th sample image. The training inter-class relationship between the sample category corresponding to the i-th sample image and the n target categories is represented by N, where N represents the number of sample images.
[0145] Relationship consistency loss L in image classification models r You can also add a margin term to impose stricter constraints on the loss of relationship consistency. Please refer to formula (3).
[0146]
[0147] As an example, after obtaining the feature consistency loss and the relation consistency loss, the weighted sum of the obtained feature consistency loss and relation consistency loss can be used as the training loss L of the image classification model, please refer to formula (4).
[0148] L = L f +L r (4)
[0149] Formula (4) represents a case where the weights of both feature consistency loss and relation consistency loss are 1.
[0150] If the obtained training loss fails to meet the training objective, adjust the model parameters of the image classification model until the obtained training loss meets the training objective. Then, output the image classification model as the trained target image classification model.
[0151] The following example uses a face image as a sample image, and an image classification model to determine the object corresponding to the face image, to illustrate the method for training an image classification model provided in this application.
[0152] Please refer to Figure 5 During the training phase, an auxiliary classification model can be trained first based on a sample image set. Then, combining the auxiliary image features and auxiliary category features output by the trained auxiliary classification model, the auxiliary relationship calculation module determines the auxiliary inter-class relationships between sample categories. Finally, combining the auxiliary image features and auxiliary category features output by the auxiliary classification model, as well as the auxiliary inter-class relationships output by the auxiliary relationship calculation module, an image classification model is trained to obtain the target image classification model. Therefore, during the usage phase, the target image classification model can be used to provide image classification services.
[0153] In this embodiment, a sample image set can be obtained according to the classification scenario. Based on the sample image set, the auxiliary classification model to be trained is iteratively trained multiple times to obtain a trained auxiliary classification model. For example, if the classification scenario is a face recognition scenario, then the sample image set is a set of sample images containing various face images. If there is no trained auxiliary classification model that matches the face recognition scenario, then an auxiliary classification model can be trained based on each face image. If a trained auxiliary classification model that matches the face recognition scenario already exists, then that auxiliary classification model can be used directly without needing to train it again. Before training the image classification model, prior knowledge can be obtained offline to improve the efficiency of training the image classification model.
[0154] The training process of the auxiliary classification model and the calculation process of the auxiliary relation calculation module can be referred to in the previous text, and will not be repeated here.
[0155] The training process of the image classification model is described in detail below. Please refer to [link / reference]. Figure 6 An image classification model may include a data preparation module, a feature extraction module, a training relationship calculation module, a feature consistency loss calculation module, a relationship consistency loss calculation module, and an optimization module.
[0156] The data preparation module can be used to read multiple sample images corresponding to each sample category. Taking one round of iterative training as an example, the data preparation module can read multiple sample images corresponding to one sample category, and the data preparation module can input each sample image into the feature extraction module in turn.
[0157] The feature extraction module contains fewer model parameters, significantly fewer than the parameters of the auxiliary classification model. It can be used to extract spatial features from images, obtaining spatial structure information. The feature extraction module can be structured as a convolutional neural network, for example, containing convolutional layers, non-linear activation function layers, and pooling layers. The feature extraction module can extract the training image features for each sample image. These features can then be input into the feature consistency loss calculation module. Simultaneously, the feature consistency loss calculation module can obtain the auxiliary image features corresponding to multiple sample images for that sample category from the auxiliary classification model.
[0158] The feature consistency loss calculation module can determine the feature consistency loss of the image classification model based on the errors between the obtained features of each training image and each auxiliary image.
[0159] The training relation calculation module can obtain the auxiliary image features of each sample image from the auxiliary relation calculation module, and the similarity between them and the auxiliary category features of each sample category. The training relation calculation module sorts the obtained similarities in descending order, selects the top n similarities, and uses them as auxiliary similarities, denoted as . The sample categories corresponding to each selected auxiliary similarity are then used as the target categories.
[0160] The feature extraction module can also input the obtained features of each training image into the training relationship calculation module. The training relationship calculation module can calculate the similarity between the training image features of each sample image and the auxiliary category features of the n target categories, and obtain the training similarity, denoted as .
[0161] The training relationship calculation module inputs the calculated training similarities and auxiliary similarities into the relationship consistency loss calculation module. Based on these similarities, the relationship consistency loss calculation module determines the training inter-class relationships and auxiliary inter-class relationships between the sample category and each target category. Based on the errors between these training inter-class relationships and auxiliary inter-class relationships, the relationship consistency loss of the image classification model is determined.
[0162] The feature consistency loss calculation module inputs the obtained feature consistency loss into the optimization module, and the relation consistency loss calculation module inputs the obtained relation consistency loss into the optimization module. The optimization module determines the training loss of the image classification model. If the training loss does not reach the training objective, it adjusts the model parameters of the image classification model based on algorithms such as gradient descent and enters the next round of iterative training until the obtained training loss reaches the training objective, such as when the number of iterations reaches a preset number, the training loss converges, or the training loss is less than a preset value. Finally, it outputs the trained target image classification model.
[0163] As one example, after obtaining a trained target image classification model, it can be used for image classification. Since the target image classification model contains fewer model parameters, the time required to use it is also shorter. Furthermore, because the target image classification model fully learns the image classification capabilities of the auxiliary classification model, it achieves high accuracy and reliability while maintaining short classification time.
[0164] When using a target image classification model for image classification, after obtaining the image to be classified, the model can be used to extract features from the image to obtain target image features. Based on these target image features, the target category of the image to be classified is determined from multiple sample categories using the target image classification model.
[0165] Please refer to Figure 7A When the image to be classified is a face image, a target image classification model with a small number of model parameters is used to extract features from the face image to obtain target image features. The target image classification model is then used again to classify the face image based on these target image features, determining the target category of the face image as "Xiao Hong," indicating that the face in the image belongs to "Xiao Hong."
[0166] Similarly, a trained auxiliary classification model containing a large number of model parameters is used to extract features from the face image to obtain auxiliary image features. The auxiliary classification model is then used again to classify the face image based on these auxiliary image features, determining the target category of the face image as "Xiao Hong," indicating that the face in the image belongs to "Xiao Hong."
[0167] The target image features extracted by the target image classification model and the auxiliary image features are quite similar. Therefore, the target category determined by the target image classification model and the auxiliary image classification model for the face image will also be the same. Since the number of parameters in the target image classification model is much smaller than that in the auxiliary image classification model, using the target image classification model can improve the accuracy and reliability of image classification while maintaining lower classification efficiency; alternatively, it can reduce the efficiency of image classification while maintaining high accuracy and reliability, making the target image classification model suitable for portable devices such as mobile devices.
[0168] Please refer to Figure 7B The target image classification model can include an image acquisition module, a feature extraction module, and an image classification module.
[0169] The image acquisition module acquires the image to be classified and sends it to the feature extraction module. The feature extraction module also extracts target image features from the image to be classified and sends these features to the image classification module. The image classification module determines the target category of the image based on these target image features. The feature extraction module utilizes knowledge transfer from a larger, more expressive auxiliary classification model. The feature distribution extracted by the feature extraction module is highly similar to the feature distribution extracted by the auxiliary classification model, thus resulting in higher classification accuracy for the target image classification model.
[0170] Based on the same inventive concept, embodiments of this application provide an apparatus for training an image classification model, capable of achieving the functions corresponding to the aforementioned method for training an image classification model. Please refer to... Figure 8 The device includes an acquisition module 801 and a processing module 802, wherein:
[0171] Acquisition module 801: used to acquire a set of sample images associated with multiple sample categories, with each sample category associated with multiple sample images;
[0172] Processing module 802: Used to extract features from each sample image using a trained auxiliary classification model to obtain corresponding auxiliary image features, and to perform the following operations for each sample category: perform feature fusion on the auxiliary image features of multiple sample images associated with a sample category to obtain corresponding auxiliary category features;
[0173] Processing module 802 is also used to: perform multiple rounds of iterative training on the image classification model to be trained based on the sample image set, and output the trained target image classification model, wherein each round of iteration includes:
[0174] The processing module 802 is also used to: use an image classification model to extract features from multiple sample images associated with a sample category to obtain the corresponding training image features;
[0175] The model parameters of the image classification model are adjusted based on the training image features and auxiliary image features corresponding to multiple sample images, as well as the auxiliary category features of multiple sample categories.
[0176] In one possible embodiment, the processing module 802 is specifically used for:
[0177] Based on the training image features and auxiliary image features corresponding to multiple sample images, the feature consistency loss of the image classification model is determined. The feature consistency loss represents the consistency of feature extraction using the auxiliary classification model and the image classification model.
[0178] Based on the training image features and auxiliary image features corresponding to multiple sample images, as well as the auxiliary category features of multiple sample categories, the relation consistency loss of the image classification model is determined. The relation consistency loss represents the consistency of the inter-class relations between multiple sample categories determined by the auxiliary classification model and the image classification model.
[0179] Based on the obtained feature consistency loss and relation consistency loss, the model parameters of the image classification model are adjusted.
[0180] In one possible embodiment, the processing module 802 is specifically used for:
[0181] For multiple sample images, perform the following operations respectively:
[0182] Multiple sample categories are used as target categories. The similarity between the training image features corresponding to a sample image and the auxiliary category features of each of the multiple target categories is determined to obtain the corresponding training similarity.
[0183] The auxiliary image features corresponding to a sample image are determined, and the similarity between them and the auxiliary category features of various target categories is obtained to obtain the corresponding auxiliary similarity.
[0184] Based on the obtained multiple training similarities and multiple auxiliary similarities, the relation consistency loss of the image classification model is determined.
[0185] In one possible embodiment, the processing module 802 is specifically used for:
[0186] For multiple sample images, perform the following operations respectively:
[0187] The similarity between the auxiliary image features corresponding to a sample image and the auxiliary category features of each of the multiple sample categories is determined. Among the obtained similarities, the multiple similarities that are greater than the preset similarity threshold are respectively used as auxiliary similarities, and the sample categories corresponding to the multiple obtained auxiliary similarities are respectively used as target categories.
[0188] The similarity between the training image features corresponding to a sample image and the auxiliary category features of each of the multiple target categories is determined to obtain the corresponding training similarity.
[0189] Based on the obtained multiple training similarities and multiple auxiliary similarities, the relation consistency loss of the image classification model is determined.
[0190] In one possible embodiment, the processing module 802 is specifically used for:
[0191] For different target categories, perform the following operations respectively:
[0192] The average training similarity between multiple sample images and auxiliary category features of a target category is used as a training inter-class relationship between a sample category and a target category.
[0193] The average of the auxiliary similarities between multiple sample images and auxiliary category features of a target category is used as an auxiliary inter-class relationship between a sample category and a target category.
[0194] Based on the errors between multiple training class relationships and multiple auxiliary class relationships obtained, the relationship consistency loss of the image classification model is determined.
[0195] In one possible embodiment, the processing module 802 is specifically used for:
[0196] The errors between the training image features and auxiliary image features corresponding to each of the multiple sample images are determined to obtain the corresponding feature consistency errors.
[0197] The feature consistency loss of the image classification model is determined based on the weighted average of the obtained feature consistency errors.
[0198] In one possible embodiment, the processing module 802 is specifically used for:
[0199] The weighted sum of the obtained feature consistency loss and relation consistency loss is used as the training loss of the image classification model.
[0200] If the obtained training loss fails to achieve the training objective, adjust the model parameters of the image classification model.
[0201] In one possible embodiment, the processing module 802 is further configured to:
[0202] Obtain the image to be classified;
[0203] A target image classification model is used to extract features from the image to be classified, thereby obtaining the target image features;
[0204] A target image classification model is used to determine the target category of the image to be classified from multiple sample categories based on the target image features.
[0205] Please refer to Figure 9 The aforementioned apparatus for training the image classification model can run on a computer device 900. The current and historical versions of the data storage program, as well as the application software corresponding to the data storage program, can be installed on the computer device 900, which includes a processor 980 and a memory 920. In some embodiments, the computer device 900 may include a display unit 940, which includes a display panel 941 for displaying a user-interactive interface, etc.
[0206] In one possible embodiment, the display panel 941 may be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED).
[0207] The processor 980 is used to read a computer program and then execute the methods defined by the computer program. For example, the processor 980 reads a data storage program or file, thereby running the data storage program on the computer device 900 and displaying the corresponding interface on the display unit 940. The processor 980 may include one or more general-purpose processors, and may also include one or more DSPs (Digital Signal Processors) for performing related operations to implement the technical solutions provided in the embodiments of this application.
[0208] The memory 920 generally includes main memory and secondary storage. Main memory can be random access memory (RAM), read-only memory (ROM), and cache, etc. Secondary storage can be a hard disk, optical disk, USB flash drive, floppy disk, or magnetic tape drive, etc. The memory 920 is used to store computer programs and other data. The computer programs include applications corresponding to each client, and other data may include data generated after the operating system or applications are run, including system data (e.g., operating system configuration parameters) and user data. In this embodiment, program instructions are stored in the memory 920, and the processor 980 executes the program instructions in the memory 920 to implement any of the methods described in the preceding figures.
[0209] The aforementioned display unit 940 is used to receive input digital information, character information, or contact touch operations / non-contact gestures, and to generate signal inputs related to user settings and function control of the computer device 900. Specifically, in this embodiment, the display unit 940 may include a display panel 941. The display panel 941, for example, is a touch screen, which can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or on the display panel 941), and drive corresponding connection devices according to a pre-set program.
[0210] In one possible embodiment, the display panel 941 may include two parts: a touch detection device and a touch controller. The touch detection device detects the player's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 980. It can also receive and execute commands from the processor 980.
[0211] The display panel 941 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 940, in some embodiments, the computer device 900 may also include an input unit 930. The input unit 930 may include an image input device 931 and other input devices 932, wherein the other input devices may include, but are not limited to, one or more of the following: a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick.
[0212] In addition to the above, the computer device 900 may also include a power supply 990 for powering other modules, an audio circuit 960, a near-field communication module 970, and an RF circuit 910. The computer device 900 may also include one or more sensors 950, such as an accelerometer, a light sensor, and a pressure sensor. The audio circuit 960 specifically includes a speaker 961 and a microphone 962, for example, the computer device 900 can use the microphone 962 to collect the user's voice and perform corresponding operations.
[0213] As one embodiment, the number of processors 980 can be one or more, and the processors 980 and the memory 920 can be coupled together or relatively independent.
[0214] As one example, Figure 9 The processor 980 in the middle can be used to implement, for example Figure 8 The functions of the acquisition module 801 and the processing module 802 in the process.
[0215] As one example, Figure 9 The processor 980 in the text can be used to implement the functions of the server or terminal devices discussed above.
[0216] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0217] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of software products, for example, through a computer program product. This computer program product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
[0218] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for training an image classification model, characterized in that, include: Obtain a set of sample images associated with multiple sample categories, with each sample category associated with multiple sample images; The trained auxiliary classification model is used to extract features from each sample image to obtain the corresponding auxiliary image features. For each sample category, the following operations are performed: the auxiliary image features of multiple sample images associated with a sample category are fused to obtain the corresponding auxiliary category features. Based on the sample image set, the image classification model to be trained is subjected to multiple rounds of iterative training, and the trained target image classification model is output. Each round of iteration includes: Using the image classification model, feature extraction is performed on multiple sample images associated with a sample category to obtain the corresponding training image features; Based on the training image features and auxiliary image features corresponding to each of the multiple sample images, the feature consistency loss of the image classification model is determined, wherein the feature consistency loss characterizes the consistency of feature extraction using the auxiliary classification model and the image classification model; Based on the training image features and auxiliary image features corresponding to each of the multiple sample images, and the auxiliary category features of each of the multiple sample categories, the relation consistency loss of the image classification model is determined, wherein the relation consistency loss represents the consistency of the inter-class relations between the multiple sample categories determined by the auxiliary classification model and the image classification model; Based on the obtained feature consistency loss and relation consistency loss, the model parameters of the image classification model are adjusted.
2. The method according to claim 1, characterized in that, The determination of the relationship consistency loss of the image classification model based on the training image features and auxiliary image features corresponding to each of the multiple sample images, and the auxiliary category features of each of the multiple sample categories, includes: For each of the multiple sample images, perform the following operations: The various sample categories are used as target categories, and the similarity between the training image features corresponding to a sample image and the auxiliary category features of each of the various target categories is determined to obtain the corresponding training similarity. The similarity between the auxiliary image features corresponding to each sample image and the auxiliary category features of each of the multiple target categories is determined to obtain the corresponding auxiliary similarity. Based on the obtained multiple training similarities and multiple auxiliary similarities, the relation consistency loss of the image classification model is determined.
3. The method according to claim 1, characterized in that, The determination of the relationship consistency loss of the image classification model based on the training image features and auxiliary image features corresponding to each of the multiple sample images, and the auxiliary category features of each of the multiple sample categories, includes: For each of the multiple sample images, perform the following operations: The similarity between the auxiliary image features corresponding to a sample image and the auxiliary category features of each of the multiple sample categories is determined. Among the obtained similarities, the multiple similarities that are greater than a preset similarity threshold are respectively used as auxiliary similarities, and the sample categories corresponding to the multiple obtained auxiliary similarities are respectively used as target categories. The similarity between the training image features corresponding to each sample image and the auxiliary category features of each of the multiple target categories is determined to obtain the corresponding training similarity. Based on the obtained multiple training similarities and multiple auxiliary similarities, the relation consistency loss of the image classification model is determined.
4. The method according to claim 2 or 3, characterized in that, The determination of the relation consistency loss of the image classification model based on multiple obtained training similarities and multiple auxiliary similarities includes: For each of the aforementioned target categories, perform the following operations respectively: The average training similarity between the multiple sample images and the auxiliary category features of a target category is used as the training inter-class relationship between the sample category and the target category. The average value of the auxiliary similarities between the plurality of sample images and the auxiliary category features of the target category is taken as the auxiliary inter-class relationship between the sample category and the target category. Based on the errors between multiple training class relationships and multiple auxiliary class relationships obtained, the relationship consistency loss of the image classification model is determined.
5. The method according to claim 1, characterized in that, The step of determining the feature consistency loss of the image classification model based on the training image features and auxiliary image features corresponding to each of the multiple sample images includes: The errors between the training image features and auxiliary image features corresponding to each of the multiple sample images are determined to obtain the corresponding feature consistency errors; The feature consistency loss of the image classification model is determined based on the weighted average of the obtained feature consistency errors.
6. The method according to claim 1, characterized in that, The adjustment of the model parameters of the image classification model based on the obtained feature consistency loss and relation consistency loss includes: The weighted sum of the obtained feature consistency loss and relation consistency loss is used as the training loss of the image classification model. If the obtained training loss fails to achieve the training objective, the model parameters of the image classification model are adjusted.
7. The method according to any one of claims 1 to 6, characterized in that, After performing multiple rounds of iterative training on the image classification model to be trained based on the sample image set, and outputting the trained target image classification model, the process further includes: Obtain the image to be classified; The target image classification model is used to extract features from the image to be classified to obtain target image features; Using the target image classification model, the target category of the image to be classified is determined from the multiple sample categories based on the target image features.
8. An apparatus for training an image classification model, characterized in that, include: Acquisition module: Used to acquire sample image sets associated with multiple sample categories, with each sample category associated with multiple sample images; Processing module: Used to extract features from each sample image using a trained auxiliary classification model to obtain corresponding auxiliary image features, and to perform the following operations for each sample category: perform feature fusion on the auxiliary image features of multiple sample images associated with a sample category to obtain corresponding auxiliary category features; The processing module is further configured to: perform multiple rounds of iterative training on the image classification model to be trained based on the sample image set, and output the trained target image classification model, wherein each round of iteration includes: The processing module is further configured to: use the image classification model to extract features from multiple sample images associated with a sample category to obtain corresponding training image features; Based on the training image features and auxiliary image features corresponding to each of the multiple sample images, the feature consistency loss of the image classification model is determined, wherein the feature consistency loss characterizes the consistency of feature extraction using the auxiliary classification model and the image classification model; Based on the training image features and auxiliary image features corresponding to each of the multiple sample images, and the auxiliary category features of each of the multiple sample categories, the relation consistency loss of the image classification model is determined, wherein the relation consistency loss represents the consistency of the inter-class relations between the multiple sample categories determined by the auxiliary classification model and the image classification model; Based on the obtained feature consistency loss and relation consistency loss, the model parameters of the image classification model are adjusted.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method as described in any one of claims 1 to 7.
10. A computer device, characterized in that, include: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the method as described in any one of claims 1 to 7 according to the obtained program instructions.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Image category identification method and device based on model distillation, storage medium and terminal
CN113408570A
Fine-grained image classification method and device, storage medium and terminal
CN113836338A