Training method, classification method, device and related device of image classification model
By acquiring the original images and categories of edge-end devices, using crawler technology to expand samples and combining deep residual network training models, the problem that existing image classification models cannot meet the needs of edge-end devices to acquire image classification, and the accuracy of image classification and iterative optimization of the model are achieved.
Patent Information
- Application Number
- CN202310147722.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-02-20
AI Technical Summary
In the cloud-edge collaboration scenario, the existing image classification model trained based on public data sources cannot meet the classification needs of edge devices to acquire images, resulting in low accuracy of classification results.
By obtaining the original image and its categories of edge-end devices, crawling technology expands the samples, and combining deep residual network (ResNet) model training, an image classification model that fits the actual scene at the edge-end is generated.
Improve the accuracy of image classification, form forward loop iteration, and can update the model according to actual scene changes and improve classification accuracy.
Smart Images

Figure CN116310518B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to the field of deep learning technology, and specifically to a training method, a classification method, a device, and related apparatuses for an image classification model. Background Art
[0002] In a cloud-edge collaboration scenario, edge devices usually collect image data and upload it to a database in the cloud. A management device in the cloud classifies the images in the database based on a classification model. Among them, the classification model is usually trained based on public data sources such as ImageNet. Summary of the Invention
[0003] The present disclosure provides a training method, a classification method, a device, and related apparatuses for an image classification model.
[0004] According to one aspect of the present disclosure, there is provided a training method for an image classification model, the method comprising:
[0005] Obtaining original training samples; wherein the original training samples include: at least one original image from an edge device managed by a management device, and the category of the at least one original image;
[0006] Performing web crawling according to the category of the at least one original image to obtain target training samples; wherein the target training samples include at least one target image and the category of the at least one target image; the category of the target image obtained by performing web crawling based on the category of one original image is the same as the category of the original image;
[0007] Training a first model based on the original training samples and the target training samples to obtain a second model; wherein the first model is an image classification model trained based on images in a public data source; the second model is used to classify images collected by an edge device.
[0008] According to another aspect of the present disclosure, there is provided an image classification method, the method comprising:
[0009] Obtaining a second model; wherein the second model is a model obtained by training a first model based on original training samples and target training samples; the first model is an image classification model trained based on images in a public data source; the original training samples include: at least one original image from an edge device managed by a management device, and the category of the at least one original image; the target training samples include: at least one target image obtained by performing web crawling based on the category of the at least one original image, and the category of the at least one target image; wherein the category of the target image obtained by performing web crawling based on the category of one original image is the same as the category of the original image;
[0010] Classify the images collected by the edge device based on the second model.
[0011] According to another aspect of the present disclosure, there is provided a training device for an image classification model, the device comprising:
[0012] An acquisition unit for acquiring original training samples; wherein the original training samples include: at least one original image from an edge device, and the category of at least one original image;
[0013] An augmentation unit for performing web crawling according to the category of at least one original image to obtain target training samples; wherein the target training samples include at least one target image and the category of at least one target image; the category of the target image obtained by performing web crawling based on the category of one original image is the same as the category of this original image;
[0014] A training unit for training a first model based on the original training samples and the target training samples to obtain a second model; wherein the first model is an image classification model trained based on the images in the public data source; the second model is used for classifying the images collected by the edge device.
[0015] According to another aspect of the present disclosure, there is provided an image classification device, the device comprising:
[0016] An acquisition unit for acquiring a second model; wherein the second model is a model obtained by training a first model based on the original training samples and the target training samples; wherein the first model is an image classification model trained based on the images in the public data source; the original training samples include: at least one original image from an edge device, and the category of at least one original image; the target training samples include: at least one target image obtained by performing web crawling based on the category of at least one original image, and the category of at least one target image; wherein the category of the target image obtained by performing web crawling based on the category of one original image is the same as the category of this original image;
[0017] A classification unit for classifying the images collected by the edge device based on the second model.
[0018] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0019] At least one processor; and
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores instructions executable by the at least one processor. When executed by the at least one processor, the instructions enable the at least one processor to execute the training method or the image classification method of the image classification model provided by the present disclosure.
[0022] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the training method or the image classification method of the image classification model provided by the present disclosure.
[0023] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the training method or the image classification method of the image classification model provided by the present disclosure.
[0024] The technical solution provided by the present disclosure obtains original training samples, and based on the original training samples and in combination with web crawler technology, obtains target training samples. Using the original training samples and the target training samples as sample data to implement the training of the model helps to generate an image classification model that fits the actual scenario of the images collected at the edge side. When performing image classification based on this image classification model, the accuracy of the classification result is improved. Further, according to the training method of this image classification model, iterative optimization of this image classification model can be achieved based on the actually collected image data, forming a positive cyclic iteration, which helps to update the image classification model following the changes in the actual scenario, and thus improves the accuracy of image classification.
[0025] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0027] Figure 1 FIG. is a schematic diagram of a cloud-edge collaboration scenario shown in an embodiment of the present disclosure;
[0028] Figure 2 FIG. is a schematic diagram of a cloud-edge collaboration scenario shown in an embodiment of the present disclosure;
[0029] Figure 3 FIG. is a schematic flowchart of a method for training an image classification model shown in an embodiment of the present disclosure;
[0030] Figure 4a FIG. is a schematic diagram of a statistical distribution shown in an embodiment of the present disclosure;
[0031] Figure 4bIt is a schematic diagram of a statistical distribution based on a first model shown in an embodiment of the present disclosure;
[0032] Figure 5 It is a schematic diagram of an original image and a target image shown in an embodiment of the present disclosure;
[0033] Figure 6 It is a schematic diagram of a comparison result of image deblurring shown in an embodiment of the present disclosure;
[0034] Figure 7 It is a schematic diagram of a training process of an image classification model shown in an embodiment of the present disclosure;
[0035] Figure 8 It is a schematic diagram of a process of an image classification method shown in an embodiment of the present disclosure;
[0036] Figure 9 It is a block diagram of a structure of a training device of an image classification model shown in an embodiment of the present disclosure;
[0037] Figure 10 It is a block diagram of a structure of an image classification device shown in an embodiment of the present disclosure;
[0038] Figure 11 It is a block diagram of an electronic device for implementing a training method of an image classification model or an image classification method in an embodiment of the present disclosure. Specific embodiments
[0039] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0040] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information and other processes all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0041] Such as Figure 1As shown in the figure, it is a schematic diagram of a cloud-edge collaboration scenario provided by an embodiment of the present application. Among them, the cloud includes a management device 10 and a database 11, and the edge side includes edge devices 12, 13, and 14. The edge devices (such as edge devices 12, 13, and 14) are used to collect data and upload it to the database 11 in the cloud, and the management device 10 is used to process the data in the database 11. Further, the management device 11 can perform tasks such as prediction and classification in real time according to the collected data and send the results to the edge devices; or, the management device 10 sends the task model to the edge devices to instruct the edge devices to perform data processing tasks.
[0042] Among them, during the process of the edge device uploading data to the cloud, the data traffic is uncontrollable, and a large amount of data may be uploaded simultaneously, and the management device 10 may not be able to process it in time. In this regard, the embodiment of the present application also provides a schematic diagram of a scenario. As Figure 2 shown, this scenario includes: edge devices (such as edge devices 12, 13, and 14), a message queue manager, and a management device. Among them, the edge device may include: a sensor and a gateway device. The sensor is used to collect data, and the gateway device is responsible for processing and reporting the data.
[0043] Specifically, the gateway device may be an industrial computer, an edge gateway, etc. The gateway device can use a high-performance non-blocking communication framework, such as the netty communication framework, and connect to the sensor based on the transmission control protocol (TCP). The sensor can upload the collected data to the gateway device based on the hyper text transfer protocol (HTTP), and the gateway device can communicate based on the TCP socket (TCPSocket).
[0044] In the link from the gateway device to the management device, if communication is through HTTP or TCP, when the link establishment fails, data may be lost. Therefore, a message queue manager is used to ensure the reliability of data transmission. For example, the message queue manager balances data transmission through RabbitMQ. Specifically, the message queue manager receives the data sent by the gateway device and pushes it to the management device in a cached batch manner, so that the data is transmitted in a relatively stable form. Similarly, the management device issues control instructions to the message queue manager, and divides resources through the message queue, and can plan the mapping relationship between the gateway device and the message queue, and can quickly locate problems when they occur.
[0045] The above management device or edge device can be an independent communication device. This application does not limit the specific form of the communication device, which can be a terminal or a server. Among them, the terminal can specifically be a mobile phone, an augmented reality (AR) device, a virtual reality (VR) device, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The server can be a physical or logical server. When the communication device is an edge device, it can specifically be a sensor.
[0046] Currently, the management device usually performs image classification based on a pre-trained model obtained by training on public data sources such as ImageNet. Since the image categories in the public data source are limited and may not cover the categories of the images collected by the edge device, the current classification model does not meet the classification requirements of the collected images, and thus the accuracy of the classification results obtained by performing image classification based on the pre-trained model is relatively low.
[0047] In response to this, the embodiment of this application provides a method for training an image classification model, so that the trained image classification model fits the actual application scenario and improves the accuracy of image classification. This method is applied to the management device in the cloud, such as Figure 3 as shown, and includes steps S301 - S303.
[0048] S301. The management device obtains the original training samples.
[0049] Among them, the original training samples include: at least one original image from the edge device managed by the management device in the cloud, and the category of at least one original image.
[0050] In a possible implementation, the management device obtains at least one original image through the following steps S11 - S13.
[0051] S11. The management device obtains at least one historical image collected by the edge device.
[0052] Among them, the historical image refers to the image collected and uploaded by the edge device in the current database.
[0053] S12. The management device sorts at least one historical image in reverse order according to the collection time to obtain a historical image sequence.
[0054] It can be understood that in the historical image sequence sorted in reverse order of the acquisition time, the most recently acquired and uploaded image is at the front of the sequence. Based on this sorting method, it is helpful to obtain the images acquired in the most recent time period.
[0055] S13. The management device selects at least one image from the first preset number of historical images in the historical image sequence as at least one original image.
[0056] Among them, the acquisition times of the first preset number of historical images in the historical image sequence are relatively close to the current moment. Specifically, the acquisition times of the first preset number of historical images in the historical image sequence are closer to the current moment than the acquisition times of the remaining historical images in the historical image sequence. Exemplarily, if the current historical image sequence consists of 500 historical images and 100 original images need to be selected, the management device can select 100 images from the first 300 images in this historical image sequence.
[0057] It can be understood that in order to approximate the actual distribution of the categories of the acquired images as much as possible, the management device can randomly select at least one original image from the first preset number of historical images in the historical image sequence.
[0058] It should be noted that this application does not limit the number of original images selected. In one example, the number of the original images can be the first preset percentage of the number of historical images in the above database. For example, the first preset percentage is 5%. Combining the above example, if there are 500 images in the historical image sequence, then according to the first preset percentage, the number of original images to be selected is 25.
[0059] It should be noted that the methods and quantities for obtaining the original images in the above steps S11 - S13 are only examples. In actual applications, the original images can also be determined in other ways, and this is not limited.
[0060] In one possible implementation, the categories of the above at least one original image can be obtained by manual annotation. Among them, manual annotation means that technicians determine image information, such as the category of the image, by observing the image and perform annotation. Among them, technicians can annotate the categories of the at least one original image on the management device, or can also annotate the categories of the at least one original image through other computing devices. Based on this, the management device can obtain the categories of the at least one original image obtained by manual annotation. For example, technicians can upload the original images with annotated categories to the database in the cloud for the management device to read, or directly upload them to the management device.
[0061] It is understandable that the original training samples are used as training data to obtain an image classification model. To make the image classification model fit the actual scenario of the currently acquired images, through manual annotation, it is helpful to obtain the actual categories of the original images in the original training samples.
[0062] In another possible implementation, the management device can also obtain the categories of at least one original image in other ways, which is not limited here. For example, through machine learning and other means, the category of the original image is determined based on the original image. Or, on the basis of obtaining the category of the original image through a machine learning algorithm, the result obtained by the machine learning algorithm is further checked by manual annotation to reduce the annotation workload and improve work efficiency.
[0063] S302. The management device crawls to obtain target training samples according to the categories of at least one original image.
[0064] Among them, the target training samples include at least one target image and the category of at least one target image; the category of the target image obtained by crawling based on the category of one original image is the same as the category of the original image.
[0065] It is understandable that crawling based on the categories of at least one original image is essentially to perform sample augmentation for the categories of at least one original image so that the number of augmented images meets the requirements of the sample quantity.
[0066] In a possible implementation, the above crawling specifically uses the multi-threaded thread module and lock based on Python to implement a multi-concurrent image crawler.
[0067] Optionally, when at least one original image meets one or more of the following conditions, sample augmentation is performed based on the category of the at least one original image.
[0068] Condition 1: Among at least one original image, the number of original images of the same category is greater than or equal to the first threshold.
[0069] Among them, the first threshold is a preset threshold for the number of images, and its specific value can be adjusted according to the actual application scenario. Exemplarily, assuming there are images of 5 categories and the first threshold is 50, then the categories in which the number of images included in the 5 categories is greater than or equal to 50 are used as target categories, and crawling is performed based on the target categories to obtain target images. Among them, the target category refers to the category used for crawling.
[0070] Specifically, in combination with Condition 1, the above step S302 is implemented through the following steps S21 - S23.
[0071] S21. The management device counts the number of initial images in each current category.
[0072] Among them, the initial images in the current respective categories can be the original images determined based on the historical image sequence in the above step S301.
[0073] Exemplarily, as Figure 4a shown, it is a schematic diagram of a statistical distribution provided by an embodiment of the present application.
[0074] S22. The management device determines, among the respective categories, target categories in which the number of initial images is greater than or equal to a first threshold.
[0075] Combined with the above example, Figure 4a shows four categories of clocks, cans, sunflowers, and calendars, as well as the distribution of the number of initial images in each category. If the first threshold is 50, then in Figure 6 the target categories in which the number of initial images is greater than or equal to the first threshold include clocks and cans.
[0076] S23. Crawl based on the target categories to obtain at least one target image.
[0077] Combined with the above example, crawl based on the two categories of clocks and cans to expand the images including clocks and cans.
[0078] Such as Figure 4b shown, it is a schematic diagram of a statistical distribution based on a first model provided by an embodiment of the present application. Among them, the first model is an image classification model trained based on the images in the public data source. It can be seen that the number of blurred images accounts for a relatively large proportion in all image categories. That is to say, based on the current image classification model, the image categories cannot be accurately obtained, and thus the actual distribution of the image categories cannot be accurately obtained. The original images obtained through the above step S301, and obtaining the categories of this part of the original images manually annotated, and then based on the above steps S21 - S23, helps to obtain the actual distribution of the image categories to achieve sample expansion.
[0079] It can be understood that through the above condition one, it helps to determine the categories whose image numbers meet the above conditions as target categories for sample expansion. This method helps to improve the accuracy of feature extraction for the target categories during subsequent training of the model, and further improve the accuracy of the image classification model for identifying images of the target categories.
[0080] Condition two: The clarity of at least one original image is less than or equal to a second threshold.
[0081] Among them, the second threshold is a preset threshold for clarity, and its specific value can be adjusted according to the actual application scenario.
[0082] In a possible implementation, the clarity of the image can be manually annotated. After the technician annotates the clarity of the image, the image with the annotated clarity can be uploaded to the management device so that the management device can obtain the image including the clarity. Exemplarily, in the above step S301, when the technician annotates the category of the image, the clarity of the image can also be manually annotated. Specifically, "0" can be used to indicate clarity and "1" can be used to indicate lack of clarity. Further, the management device can obtain the clarity of the image according to the information manually annotated for indicating the clarity of the image.
[0083] In another possible implementation, the clarity of the image is obtained by the computing device identifying the image information. For example, the clarity of the image can be obtained through the gradient difference of the gray-scale features between adjacent pixels, or can be distinguished through the frequency components of the image. Among them, the second threshold can be adjusted according to the specific algorithm, and no limitation is imposed thereon.
[0084] Specifically, in combination with condition two, the above step S302 is implemented through the following steps S31 - S33.
[0085] S31. The management device obtains the clarity of the initial image.
[0086] Among them, the initial image can be the original image determined based on the historical image sequence in the above step S301.
[0087] S32. The management device takes the initial images with clarity less than or equal to the second threshold as at least one original image.
[0088] S33. The management device crawls based on the categories of the above at least one original image to obtain at least one target image.
[0089] Exemplarily, the management device can obtain the images with clarity less than or equal to the second threshold, and then determine the categories of the images with clarity less than or equal to the second threshold, that is, obtain the categories of at least one original image. The management device can crawl based on this category for sample expansion. Further, the management device can also combine the number of images to judge the number of images with clarity less than or equal to the second threshold in each category. For example, count the number of images with lower clarity in each category, take the categories with the number of images greater than or equal to the preset threshold as the target categories, and perform sample expansion based on the target categories.
[0090] As Figure 5 shown, it is a schematic diagram of an original image and a target image. Among them, Figure 5 in the figure (a) is the original image, and the figure (b) is the target image obtained by crawling according to the category of the figure (a).
[0091] It can be understood that through the above method, it helps to expand the samples for the category to which the images with lower clarity in the original training samples belong, avoiding the problem that the model cannot accurately extract features based on blurred images during training and affecting the subsequent image classification accuracy. When expanding samples through web crawling, the images crawled from the network are usually clearer, which helps to improve the accuracy of feature extraction.
[0092] Condition three: The accuracy of at least one original image is less than or equal to a third threshold.
[0093] Among them, the accuracy of at least one original image is the classification result obtained by classifying at least one original image using a first model, compared with the accuracy of the category of at least one original image; the category of at least one original image is the category after manual annotation. The first model is an image classification model trained based on images in a public data source.
[0094] Among them, the third threshold is a preset threshold for accuracy, and its specific value can be determined according to the similarity between the classification result obtained based on the first model and the category after manual annotation. Exemplarily, in step S301 above, at least one original image is obtained based on historical images in a database. Optionally, the database may also include the results of image classification of the above historical images based on the first model. By comparing the image category obtained through manual annotation with the image category obtained based on the first model, if they are different, it is considered that the accuracy is low; if they are the same or similar, it is considered that the accuracy is high.
[0095] Specifically, in combination with condition three, step S302 is implemented through the following steps S41 - S42.
[0096] S41. The management device obtains the classification results of the initial images in each current category based on the first model.
[0097] S42. The management device determines the similarity between the classification results of the above initial images based on the first model and the categories after manual annotation.
[0098] Exemplarily, the above similarity can be directly determined by a technician based on the classification result of the first model through manual annotation. Further, the management device can obtain the similarity between the classification result of the initial image based on the first model and the manually annotated category by acquiring the information judged by the technician. For example, for an image, the classification result of the first model is "computer". The technician observes the image to determine whether the classification result is accurate. If it is accurate, it means that the similarity between the classification result based on the first model and the manually annotated category is high; otherwise, the similarity is low. Alternatively, it can also be obtained by calculating the device to calculate the similarity between the manually annotated category and the above classification result. Specifically, the management device determines the semantic similarity between the classification result and the manually annotated category. For example, when the manually annotated image category is "computer" and the classification result based on the first model is "computer", the management device determines that the similarity is high through semantic judgment between the two; while when the manually annotated image category is "table tennis" and the classification result based on the first model is "football", the management device determines that the similarity is low through semantic judgment between the two.
[0099] S43. The management device uses the images with the above similarity less than or equal to the third threshold as at least one original image.
[0100] S44. The management device performs web crawling based on the categories of the above at least one original image to obtain at least one target image.
[0101] Combined with the above example, the images with a low similarity between the above classification result and the manually annotated category are used as at least one original image, and sample augmentation is performed based on the categories of the at least one original image.
[0102] It can be understood that through the above method, it helps to perform sample augmentation for the categories to which the images with low accuracy of the classification results obtained based on the first model in the original training samples belong, so as to extract more features of the images of this category based on the augmented samples and improve the accuracy of the image classification model.
[0103] It should be noted that in practical applications, any one or more of the above three conditions can be selected to determine the categories for which web crawling is to be performed. In addition, the above three conditions can be screened one by one in the order of priority. For example, when condition one is satisfied, conditions two and three are further considered, and there is no restriction on this.
[0104] S303. The management device trains the first model based on the original training samples and the target training samples to obtain a second model; wherein, the first model is an image classification model trained based on the images in the public data source; the second model is used to classify the images collected by the edge device.
[0105] In a possible implementation, the above step S303 is based on the original training samples and the target training samples, and the first model is trained using a deep residual network (ResNet) model, and the trained model is used as the second model.
[0106] Among them, the ResNet model is a network model developed based on convolutional neural networks (CNN). The CNN model includes an input layer, feature extraction, a fully connected layer, and an output layer. Among them, the feature extraction layer includes a convolutional layer, an activation function, and a pooling layer. Compared with the CNN model, the ResNet model includes many units with similar structures in the feature extraction layer, and these units are called Residual Blocks. Specifically, the common point of each Residual Block is the existence of a shortcut for cross-layer direct connection. It can be understood that the main idea of the neural network is to gradually extract features from the underlying features to highly abstract features. The more layers the network has, the more helpful it is to extract richer and different levels of abstract features. However, in the CNN model, as the number of network layers increases, it may lead to gradient disappearance. In this regard, the shortcut in the ResNet model helps to solve this problem. Based on the ResNet model, the above step S303 can specifically re-initialize the last Residual Block, and then, according to the category of the above image, use some images in the original training samples and the target training samples as the training set for training. The remaining images are used as the validation set to adjust the parameters of the trained ResNet model to obtain an image classification model.
[0107] In a possible implementation, the images in the above original training samples and target training samples are subjected to image enhancement to obtain at least one processed original image and at least one target image, which are used as a sample image set, and a model is trained based on this sample image set. Among them, the image enhancement includes any one or more of random cropping, rotation, scaling, and deformation.
[0108] It can be understood that image enhancement helps to improve the accuracy of feature extraction, thereby improving the robustness of the model.
[0109] Optionally, before performing the above step S303, the method of the present disclosure further includes: the management device deletes duplicate images in at least one target image to obtain at least one updated target image.
[0110] It can be understood that there may be duplicate images in the target images obtained by the above-mentioned step S302 based on the original image crawler. When exactly the same images exist in the training set and the validation set, it will cause pseudo-high accuracy of the model. In this regard, the duplicate images after crawling can be screened through the above-mentioned method to achieve image deduplication.
[0111] In a possible implementation, the management device calculates the similarity between any two target images, and deletes any one of the two target images when the similarity meets the preset conditions.
[0112] Optionally, the above similarity is judged based on the mean squared error (MSE) between images. Specifically, the target images are converted into the same size, such as 300x300; and then the MSE is calculated based on the following formula:
[0113]
[0114] Furthermore, the images with MSE lower than 0.1 are used as similar image pairs, and one of them is randomly deleted.
[0115] Optionally, the above similarity calculation can be performed separately according to the categories of the target images. It can be understood that the similarity between images of different categories is relatively low. Therefore, by calculating according to the image categories, it helps to reduce the computational workload of the management device and save computing resources.
[0116] Through the above method, it helps to perform deduplication processing of the same images by converting the size and then calculating the MSE, avoiding exactly the same or duplicate images with different sizes but the same ratio in the target images.
[0117] It should be noted that the above image deduplication process can also be applied to the processing process of the original images. Among the images collected by the edge device, there may also be images with relatively high similarity. In this regard, the image deduplication can be performed according to the above method. Among them, in order to reduce the computational workload, the management device can perform image deduplication on the original images and the target images separately, and further perform image deduplication according to the image categories respectively.
[0118] It should be noted that the above method for judging the similarity between images is only an example. In actual application scenarios, it can also be judged by other methods, and this is not limited.
[0119] Optionally, before performing the above step S303, the method of the present disclosure may further include: the management device performs image deblurring on at least one original image and at least one target image.
[0120] In a possible implementation, the management device uses the DeblurGAN-v2 model to deblur the target sample image.
[0121] It can be understood that due to hardware cost limitations, current edge devices have low resolution and poor anti-shake functions, resulting in low clarity of the captured images, which is not convenient for direct model training. In this regard, the management device can deblur the sample images to improve image clarity and further enhance the accuracy of image classification. In addition, the DeblurGAN-v2 model can be fine-tuned for the sample images to make the model more suitable for deblurring the current sample image set.
[0122] Exemplarily, as Figure 6 shown, it is a schematic diagram of the comparison results of image deblurring. Among them, Figure 6 Figure (a) in it is the blurred image, and Figure (b) is the deblurred image.
[0123] It should be noted that the model used for the above image deblurring is only an example. In actual application scenarios, it can also be implemented in other ways, and there is no limitation in this regard.
[0124] Such as Figure 7 shown, it is a schematic diagram of the training process of an image classification model. It includes four parts: manual annotation, sample augmentation, model training, and establishing a model application programming interface (API). Manual annotation is used to annotate the images captured by the edge device, where the marking parameters include image category, image clarity, and image accuracy. After annotation, seed samples are obtained. This part is the specific implementation of the above step S301. The seed samples obtained through manual annotation include at least one of the above original images and the category of at least one original image. Sample augmentation is used to crawl pictures based on the seed samples to obtain crawler samples, including generating pictures for crawling, removing duplicate images, and storing images. This part is the specific implementation of the above step S302. Through crawling, sample augmentation is achieved, and at least one target image and the category of at least one target image are obtained. Model training and establishing the model API specifically include image deblurring, image enhancement, model training, and confidence recommendation for the pre-trained model based on the seed samples and crawler samples. This part is the specific implementation of the above step S303. Combining image processing to complete model training further realizes the optimization of the model.
[0125] Among them, the specific implementation process of the above steps can refer to the description of steps S301 - S303 above.
[0126] Such as Figure 8As shown in the figure, an embodiment of the present application further provides an image classification method, which is applied to a computing device, and the computing device may specifically be a management device in the cloud or an edge device. The method includes the following steps S801-S802.
[0127] S801. The computing device obtains a second model. The second model is a model obtained by training a first model based on original training samples and target training samples. The first model is an image classification model trained based on images in a public data source. The original training samples include at least one original image from an edge device managed by a cloud management device and the category of the at least one original image. The target training samples include at least one target image crawled based on the category of the at least one original image and the category of the at least one target image. When crawling based on the category of one original image, the category of the obtained target image is the same as the category of the original image.
[0128] In a possible implementation, when the computing device is a management device in the cloud, the at least one original image is obtained in the following manner: obtaining at least one historical image collected by the edge device; sorting the at least one historical image in reverse order according to the collection time to obtain a historical image sequence; and selecting at least one image from the first preset number of historical images in the historical image sequence as the at least one original image.
[0129] In a possible implementation, among the at least one original image, the number of original images of the same category is greater than or equal to a first threshold.
[0130] In a possible implementation, the clarity of the at least one original image is less than or equal to a second threshold; and / or, the accuracy of the at least one original image is less than or equal to a third threshold. The accuracy of the at least one original image is the accuracy of the classification result obtained by classifying the at least one original image using the first model compared to the category of the at least one original image. The category of the at least one original image is the category after manual annotation.
[0131] S802. The computing device classifies the images collected by the edge device based on the second model.
[0132] Specifically, the computing device inputs the image into the second model to obtain the output image classification result.
[0133] It is understandable that the image classification by the second model obtained by training based on the original training samples and the target training samples is helpful to realize image classification by using a model that fits the actual scene of edge device image collection, thereby improving the accuracy of image classification. Compared with the current pre-trained model based on public data sources, the above image classification method can improve the accuracy of online actual image classification by more than 10%. Figure 3 The training method of the image classification model can form a closed loop for the training of the image classification model. Humans only need to regularly annotate some images, and the management device can update the image classification model based on the annotated information. The above method can make full use of online data to form a positive cycle iteration.
[0134] It should be noted that when the above-mentioned image classification method is applied to a management device in the cloud, the second model can be generated or adjusted in combination with the above-mentioned training method for the image classification model.
[0135] It is understandable that the image classification method can also be applied to edge devices, and the image classification model is obtained through the management device in the cloud to classify the collected images. Furthermore, the edge device can also feed back the classification results to the management device in the cloud, and the management device counts the classification results and adaptively adjusts the model according to the classification results. For example, when the number of categories of images in the classification results changes, the management device can retrain the model based on the newly added or reduced categories, which can be specifically achieved through the above steps S301-S303, thereby realizing the update of the image classification model, which helps to flexibly adjust the image classification model according to the actual scene of the collected image and improve the accuracy of image classification.
[0136] Figure 9 is a structural block diagram of a training device for an image classification model shown in an embodiment of the present disclosure. Figure 9 The device comprises an acquisition unit 901, an expansion unit 902 and a training unit 903. Wherein:
[0137] The acquisition unit 901 is used to acquire an original training sample; wherein the original training sample includes: at least one original image from an edge device, and at least one category of the original image;
[0138] The expansion unit 902 is used to perform crawling based on the category of at least one original image to obtain a target training sample; wherein the target training sample includes at least one target image and at least one category of the target image; the category of the target image obtained by crawling based on the category of an original image is the same as the category of the original image;
[0139] A training unit 903 is configured to train a first model based on original training samples and target training samples to obtain a second model. The first model is an image classification model trained based on images in a public data source, and the second model is used to classify images collected by an edge device.
[0140] For the technical solution provided by the present disclosure, original training samples are obtained, and target training samples are obtained based on the original training samples in combination with web crawler technology. Using the original training samples and the target training samples as sample data to implement model training helps to generate an image classification model that fits the actual scenario of images collected at the edge. When image classification is performed based on this image classification model, the accuracy of the classification result is improved. Further, according to the training method of this image classification model, iterative optimization of this image classification model can be achieved based on the actually collected image data, forming a positive cyclic iteration, which helps to update the image classification model following the changes in the actual scenario, thereby improving the accuracy of image classification.
[0141] In some embodiments, the obtaining unit 901 specifically includes: obtaining at least one historical image collected by an edge device; sorting the at least one historical image in reverse order according to the collection time to obtain a historical image sequence; and selecting at least one image from the first preset number of historical images in the historical image sequence as at least one original image.
[0142] In some embodiments, for the at least one original image, the number of original images of the same category is greater than or equal to a first threshold.
[0143] In some embodiments, the clarity of the at least one original image is less than or equal to a second threshold; and / or the accuracy of the at least one original image is less than or equal to a third threshold. The accuracy of the at least one original image is the accuracy of the classification result obtained by using the first model to perform image classification on the at least one original image compared to the category of the at least one original image, and the category of the at least one original image is the category after manual annotation.
[0144] In some embodiments, the training unit 903 is further configured to delete duplicate images in the at least one target image to obtain the updated at least one target image; train the first model based on the original training samples and the updated target training samples to obtain a second model; and the target images in the updated target training samples are the updated at least one target image.
[0145] In some embodiments, the training unit 903 is further configured to perform image processing on at least one original image and at least one target image to obtain at least one processed original image and at least one processed target image; the image processing includes at least one of image deblurring and image enhancement, and the image enhancement includes any one or more of random cropping, rotation, scaling, and deformation; based on the processed original training samples and target training samples, train the first model to obtain a second model; the processed original training samples and target training samples include at least one processed original image and at least one processed target image.
[0146] Figure 10 is a structural block diagram of an image classification device shown in an embodiment of the present disclosure. Refer to Figure 10 , the device includes an acquisition unit 1001 and a classification unit 1002. Among them:
[0147] The acquisition unit 1001 is configured to acquire a second model; wherein, the second model is a model obtained by training the first model based on original training samples and target training samples; wherein, the first model is an image classification model trained based on images in a public data source; the original training samples include: at least one original image from an edge device managed by a computing device, and the category of at least one original image; the target training samples include: at least one target image crawled based on the category of at least one original image, and the category of at least one target image; wherein, based on the category of one original image, the category of the target image crawled is the same as the category of the original image;
[0148] The classification unit 1002 is configured to classify the images collected by the edge device based on the second model.
[0149] In some embodiments, the acquisition unit 1001 specifically includes:
[0150] Acquire at least one historical image collected by the edge device;
[0151] Sort at least one historical image in reverse order according to the acquisition time to obtain a historical image sequence;
[0152] Select at least one image from the first preset number of historical images in the historical image sequence as at least one original image.
[0153] In some embodiments, among at least one original image, the number of original images of the same category is greater than or equal to a first threshold.
[0154] In some embodiments, the clarity of at least one original image is less than or equal to a second threshold; and / or, the accuracy of at least one original image is less than or equal to a third threshold; wherein, the accuracy of at least one original image is the accuracy of the classification result obtained by performing image classification on at least one original image using a first model, compared to the category of at least one original image; the category of at least one original image is the category after manual annotation.
[0155] The technical solution provided by the present disclosure is to obtain original training samples, and based on the original training samples and combined with web crawler technology, obtain target training samples, and use the original training samples and target training samples as sample data to implement the training of the model, which helps to generate an image classification model that fits the actual scenario of the images collected at the edge. When performing image classification based on this image classification model, the accuracy of the classification result is improved. Further, according to the training method of this image classification model, the image classification model can be iteratively optimized based on the actually collected image data to form a positive loop iteration, which helps to update the image classification model following the changes in the actual scenario, thereby improving the accuracy of image classification.
[0156] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the training method or the image classification method of the image classification model provided by the present disclosure.
[0157] According to an embodiment of the present disclosure, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause an electronic device to execute the training method or the image classification method of the image classification model provided by the present disclosure.
[0158] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, including a computer program, which when executed by a processor, implements the training method or the image classification method of the image classification model provided by the present disclosure.
[0159] In some embodiments, the electronic device may be the management device shown in the above Figure 1 or when the electronic device is used to execute the image classification method provided by the present disclosure, it may also be the edge device shown in the above Figure 1 Figure 11 FIG. shows a schematic block diagram of an exemplary electronic device 1100 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0160] As Figure 11 shown, the electronic device 1100 includes a computing unit 1101 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0161] A plurality of components in the electronic device 1100 are connected to the I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, an optical disk, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows the electronic device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0162] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as the training method or the image classification method of an image classification model. For example, in some embodiments, the training method or the image classification method of an image classification model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the training method or the image classification method of the image classification model described above can be executed. Alternatively, in other embodiments, the computing unit 1101 can be configured to execute the training method or the image classification method of the image classification model by any other suitable means (e.g., by means of firmware).
[0163] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0164] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0165] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0166] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0167] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0168] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0169] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0170] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A training method for an image classification model, the method comprising: Obtaining at least one historical image collected by an edge device, and sorting the at least one historical image in reverse order of the collection time to obtain a historical image sequence; Selecting the at least one image from the first preset number of historical images in the historical image sequence as the at least one original image; wherein, the original training sample includes: at least one original image from the edge device managed by the management device, and the category of the at least one original image; among the at least one original image, the number of original images of the same category is greater than or equal to the first threshold; Wherein, the clarity of the at least one original image is less than or equal to the second threshold; and / or, the accuracy of the at least one original image is less than or equal to the third threshold; wherein, the accuracy of the at least one original image is the accuracy of the classification result obtained by classifying the at least one original image using the first model compared to the category of the at least one original image; the category of the at least one original image is the category after manual annotation; Performing web crawling according to the category of the at least one original image to obtain a target training sample; wherein, the target training sample includes at least one target image and the category of the at least one target image; the category of the target image obtained by performing web crawling based on the category of one original image is the same as the category of this original image; Training the first model based on the original training sample and the target training sample to obtain a second model; wherein, the first model is an image classification model trained based on images in a public data source; the second model is used to classify images collected by the edge device.
2. The method according to claim 1, the method further comprising: Deleting duplicate images in the at least one target image to obtain the updated at least one target image; The training the first model based on the original training sample and the target training sample to obtain a second model includes: Training the first model based on the original training sample and the updated target training sample to obtain the second model; The target images in the updated target training sample are the updated at least one target image.
3. The method according to claim 1 or 2, the method further comprising: Performing image processing on the at least one original image and the at least one target image to obtain the processed at least one original image and at least one target image; the image processing includes at least one of image deblurring and image enhancement, and the image enhancement includes any one or more of random cropping, rotation, scaling, and distortion; The training the first model based on the original training sample and the target training sample to obtain a second model includes: Training the first model based on the processed original training sample and the target training sample to obtain the second model; The processed original training sample and the target training sample include the processed at least one original image and at least one target image.
4. An image classification method, the method comprising: Obtaining a second model; wherein, the second model is a model obtained by training a first model based on original training samples and target training samples; wherein, the first model is an image classification model trained based on images in a public data source; the original training samples include: at least one original image from an edge device managed by a management device, and the category of the at least one original image; the target training samples include: at least one target image obtained by crawling based on the category of the at least one original image, and the category of the at least one target image; wherein, the category of the target image obtained by crawling based on the category of one original image is the same as the category of this original image; Classifying the images collected by the edge device based on the second model; Wherein, the at least one original image is obtained based on the following method: Obtaining at least one historical image collected by the edge device; Sorting the at least one historical image in reverse order according to the collection time to obtain a historical image sequence; Selecting the at least one image from the first preset number of historical images in the historical image sequence as the at least one original image, and the number of original images of the same category in the at least one original image is greater than or equal to a first threshold; Wherein, the clarity of the at least one original image is less than or equal to a second threshold; and / or, the accuracy of the at least one original image is less than or equal to a third threshold; wherein, the accuracy of the at least one original image is the accuracy of the classification result obtained by classifying the at least one original image using the first model compared to the category of the at least one original image; the category of the at least one original image is the category after manual annotation.
5. A training device for an image classification model, comprising: An obtaining unit, configured to obtain at least one historical image collected by an edge device, and sort the at least one historical image in reverse order according to the collection time to obtain a historical image sequence; The obtaining unit is configured to select the at least one image from the first preset number of historical images in the historical image sequence as the at least one original image; wherein, the original training samples include: at least one original image from an edge device, and the category of the at least one original image; the number of original images of the same category in the at least one original image is greater than or equal to a first threshold; Wherein, the clarity of the at least one original image is less than or equal to a second threshold; and / or, the accuracy of the at least one original image is less than or equal to a third threshold; wherein, the accuracy of the at least one original image is the accuracy of the classification result obtained by classifying the at least one original image using the first model compared to the category of the at least one original image; the category of the at least one original image is the category after manual annotation; An expansion unit for crawling according to the category of the at least one original image to obtain target training samples; wherein the target training samples include at least one target image and the category of the at least one target image; the category of the target image obtained by crawling based on the category of one original image is the same as the category of this original image; A training unit for training a first model based on the original training samples and the target training samples to obtain a second model; wherein the first model is an image classification model trained based on images in a public data source; the second model is used for classifying images collected by the edge device.
6. An image classification device, comprising: An acquisition unit for acquiring a second model; wherein the second model is a model obtained by training a first model based on original training samples and target training samples; wherein the first model is an image classification model trained based on images in a public data source; the original training samples include: at least one original image from an edge device and the category of the at least one original image; the target training samples include: at least one target image obtained by crawling based on the category of the at least one original image and the category of the at least one target image; wherein the category of the target image obtained by crawling based on the category of one original image is the same as the category of this original image; A classification unit for classifying images collected by the edge device based on the second model; Wherein the at least one original image is obtained in the following manner: Obtaining at least one historical image collected by the edge device; Sorting the at least one historical image in reverse order according to the collection time to obtain a historical image sequence; Selecting the at least one image from the first preset number of historical images in the historical image sequence as the at least one original image, and the number of original images of the same category in the at least one original image is greater than or equal to a first threshold; Wherein the clarity of the at least one original image is less than or equal to a second threshold; and / or the accuracy of the at least one original image is less than or equal to a third threshold; wherein the accuracy of the at least one original image is the accuracy of the classification result obtained by classifying the at least one original image using the first model compared to the category of the at least one original image; the category of the at least one original image is the category after manual annotation.
7. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-3 or 4.
8. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-3 or 4.
9. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 - 3 or 4.
Citation Information
Patent Citations
Image classification model training method and device and electronic equipment
CN114972877A
Image classification method and training method of image classification model
CN115564992A