Method, Medium and System for Classifying Images on a Terminal Device
By combining the image classification model of terminal devices and cloud servers, the problem of limited classification capabilities of terminal devices' image classification model is solved, and the increase of image classification categories and the improvement of accuracy are achieved, avoiding the leakage of image privacy.
Patent Information
- Application Number
- CN202010652546.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-07-08
AI Technical Summary
The limited computing power, memory and storage space of terminal devices leads to limited classification capabilities of image classification models, fewer categories supported and less accuracy.
By combining the image classification model in the terminal device and the image classification model on the cloud server side temporarily acquired, detailed classification or classification calibration of images can be achieved, classification categories and accuracy can be improved.
Without revealing the image privacy of terminal devices, improve the classification capabilities of image classification, increase the number of supported categories and improve classification accuracy.
Smart Images

Figure CN113936206B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and particularly to a method, medium, and system for classifying images on a terminal device. Background Art
[0002] A large number of image files, such as photos taken by users daily, are often stored in terminal devices (such as mobile phones, tablets, etc.). For the large number of image files in the terminal device, the terminal device generally uses image classification technology to classify the images to support the user to find a certain category of images through the image search function, such as images of landscapes, flowers, animals, cars, etc.
[0003] Currently, image classification technology usually uses a neural network model (specifically, an image classification model) to identify images and uses the recognition result as the label of the image. For example, the recognition result of an image can be one or more labels, that is, an image can have one or more labels. For example, an image containing a rose has two labels, namely the label "flower" and the label "rose". Among them, since the images stored in the terminal device are the user's privacy data, the images in the terminal device cannot be uploaded to the cloud server for classification processing by the cloud server, but are generally classified by the image classification model stored in the terminal device.
[0004] However, due to the limited computing power, memory, and storage space of the terminal device, the storage space for storing the image classification model on the terminal device cannot be very large (generally less than 100 megabytes (MB)), so the computing power (i.e., classification ability) of this image classification model is relatively limited, generally only supporting a few hundred label categories, and the accuracy of image classification is also relatively low. Summary of the Invention
[0005] Embodiments of this application provide a method, medium, and system for classifying images on a terminal device, which can increase the number of classification categories and improve the classification accuracy of image classification on the terminal device by combining the classification ability of the image classification model in the terminal device and the classification ability of the image classification model temporarily obtained on the cloud server side, and avoid the privacy leakage of the images in the terminal device.
[0006] In a first aspect, an embodiment of the present application provides a method for classifying images on a terminal device. The method includes: the terminal device classifies multiple images on the terminal device into a first type through an image classification model in the terminal device; determines that the condition for reclassifying the multiple images is met, and then obtains a first image sub-classification model corresponding to the first type from the server; further, classifies the multiple images into multiple first sub-types of the first type through the first image sub-classification model, and finally deletes the first image sub-classification model. For example, the terminal device can classify multiple images into large categories such as cars, flowers, and animals through the image classification model in the terminal device. The terminal device can obtain an image sub-classification model corresponding to flowers from the server and use the image sub-classification model to classify the images classified into the large category of flowers into small categories such as roses, chrysanthemums, and sunflowers. Also, the terminal device can further classify the images of the chrysanthemum category into small categories such as white chrysanthemums, red chrysanthemums, and white chrysanthemums through the first image sub-classification model. In this way, even if the classification ability of the image classification model fixedly stored on the terminal device is low, without uploading the images in the terminal device to the server, the terminal device can still perform detailed classification or classification calibration on the images on the terminal device through the first image sub-classification model temporarily obtained and stored on the server side. And the classification ability of the image sub-classification model on the server side is usually strong, such as supporting more classification categories and having higher classification accuracy. In this way, while avoiding the privacy leakage of the images on the terminal device, the terminal device can combine the classification ability of the image classification model in the terminal device and the classification ability of the image sub-classification model on the server side (such as the first image sub-classification model) to improve the classification ability of the terminal device for image classification on the terminal device. Additionally, since the terminal device can timely delete the image classification model from the storage space of the terminal device after using the image sub-classification model on the server side, it can timely release the storage space of the terminal device, which is beneficial to the normal operation of the image classification method of the terminal device and other conventional services such as call services or Internet services.
[0007] It should be noted that the image classification model in the terminal device can be the resident model described below, and the image sub-classification model in the server (such as the first image sub-classification model) can be the non-resident model described below, and the server can be the cloud server described below. Additionally, for the convenience of description, different descriptions are used for an image classification at different positions in this article, such as "type", "category", "large category", and "small category", but their essence is the same.
[0008] In a possible implementation of the first aspect described above, the conditions for reclassifying the multiple images (such as items "1)" and "2)" in trigger condition D2 below) include at least one of the following: the number of the multiple images is greater than or equal to a first quantity threshold (such as 50); the ratio of the number of the multiple images to all the images to be classified on the terminal device is greater than or equal to a second quantity threshold (such as 10%), where all the images to be classified are all or part of the images stored on the terminal device. It can be understood that at this time, there are more images in the multiple images, and the multiple images often need to be further refined in classification, such as reclassification or classification calibration.
[0009] Specifically, after the terminal device classifies the images to be classified into each major category (including the first type) through the image classification model in the terminal device, it respectively determines whether the images in each major category meet the above "conditions for reclassification", that is, whether the number of images in each major category is greater than or equal to the first quantity threshold, and / or whether the ratio of the number of images in each major category to all the images to be classified is greater than or equal to the second quantity threshold, so as to determine that the multiple images classified as the first type meet the conditions for reclassification. For example, if the number of images in the flower major category is 150 (greater than 50), or the ratio of the number of 150 images in the flower major category to the number of 500 all images to be classified is 30% (greater than 10%), meeting the above "conditions for reclassification", then the images in the flower major category meet the above "conditions for reclassification".
[0010] In this way, since the terminal device requests the image subdivision model from the server for the images that need to be refined in classification or classification calibration, such as the first image subdivision model requested for the multiple images of the first type, rather than generally requesting the image sub-classification model for all the images to be classified in the terminal device. Thus, it is possible to avoid excessive computational load during the image classification process of the terminal device, and the image subdivision model temporarily obtained and stored by the terminal device occupying too much storage space of the terminal device.
[0011] In a possible implementation of the first aspect described above, the terminal device classifies multiple images on the terminal device into the first type through the image classification model in the terminal device, including: the terminal device classifies the multiple images into the first type through the image classification model and classifies the multiple images into multiple second subtypes of the first type. It can be understood that the image classification model in the terminal device can not only classify the images to be classified into major categories such as the first type, but also can, to a certain extent, subdivide the images of a major category into images of different sub-categories. For example, through the image classification model in the terminal device, the terminal device can subdivide the images in the flower major category into sub-categories such as roses and chrysanthemums.
[0012] In a possible implementation of the above first aspect, the condition for re - classifying the multiple images (such as item "3)" in trigger condition D2 below) includes: the ratio of the number of images of the target subtype in the multiple images to the number of the multiple images is greater than or equal to a third quantity threshold (such as 10%), where the target subtype is the subtype of the images to be re - classified among the multiple images. It can be understood that at this time, the multiple images often need to be further refined in classification, such as re - classification or classification calibration. For example, the target subtype is the images in the terminal device that are not classified in detail in the large categories classified by the image classification model, such as other small categories in the large category of flowers, that is, the small categories for which it is not detailed which kind of flower it is. Specifically, after all the images to be classified are classified into each large category (such as the first type) by the image classification model in the terminal device, it can be determined whether the images in each large category meet the above "condition for re - classification", that is, whether the ratio of the number of images of the target subtype in the images of each large category to the number of all images in that large category is greater than or equal to the third quantity threshold, so as to determine that the multiple images classified as the first type meet the condition for re - classification. For example, the ratio of the number of 60 images of other small categories in the large category of flowers to the number of 150 images in the large category of flowers is 40% which is greater than 10%, that is, the images in the large category of flowers meet the above "condition for re - classification". In this way, the user can screen out the images that need to be refined in classification or classified and calibrated, and request the image segmentation model from the server for these images.
[0013] In a possible implementation of the above first aspect, the condition for re - classifying the multiple images (such as trigger condition D1 below) includes at least one of the following: the terminal device is in a charging state; the battery power of the terminal device is greater than or equal to a power threshold; the terminal device is connected to a wireless network; the terminal device receives an instruction from the user allowing re - classification. It can be understood that the terminal device meeting this "re - classification condition" indicates that the operating state of the terminal device is good, and there is relatively sufficient storage space and network resources on the terminal device to request the temporary first - image segmentation model from the server, ensuring the normal execution of the image classification method on the terminal device. On the contrary, it indicates that the operating state of the terminal device is poor. To ensure the normal execution of the regular services of the terminal device, such as call services or Internet services, the terminal device will not request the temporary image classification model from the server.
[0014] In a possible implementation of the above first aspect, the above terminal device classifies multiple images on the terminal device into a first type through an image classification model on the terminal device, including: the terminal device labels the multiple images classified as the first type with a first main label for identifying the first type through the image classification model. For example, the first main label of the images in the flower category is "flower", which is used to indicate this large category of flowers. In this way, it is convenient for the subsequent terminal device to query or distinguish different categories (or types), such as distinguishing the images of the first type from the images of other types. For example, it is convenient for the user to subsequently search for the images of the flower category by inputting the main label "flower". Among them, a main label can be a first-level label in the following text.
[0015] In a possible implementation of the above first aspect, the above terminal device obtains a first image sub-classification model corresponding to the first type from the server, including: the terminal device sends a re-classification request (i.e., the model application request in the following text) to the server, and this re-classification request is used to request to obtain the first image sub-classification model from the server; the terminal device receives a re-classification response from the server (i.e., the message carrying the information of the non-resident model sent by the server in the following text), and this re-classification response includes the first image sub-classification model. In this way, triggered by the re-classification request of the terminal device, the server returns the first image sub-classification model to the terminal device corresponding to the request of the terminal device, and the first image sub-classification model can be obtained from the server.
[0016] In a possible implementation of the above first aspect, the above re-classification request includes a first main label, and the first main label is used to identify the first type and corresponds to the first image sub-classification model. In this way, when the server receives the re-classification request, it can know that the terminal device requests the first image sub-classification model corresponding to the first main label.
[0017] In a possible implementation of the above first aspect, the above re-classification request further includes a second main label and a third main label, and the re-classification request is further used to request to obtain a second image sub-classification model corresponding to the second main label and a third image sub-classification model corresponding to the third main label from the server; among them, the second main label is used to identify the images classified as the second type by the image classification model, and the third main label is used to identify the images classified as the third type by the image classification model. It can be understood that the terminal device can first determine that the images of the second type and the third type both meet the above "re-classification conditions", and for this, reference can be made to the above description of the multiple images of the first type meeting the re-classification conditions, which will not be elaborated here. In addition, the second image sub-classification model is used to classify all the images of the second type into multiple sub-types of the second type, and the third image sub-classification model is used to classify all the images of the third type into multiple sub-types of the third type. For example, the image sub-classification model corresponding to the car category is used to classify all the images of the car category into sub-categories (i.e., sub-types) such as sedans, taxis, and trucks.
[0018] In a possible implementation of the above first aspect, the reclassification response further includes the second image segmentation model, or the reclassification response further includes the second image segmentation model and the third image segmentation model. That is, the image segmentation models returned by the server to the terminal device can be all the image segmentation models requested by the terminal device, or some of the image segmentation models requested by the terminal device. In this way, it is possible to avoid an excessive number of image segmentation models and a large amount of data obtained by the terminal device, so as to avoid these image segmentation models occupying a large amount of storage capacity of the terminal device, which is conducive to the normal operation of the image classification method and other conventional services (such as call services and Internet services) on the terminal device.
[0019] In a possible implementation of the above first aspect, the reclassification request further includes a target capacity (such as 200M), where the target capacity represents the storage capacity of the terminal device that can be used to store the image segmentation models obtained from the server; the sum of the data amounts of all the image segmentation models in the reclassification response is less than or equal to the storage capacity represented by the target capacity. For example, the reclassification response only includes the first image segmentation model and the second image segmentation model, and the sum of the data amounts of these two image segmentation models is less than or equal to the storage capacity represented by the target capacity. Among them, the target capacity enables the server to clearly know how much the data amount of the image segmentation model sent to the terminal device is, which can avoid the image segmentation models obtained by the terminal device and is conducive to the normal operation of the image classification method on the terminal device.
[0020] In a possible implementation of the above first aspect, the terminal device classifies multiple images into multiple first subtypes of the first type through the first image segmentation model, including: the terminal device reclassifies multiple images from multiple second subtypes into multiple first subtypes through the first image segmentation model. It can be understood that due to the limited classification ability of the image classification model in the terminal device, the number of subcategories of the images classified by this image classification model is limited and the accuracy is low. For example, in a scenario where the image classification model in the terminal device cannot classify images of subcategories such as sunflowers in the large category of flowers, the terminal device obtains an image segmentation model corresponding to the flowers (such as the first image segmentation model) from the server, which can finely classify subcategories such as sunflowers in the images of the large category of flowers. Or, in a scenario where the image classification model in the terminal device cannot classify images of subcategories in the category of chrysanthemums, the terminal device obtains an image segmentation model corresponding to the flowers from the server, which can finely classify subcategories such as yellow chrysanthemums, white chrysanthemums, and red chrysanthemums in the images of the category of chrysanthemums. In this way, it is possible to reclassify or calibrate the classification of the images classified by the terminal device, improving the classification ability of the image classification on the terminal device side, such as an increase in the supported categories and an improvement in accuracy.
[0021] In a possible implementation of the foregoing first aspect, the terminal device classifies the multiple images into multiple first subtypes of the first type through the first image segmentation model, including: the terminal device labels different first sub-labels for images belonging to different first subtypes through the image segmentation model. It can be understood that in this application, by adding a main label and a sub-label to an image, the terminal device can divide an image into different image sets or image subsets to classify the image in detail. In this way, it is convenient for the terminal device to further query or distinguish different categories (or types) subsequently, such as distinguishing images of small chrysanthemum categories from images of small rose categories. For example, it is convenient for the user to subsequently search for images of the small rose category by inputting the sub-label "rose". Among them, a sub-label can be a secondary label, a tertiary label, etc. in the following text.
[0022] In a second aspect, an embodiment of the present application provides a method for classifying images on a terminal device. The method includes: the server receives a reclassification request from the terminal device, and the reclassification request is used to request to obtain a first image fine classification model from the server, where the first image fine classification model is used to classify multiple images classified as the first type by the image classification model of the terminal device into multiple first subtypes of the first type; the server searches for the first image segmentation model corresponding to the first type according to the reclassification request; the server sends a reclassification response to the terminal device, and the reclassification response includes the first image segmentation model. In this way, even if the classification ability of the image classification model fixedly stored on the terminal device is low, without uploading the images in the terminal device to the server, the terminal device can still perform detailed classification or classification calibration on the images on the terminal device through the first image segmentation model temporarily obtained and stored on the server side. The classification ability of the image segmentation model on the server side is usually strong, such as supporting more classification categories and higher classification accuracy. In this way, while avoiding the privacy leakage of the images on the terminal device, the terminal device can combine the classification ability of the image classification model in the terminal device and the classification ability of the image segmentation model (such as the first image segmentation model) on the server side to improve the classification ability of the terminal device for image classification on the terminal device. In addition, since the terminal device can timely delete the image classification model from the storage space of the terminal device after using the image segmentation model on the server side, the storage space of the terminal device can be released in time, which is beneficial to the normal operation of the image classification method of the terminal device and other conventional services such as call services or Internet services.
[0023] In a possible implementation of the second aspect described above, the reclassification request includes a first main tag, and the first main tag is used to identify the first type and corresponds to the first image segmentation model. For the specific description thereof, reference may be made to the relevant description of the first aspect above, which will not be elaborated here. Additionally, different image segmentation models in the server may correspond to different main tags, such as different image segmentation models having different tags. And the tags of the image segmentation models may correspond to the same classification tags to be classified, or to the classification tags to be classified that are pre-associated. For example, the classification tag to be classified "flower" corresponds to the image segmentation model with the tag "flower", or the classification tag to be classified "flower" corresponds to the image segmentation model with the tag "flower" and the image segmentation model with the tag "tree". In this way, the server can search for and send the first image segmentation model based on the first main tag.
[0024] In a possible implementation of the second aspect described above, the reclassification request further includes a target capacity, where the target capacity represents the storage capacity (such as 200M) that the terminal device can use to store the image segmentation model obtained from the server. For the specific description thereof, reference may be made to the relevant description of the first aspect above, which will not be elaborated here.
[0025] In a possible implementation of the second aspect described above, the reclassification request further includes a second main tag and a third main tag; and the method further includes: the server determines that the image segmentation models in the reclassification response sent to the terminal device only include the first image segmentation model and the second image segmentation model corresponding to the second main tag according to the target capacity in the reclassification request, where the sum of the data volumes of the first image segmentation model and the second image segmentation model is less than or equal to the storage capacity represented by the target capacity. For the specific description thereof, reference may be made to the relevant description of the first aspect above, which will not be elaborated here. Additionally, the image segmentation models returned by the server to the terminal device may be partial image segmentation models requested by the terminal device. Among them, the server may randomly select several image segmentation models whose sum of data volumes is less than or equal to the target capacity from the image segmentation models requested by the terminal device based on the target capacity. Or, the server may select the non-resident models whose sum of data volumes is less than or equal to the target capacity from the image segmentation models requested by the terminal device according to the model priority based on the target capacity. Similarly, after the terminal device uses the second image segmentation model, the third image segmentation model, etc., these image segmentation models can also be deleted.
[0026] In a possible implementation of the second aspect described above, the reclassification request further includes a second main tag and a third main tag; and the method further includes: The server determines, according to the target capacity in the reclassification request, that the image segmentation models in the reclassification response sent to the terminal device include a first image segmentation model, a second image segmentation model corresponding to the second main tag, and a third image segmentation model corresponding to the third main tag, where the sum of the data volumes of the first image segmentation model, the second image segmentation model, and the third image segmentation model is less than or equal to the storage capacity represented by the target capacity. For the specific description of this, reference can be made to the relevant description of the first aspect above, and details will not be elaborated here. In addition, the image segmentation models returned by the server to the terminal device can be all the image segmentation models requested by the terminal device, or some of the image segmentation models requested by the terminal device. Similarly, referring to the previous implementation of the second aspect, the server can screen out the non-resident models whose sum of data volumes is less than or equal to the target capacity from the image segmentation models requested by the terminal device based on the target capacity.
[0027] In a third aspect, according to an embodiment of the present application, a computer-readable medium is disclosed, on which instructions are stored, and when the instructions are executed on a machine, the machine executes the method for classifying images on a terminal device described in any one of the first aspect or the second aspect above.
[0028] In a fourth aspect, according to an embodiment of the present application, a system is disclosed, which is applied to a terminal device or a server, and the system includes:
[0029] A memory for storing instructions executed by one or more processors of the system, and
[0030] A processor, which is one of the processors of the system, for executing the method for classifying images on a terminal device described in any one of the first aspect or the second aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 A schematic diagram showing the relationship between an image classification model and an image set disclosed in an embodiment of the present application;
[0032] Figure 2 A schematic diagram of a scenario to which an image classification method disclosed in an embodiment of the present application is applied;
[0033] Figure 3 A schematic flowchart of an image classification method disclosed in an embodiment of the present application;
[0034] Figure 4 A schematic diagram of a UI interface of an image classification result disclosed in an embodiment of the present application;
[0035] Figure 5Schematic diagram II of a UI interface for an image classification result disclosed in an embodiment of the present application;
[0036] Figure 6 Schematic diagram III of a UI interface for an image classification result disclosed in an embodiment of the present application;
[0037] Figure 7 Schematic diagram of the structure of a terminal device disclosed in an embodiment of the present application;
[0038] Figure 8 Block diagram of a system disclosed in an embodiment of the present application. Detailed implementation manners
[0039] The present application will be further described below in conjunction with specific embodiments and the accompanying drawings. It can be understood that the specific embodiments described herein are only for explaining the present application, rather than limiting the present application. In addition, for the sake of description, only parts related to the present application rather than all structures or processes are shown in the accompanying drawings. It should be noted that in this specification, similar reference numerals and letters denote similar items in the following drawings.
[0040] Exemplary embodiments of the present application include, but are not limited to, image classification methods and their terminal devices and media, etc.
[0041] In an embodiment of the present application, the terminal device may first use an image classification model stored therein to perform a rough classification on an image to obtain a plurality of image sets. For example, the rough classification obtains image sets of large categories such as cars, flowers, and animals. Subsequently, the terminal device may request a temporary image classification model for detailed classification of the image from the cloud server according to the rough classification result. Furthermore, these image classification models are used to perform detailed classification on the images in the existing large-category image sets, or to calibrate the existing classifications of these images to obtain a plurality of sub-image sets of refined categories. For example, the images in the image set roughly classified as flowers are refined and classified into sub-image sets such as roses, chrysanthemums, sunflowers, and other flowers.
[0042] In this way, the present application performs image classification by combining the image classification capabilities of the image classification model stored fixedly in the terminal device and the image classification capabilities of the image classification model stored temporarily on the cloud server side, thereby improving the image classification capabilities in the terminal device, such as increasing the classification categories supported by the image classification, and improving the accuracy of the image classification. In this way, the problem that the single image classification model on the terminal device with limited computing power, memory and storage space does not support enough types and the accuracy is not accurate enough is solved. In addition, the terminal device obtains a small amount of image classification models from the cloud server side on demand each time, which will not occupy too much storage space of the terminal device. Moreover, after the terminal device uses an image classification model on the cloud server side, the image classification model can be deleted from the storage space of the terminal device to release the storage space of the terminal device in time, thereby facilitating the normal operation of the conventional business of the terminal device. Furthermore, since the terminal device does not need to send the image to other devices such as the cloud server, the privacy leakage of the image in the terminal device is avoided.
[0043] According to some embodiments of the present application, we refer to the image classification model (i.e., the neural network model used for image classification) that is fixedly stored in the terminal device as a resident model; correspondingly, the image classification model stored in the cloud server is referred to as a non-resident model. Obviously, whether the image classification model is resident in the present application refers to whether the image classification model is stored in the terminal device for a long time (or fixedly). It can be understood that the "resident model" and "non-resident model" mentioned in the embodiments of the present application are only an exemplary description of the present application, and the present application is not limited to this. In other embodiments, the image classification model fixedly stored in the terminal device may also be called by other names, such as "end-side model".
[0044] For example, Figure 1 FIG. 1 is a schematic diagram showing the relationship between an image model and an image set in some embodiments of the present application. Figure 1 , the images that need to be classified in the terminal device are taken as the image set, and the images in the image set are input into the resident model to obtain the rough classification results, for example, image set 1, image set 2, image set 3...image set N (i.e., image set 1 to image set N) are obtained. For example, image set 1 to image set N are respectively image sets of a large category, such as image set 1 of the large category of "flowers", image set 2 of the large category of "cars", image set 3 of the large category of "animals", and set 4 of the large category of "others". Corresponding to each image set 1-N, there is a corresponding non-resident model F1-FN, which is used to finely classify the image set classified by the resident model, for example, classifying the image set 1 of the large category of "flowers" into "roses", "chrysanthemums", etc.
[0045] The data volume of the resident model fixedly stored in the terminal device is usually small, resulting in fewer categories supported by image classification using the resident model and a lower accuracy rate of image classification. For example, the currently most typical MobileNetV3_Large model, the size of this model (i.e., the data volume size) is 21.6 megabytes (MB), and the TOP1 classification accuracy in the 1000-class image classification task of the ImageNet dataset is 75.2%. However, the labels of 1000 categories are far from meeting the various image classifications on the terminal device, and the classification accuracy of 75.2% is not high enough, resulting in fewer categories supported by image classification in the terminal device and a lower accuracy rate.
[0046] The data volume of the non-resident model stored in the cloud server is usually large, resulting in more categories supported by image classification using the non-resident model and a higher accuracy rate of image classification. For example, compared with the above-mentioned MobileNetV3_Large model on the edge side, an image classification model in the cloud server, the model size is 1.92 gigabytes (GB), and the TOP1 classification accuracy in the 1000-class image classification task of the ImageNet dataset is 88.4%, which is much higher than the above 75.2%.
[0047] The following combines Figures 1 to 8 to further describe the embodiments of the present application in detail.
[0048] Figure 2 This is a scenario to which an image classification method disclosed in an embodiment of the present application is applied. As Figure 2 shown, this scenario 10 includes: a terminal device 100 and a cloud server 200.
[0049] Among them, the terminal device 100 can be various electronic devices capable of storing images and performing image classification, such as: mobile phones, computers, laptop computers, tablets, TVs, display devices, outdoor display screens, vehicle-mounted terminals, and other devices. The terminal device 100 performs image classification operations through the edge-side image classification method to facilitate the user to find different categories of images in the terminal device 100.
[0050] In some embodiments, the terminal device 100 may perform wireless communication with other electronic devices through various wireless means. For example, it may perform wireless communication with the cloud server 200. For example, the terminal device 100 may send wireless signals to the cloud server 200 through its own radio frequency circuit and antenna via a wireless communication link, or obtain data for processing the specific services of the terminal device 100 from the cloud server 200. For example, the terminal device 100 sends the classification result of classifying an image using a resident model to the cloud server 200, or the terminal device 100 requests a message of a non-resident model from the cloud server 200, etc. Of course, the terminal device 100 may also perform data communication with the cloud server 200 through other wireless communication means, such as radio frequency identification technology, Wi-Fi network, etc., and the embodiments of the present application do not limit this.
[0051] The cloud server 200 may be a hardware server or may be implanted in a virtualized environment. For example, according to some embodiments of the present application, the cloud server 200 may be a virtual machine executed on a hardware server including one or more other virtual machines. According to some embodiments of the present application, the cloud server 200 may interact with the terminal device 100 through a network, such as sending data to the terminal device 100 and / or receiving data from the terminal device 100, such as receiving a model application request for a non-resident model from the terminal device 100 and sending information of the corresponding non-resident model to the terminal device 100.
[0052] It can be understood that in the embodiments of the present application, the server that interacts with the terminal device 100 to implement the image classification method includes, but is not limited to, the above-mentioned cloud server 200, and may also be an independent server. Among them, the independent server has all the software and hardware resources of the entire server and can allocate and implement various network function services by itself. For example, it can independently interact with the terminal device 100 to implement the image classification method in the present application.
[0053] For the convenience of description, hereinafter, taking the terminal device 100 as a mobile phone as an example, the image classification method provided by the embodiments of the application will be introduced in detail.
[0054] Specifically, according to some embodiments of the present application, in Figure 2 the scenario shown, the image classification method of the present application includes:
[0055] Step 301: The mobile phone 100 classifies the to-be-classified images stored on the mobile phone 100 through a resident model, and obtains image sets 1 to image set N, where N is a positive integer. Among them, the to-be-classified images may be all the images stored on the mobile phone 100, or all the images in the album of the mobile phone 100, or the latest unclassified images stored on the mobile phone 100, or a part of the images in the album.
[0056] Specifically, in some embodiments, before the mobile phone 100 invokes the resident model, the resident model is always stored in the mobile phone 100, such as in the read-only memory (ROM) of the mobile phone 100 (i.e., the storage space of the mobile phone 100). Subsequently, when the mobile phone 100 needs to execute step 301 above, that is, classify all the images in the image set in the mobile phone 100, the resident model can be invoked into the random-access memory (RAM) (i.e., the memory of the mobile phone 100). Furthermore, the mobile phone 100 can classify the images based on the resident model invoked by the RAM. For example, as described above, it is divided into image set 1 to image set N.
[0057] In addition, in some embodiments, after the mobile phone 100 inputs the image to be classified into the resident model, the resident model can label the image after classification by the resident model to distinguish images of different major categories. The label labeled by the resident model can be called a first-level label or a main label. For example, the first-level label of an image containing a rose is "flower", and the first-level label of an image containing a bus is "car".
[0058] It can be understood that for images with the same first-level label or main label, after being classified by the resident model, they belong to the same image set, such as the above image set 1, image set 2,... image set N, etc. For example, the first-level label of an image A containing a rose is "flower", and the first-level label of an image B containing a chrysanthemum is also "flower". Therefore, image A and image B both belong to the "flower" image set. In addition, as described below, the images after being re-classified by the non-resident model can be labeled with second-level labels, third-level labels, and so on.
[0059] In other embodiments, the label labeled by the resident model for the image after classification by the resident model can not only distinguish images of different major categories, but also distinguish different sub-categories in some major categories to a certain extent. At this time, the label labeled by the resident model can include a first-level label, a second-level label, a third-level label, etc. For example, after an image containing a rose is classified by the resident model, it not only has a first-level label "flower", but also has a second-level label "rose". That is, the classification ability of the resident model can be adjusted according to the computing power and storage capacity of the mobile phone 100. For a mobile phone 100 with stronger computing power and larger storage space, compared with a mobile phone 100 with weaker computing power or smaller storage space, a resident model with more powerful classification ability can be set. It can be understood that although the resident model on the mobile phone 100 has the ability of fine classification, the ability of the resident model to classify images into small categories is limited. Therefore, after the image to be classified is input into the resident model, more images that are not classified in detail are obtained, such as images classified into the "other" category.
[0060] In some embodiments of the present application, the tags of each level of an image are stored in the image attribute information, and are arranged in descending order of level in the image attribute information to form a tag set (or sequence). For example, the tag set of an image containing a rose is {flower, chrysanthemum} composed of the first-level tag "flower" and the second-level tag "chrysanthemum". In some embodiments of the present application, the image set of a large category can be identified by the first-level tag or the tag set of the images therein. For example, the tag corresponding to the image set of the large category "flower" is the tag "flower", or the sequence {flower, chrysanthemum} is used for identification. It can be understood that the attribute information of the image generally includes the width, height, color concentration, coding method, etc. of the image.
[0061] Step 302: The mobile phone 100 determines whether the running state of the mobile phone 100 meets the trigger condition D1 for requesting the non-resident model.
[0062] If the trigger condition D1 is met, the following step 303 is executed; if the trigger condition D1 is not met, step 302 is continued. In addition, in other embodiments, if the mobile phone 100 does not meet the trigger condition D1, the mobile phone 100 can end the method flow of image classification.
[0063] In some embodiments, the trigger condition D1 includes at least one of the following:
[0064] 1) The mobile phone 100 is in a charging state (i.e., the mobile phone 100 is connected to a charger).
[0065] 2) The battery power of the mobile phone 100 (i.e., the battery power of the mobile phone 100) exceeds the power threshold (such as 70%). It can be understood that the specific threshold value here (such as the power threshold of 70%, or the quantity threshold of 50 below) is only exemplary, and the size of the specific threshold can be modified according to the actual situation, and is not limited in the embodiments of the present application.
[0066] 3) The mobile phone 100 is connected to a wireless network, such as a Wi-Fi network.
[0067] 4) The mobile phone 100 is connected to a mobile data network and the user allows to request the non-resident model under the mobile data network.
[0068] 5) The user triggers to select to allow the mobile phone 100 to interact with the cloud server 200 to obtain the non-resident model. For example, in some embodiments, a trigger control can be provided in the UI interface of the mobile phone 100, and the trigger control can be floating or superimposed on any UI interface of the "Gallery" APP, or the trigger control can be displayed on the search interface in the "Gallery" APP, and the user can click the trigger control to allow obtaining the non-resident model. As Figure 4As shown, the search interface 11 includes a "cloud classification" control 11a, which is the above-mentioned trigger control and is used to trigger the mobile phone 100 to request a non-resident model from the cloud server 200, such as requesting the non-resident model corresponding to the image set 1 of the "flower" category.
[0069] It can be understood that if the running state of the mobile phone 100 meets any one of 1) to 3) in the trigger condition D1, it indicates that the running state of the mobile phone 100 is good, and the mobile phone 100 usually has relatively sufficient storage space and network resources to request a non-resident model from the cloud server 200. On the contrary, it indicates that the running state of the mobile phone 100 is poor. Then, in order to ensure the normal execution of the regular services of the mobile phone 100, such as call services or Internet services, the mobile phone 100 will not request a non-resident model from the cloud server 200.
[0070] Step 303: The mobile phone 100 filters out the image set to be classified in detail from the image sets 1-N classified in step 301, and determines the classification labels to be classified for the image set to be classified in detail. For example, the mobile phone 100 can filter out the image set that meets the filtering condition D2 as the image set to be classified in detail.
[0071] It can be understood that in order to avoid excessive computational complexity during the image classification process of the mobile phone 100 and an excessive number of non-resident models temporarily stored by the mobile phone 100 later, the mobile phone 100 needs to filter the image sets classified in step 301. Specifically, the above filtering condition D2 can be at least one of the following:
[0072] 1) The number of images in the image set is greater than or equal to the first quantity threshold (such as 50).
[0073] 2) The ratio of the number of images in the image set to the number of images in the entire image set is greater than or equal to the second quantity threshold (such as 10%). It can be understood that the entire image set here can be all the images stored on the mobile phone 100 or all the images in the photo album of the mobile phone 100.
[0074] 3) The ratio of the number of images not classified in detail in the image set to the number of all images in the image set is greater than or equal to the third quantity threshold (such as 10%).
[0075] It can be understood that the number of images in the image sets that meet conditions 1) and 2) is usually large, and the images in these image sets often need to be further refined or calibrated in classification. For example, the number of images in the entire image set is 500, the number of images in image set 1 is 100, the number of images in image set 2 is 150, and the number of images in image set 3 is 130. Obviously, the number of images in image set 1, image set 2, and image set 3 is greater than the first quantity threshold (i.e., 50), and they are image sets that meet the screening condition D2.
[0076] In addition, it can be understood that the number of images that are not classified in detail in the image sets that meet condition 3) is large, and the images in these image sets often need to be further refined in classification. For example, when the images in the large category of flowers are subdivided into small categories such as roses and chrysanthemums, the images in the large category of flowers need to be further subdivided into new small categories such as sunflowers.
[0077] Specifically, in some embodiments, in this step, the mobile phone 100 determines whether the image sets 1-N meet the screening condition D2. If the screening condition D2 is met, the mobile phone 100 obtains the label of this image set as one of the to-be-classified labels; otherwise, the mobile phone 100 continues to traverse whether other image sets meet the screening condition D2. After the mobile phone 100 filters out all the image sets that meet the screening condition D2, it can request the non-resident models corresponding to the filtered image sets from the cloud server 200. For example, the to-be-classified labels of the image sets that meet the detailed classification filtered by the mobile phone 100 = {C1, C2, C3}, and C1, C2, and C3 are the first-level labels of image set 1, image set 2, and image set 3 respectively. Then, the mobile phone 100 requests the non-resident models corresponding to image set 1, image set 2, and image set 3 from the cloud server 200 respectively.
[0078] Step 304: The mobile phone 100 determines the capacity value of the available model storage space.
[0079] Specifically, the mobile phone 100 can access its ROM and determine the capacity value of the available model storage space in the ROM, that is, the capacity value of the storage space for storing non-resident models subsequently. For example, the unit of the capacity value of the available model storage space (denoted as S1) can be MB. Of course, the specific value of S1 can be determined according to the actual needs of the user, and no limitation is made here. However, due to the limited storage space, memory, and computing power of the mobile phone 100, the value of S1 is usually not too large, which means that the number of non-resident models temporarily requested by the mobile phone 100 is not too many.
[0080] For example, in some embodiments, the value of S1 is a preset fixed value, such as 200M. In other embodiments, the value of S1 can change synchronously with the value of the free storage capacity in the mobile phone 100. For example, when the free storage capacity of the mobile phone 100 is greater than or equal to 2GB, the mobile phone 100 can set the value of S1 to 500MB; when the free storage capacity of the mobile phone 100 is less than 2GB, the mobile phone 100 can set the value of S1 to 200MB.
[0081] Step 305: The mobile phone 100 sends a model application request to the cloud server 200. The model application request carries the to-be-classified tags and the capacity value of the available model storage space. For example, the to-be-classified tags = {C1, C2, C3}, and the capacity value of the available model storage space = S1.
[0082] Specifically, when the mobile phone 100 sends a model application request to the cloud server 200, the key elements of the interface are the to-be-classified tags = {C1, C2, C3} and the capacity value of the available model storage space = S1. This interface is the interface through which the mobile phone 100 transmits the model application request to the cloud server 200 via wireless communication.
[0083] It can be understood that since the data volume of the non-resident models requested by the mobile phone 100 is limited, that is, the data volume of the non-resident models corresponding to the to-be-classified tags is limited, these non-resident models will not occupy too much storage space of the mobile phone 100.
[0084] Step 306: The cloud server 200 obtains M non-resident models corresponding to the to-be-classified tags according to the model application request. The sum of the data volumes of these M non-resident models is less than or equal to the capacity value of the available model storage space, and M is a positive integer.
[0085] It can be understood that the non-resident models corresponding to the to-be-classified tags are the non-resident models corresponding to the image set identified by the to-be-classified tags in the mobile phone 100.
[0086] In some embodiments, generally multiple non-resident models are stored in the cloud server 200. Each non-resident model corresponds to several large categories obtained from the resident model, such as the large categories classified by the resident model. Moreover, each non-resident model is obtained by training and tuning with images of the corresponding large category and has the ability to accurately classify images of that large category.
[0087] Specifically, after the cloud server 200 receives the above model application request from the mobile phone 100, it can sequentially search for the non-resident models corresponding to the tags C1, C2, and C3 respectively, and judge the model sizes (i.e., data volumes) of these non-resident models. Furthermore, it filters out M non-resident models whose sum of data volumes is less than or equal to S1.
[0088] It can be understood that each non-resident model in the cloud server 200 has one or more labels. The labels of the non-resident models corresponding to the labels sent by the mobile phone 100 in the model application request can be exactly the same or different. For example, referring to Figure 1 , the main labels of the image sets 1-3 are C1, C2, and C3 in sequence. There are non-resident models F1-F3 with labels C1, C2, and C3 stored on the cloud server. Then, the image sets 1-3 can be respectively corresponding to the non-resident models F1-F3 by searching for the labels C1, C2, and C3.
[0089] For another example, the main labels of the image sets 1-3 are C1, C2, and C3 in sequence. There are non-resident models F1-F3 with labels (C1, C5), C2, and (C3, C4, C5) stored on the cloud server. Then, the image sets 1-3 can also be respectively corresponding to the non-resident models F1-F3 by searching for the labels C1, C2, and C3. For another example, the to-be-classified label of the image set M is other (i.e., the "other" major category), and the to-be-classified label other is pre-associated with the non-resident models with labels (C1, C2). Then, the non-resident models F2 and F3 corresponding to the image set M can be obtained by searching for the label other.
[0090] For the non-resident models requested by the mobile phone 100, the cloud server 200 needs to select according to the capacity value of the available model storage space in S1. That is to say, only the non-resident models with a total storage capacity less than the capacity value will be sent to the mobile phone 100. For example, in some embodiments, M non-resident models are randomly selected from the multiple non-resident models requested by the mobile phone 100 based on S1. For example, when the value in S1 is 200MB, and the data volumes of the non-resident model F1, the non-resident model F2, and the non-resident model 3 are 80MB, 100MB, and 150MB respectively, the mobile phone 100 can randomly select the non-resident models F1 and F2 with the sum of the data volumes less than 200MB as the above M non-resident models.
[0091] In addition, the mobile phone 100 selects the non-resident models requested based on S1 according to the model priorities. For example, referring to Figure 1 , the model selection priorities of the non-resident models F1 and F2 are higher than the model selection priority of the non-resident model F3. Of course, the cloud server 200 can also select M non-resident models from the multiple non-resident models requested by the mobile phone 100 according to other selection methods, which will not be elaborated in this embodiment of the present application.
[0092] In addition, it can be understood that in other embodiments of the present application, the cloud server 200 may also send all the non-resident models requested by the mobile phone 100 to the mobile phone 100, regardless of the storage capacity of the mobile phone 100. When the mobile phone 100 sends S1, it does not need to send the capacity value of the available model storage space either.
[0093] Step 307: The cloud server 200 sends information of M non-resident models to the mobile phone 100.
[0094] For example, the information of M non-resident models includes the non-resident model F1 and the label C1, and the non-resident model F2 and the label C2, and M is equal to 2.
[0095] In some embodiments, when the cloud server 200 sends information of non-resident models to the mobile phone 100, the key elements of the interface are {non-resident model = F1, label = C1} and {non-resident model = F2, label = C3}. This interface is the interface for the cloud server 200 and the mobile phone 100 to transmit information of non-resident models through wireless communication.
[0096] It can be understood that since the mobile phone 100 does not send the image to other devices such as the cloud server 200, the privacy leakage of the image in the mobile phone 100 is avoided.
[0097] Step 308: The mobile phone 100 uses the information of M non-resident models to perform detailed classification on the images in the image set corresponding to the M non-resident models.
[0098] Specifically, after receiving the information of the non-resident models, the mobile phone 100 can store the non-resident model F1 and the non-resident model F2 in the storage space (i.e., ROM) of the mobile phone 100. Specifically, the mobile phone 100 calls the non-resident model F1 and the non-resident model F2 from the ROM to the RAM for use.
[0099] In some embodiments of the present application, the mobile phone 100 uses non-resident models to perform detailed classification on the images in the image set that have not been detailedly classified, or re-performs detailed classification on the images in the image set that have been detailedly classified to calibrate the existing classification. For example, the mobile phone 100 uses non-resident models to perform detailed classification on the images with the label "other" in the image set, and re-performs detailed classification on the images with labels other than "other" to achieve classification calibration.
[0100] Specifically, the mobile phone 100 uses a non-resident model to update the resident model with labels for image annotation, so as to achieve detailed classification of the image or calibrate the existing classification of the image. For example, when the mobile phone 100 uses the non-resident model to add a new sub-label to an image, a new sub-image set can be created in the image set based on the new sub-label, and the image can be divided into the sub-image set. Also, when the mobile phone 100 uses the non-resident model to modify the label of an image in the image set, the image can be re-divided into the sub-image set corresponding to the modified label according to the modified label.
[0101] For example, when the mobile phone 100 uses the non-resident model F1 to perform detailed classification on the images in the image set 1, the images of the large category "chrysanthemum" can be classified into small categories such as "white chrysanthemum", "red chrysanthemum" and "yellow chrysanthemum", and the corresponding third-level labels "white chrysanthemum", "red chrysanthemum" and "yellow chrysanthemum" etc. can be added to these images respectively. For the images with the same third-level label, after being classified by the non-resident model F1, they belong to the same sub-image set. For example, all the images containing white chrysanthemums belong to the same sub-image set. In this way, the categories and accuracy of image classification supported on the terminal device side are improved.
[0102] Another example is that the mobile phone 100 uses the non-resident model F1 to perform detailed classification on all the images in the image set 1 to update the labels of all the images in the image set 1. For example, the mobile phone 100 can use the non-resident model F1 to modify the second-level label of the images containing sunflowers in the image set 1 from "others" to "sunflower", and add the second-level label "sunflower" to the label set of the images. Correspondingly, the non-resident model F1 classifies the images of the small category "sunflower" in the images of the large category "flower" in detail. In this way, the images in the image set 1 that have not been classified in detail can be refined, and the categories and accuracy of image classification supported on the terminal device side are improved.
[0103] In addition, in the scenario where the mobile phone 100 uses the resident model to misclassify an image P that contains roses but does not contain chrysanthemums as the large category "chrysanthemum", through the non-resident model F1, the image P in the image set 1 can be re-classified as the large category "rose", such as modifying the second-level label of the image P from "chrysanthemum" to "rose". Thus, the existing classification of the images in the image set 1 is calibrated, and the accuracy of image classification on the terminal device side is improved.
[0104] It can be understood that labels such as second-level labels and third-level labels can all be called sub-labels. However, the main label and the sub-label are relative. For example, in some scenarios, the sub-label of the main label "flower" is "chrysanthemum", while in other scenarios, "chrysanthemum" can be the "main label", and its sub-labels are "red chrysanthemum", "yellow chrysanthemum", etc.
[0105] According to some embodiments of the present application, after the mobile phone 100 classifies the entire set of images, it can display each image set obtained by classification and the corresponding labels (such as the main labels of the images) to the user through the user interface. For example, the mobile phone 100 displays the corresponding main label below the icon of an image set of a major category. For example, it displays the main label "flowers" below the image of the image set of the "flowers" major category.
[0106] As Figure 5 shown, the UI interface 41 with an image search function in the "Gallery" APP displayed on the mobile phone 100. In this UI interface 41, there are icons 411 of the image set 1 with the main label "flowers", icons 412 of the image set 2 with the main label "cars", icons 413 of the image set 3 with the main label "animals", etc. Specifically, the icon of each image set is the entry to that image set, that is, the entry to the sub-image sets or images in that image set. In addition, the UI interface 41 also displays a search box 414, which can be used to input the label of an image of a certain category to trigger the mobile phone 100 to display the UI interface where the images of that category are located.
[0107] It can be understood that the mobile phone 100 can update and display the icons of different image sets on the page 41 under the trigger of the user to display the icons of all image sets in the mobile phone 11a to the user. For example, a touch input of the user sliding on the area where the icons 411 - 413 are located (such as sliding from left to right) triggers the mobile phone 100 to update and display the icons of the image sets.
[0108] In some embodiments, the mobile phone 100 can update its UI interface to display the images in the image set after detailed classification by the non-resident model to the user. Combining Figure 5 , when the user operates on the icon 411 of the image set 1 in the UI interface 41, such as clicking on the icon 411 or entering the label "flowers" of the image set 1 in the search control 414, the mobile phone 100 can update the UI interface 41 to the UI interface 61 as Figure 6 shown. Among them, the UI interface 61 includes icons 611 of the "chrysanthemum" sub-image set, icons 612 of the "rose" sub-image set, icons 613 of the "sunflower" sub-image set, etc. Specifically, the icon of each sub-image set is the entry to the images in that sub-image set.
[0109] Step 309: The mobile phone 100 deletes M non-resident models.
[0110] It can be understood that after the mobile phone 100 has used M non-resident models (such as non-resident model F1 and non-resident model F2), the M non-resident models can be deleted, that is, the storage space temporarily occupied by the M non-resident models in the mobile phone 100 is released to support the normal storage of instructions for the regular services executed by the mobile phone 100. Specifically, the M non-resident models usually exist in the cloud server 200 and are temporarily fetched into the mobile phone 100 as needed. Since the mobile phone 100 can release the temporarily acquired M non-resident models in a timely manner, the M non-resident models will not occupy the storage space in the mobile phone 100 for a long time. And since the data volume of the M non-resident models is less than or equal to S1, the M non-resident models will not overly occupy the storage space of the mobile phone 100. Thus, it is beneficial to the normal operation of the regular services in the mobile phone 100. It can be understood that the mobile phone can delete the M non-resident models after each use; or it can delete the non-resident models regularly, for example, at a fixed time point every day; or prompt the user to determine whether to delete these non-resident models.
[0111] In this way, by combining the image classification capabilities of the resident model and the temporarily stored non-resident models for image classification, the number of categories supported for image classification on the mobile phone 100 side is increased, the accuracy of image classification is improved, and the privacy leakage of the images in the mobile phone 100 is avoided. In addition, each time the mobile phone 100 fetches a small number of non-resident models from the cloud server 200 as needed, it will not overly occupy the storage space of the mobile phone 100. Also, after the mobile phone 100 uses the non-resident model, deleting the non-resident model can release the storage space of the mobile phone 100, which is beneficial to the normal operation of the regular services in the mobile phone 100.
[0112] In addition, in the related art, edge learning is adopted to improve the model performance (such as computing power) on a terminal device. The principle of this method is that the terminal device learns a batch of images selected by the user as training data and retrains the image classification model (such as the above-mentioned resident model) on the terminal device, so that the image classification model in the terminal device obtains the classification ability of new categories, or improves the performance of the classification ability of existing categories (i.e., the accuracy of image classification). However, edge learning generally requires the user to manually select and label a batch of images as training data. For example, the user manually adds labels to some images to select and label these images. In this way, one problem is that the workload of manually selecting and labeling images is large and it is difficult for the user to operate. Another problem is that the quantity and quality of this batch of training data are difficult to guarantee, resulting in the performance of the image classification model after retraining being difficult to guarantee. Thus, after edge learning, the classification performance of the image classification model in the terminal device for a few categories is improved, or it can have the classification performance for new categories, but the classification performance of other categories may be reduced. That is to say, edge learning may lead to a lower classification accuracy for categories that are not classified in detail.
[0113] Figure 7 FIG. 4 shows a schematic structural diagram of a terminal device 100 according to an embodiment of the present application.
[0114] The terminal device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0115] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0116] The processor 110 may include one or more processing units. For example, the processor 110 may include an Application Processor (AP), a modem processor, a Graphics Processing Unit (GPU), an Image Signal Processor (ISP), a controller, a video codec, a Digital Signal Processor (DSP), a baseband processor, and / or a Neural-Network Processing Unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. For example, the processor 110 may classify images according to a resident model to obtain a rough classification result, such as multiple image sets, and then determine the image sets that require a non-resident model based on the rough classification result, and further perform a detailed classification on the corresponding images through the non-resident model obtained by requesting from the server.
[0117] The controller may generate operation control signals according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0118] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0119] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is only illustrative and does not constitute a structural limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.
[0120] The charging management module 140 is used to receive a charging input from a charger. The power management module 141 is used to connect to the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as the battery capacity (e.g., the battery power), the number of battery charge cycles, the battery health status (leakage, impedance), etc. For example, it monitors the battery power and determines whether the battery power exceeds a power threshold (e.g., 70%).
[0121] The wireless communication function of the terminal device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc. For example, the antenna 1, the antenna 2, the mobile communication module 150, and the wireless communication module 160 can be used to interact with the cloud server 200, send a model application request to the cloud server 200, and / or receive information about the non-resident model sent by the cloud server 200.
[0122] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals.
[0123] The mobile communication module 150 can provide wireless communication solutions such as 2G / 3G / 4G / 5G, etc. applied to the terminal device 100.
[0124] The wireless communication module 160 can provide wireless communication solutions such as wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied to the terminal device 100.
[0125] The terminal device 100 implements the display function through the GPU, the display screen 194, the application processor, etc. The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal device 100 may include one or K display screens 194, where K is a positive integer greater than 1. For example, the display function of the terminal device 100 can display the classification results of image classification.
[0126] The terminal device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, the application processor, etc.
[0127] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0128] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the terminal device 100 (such as audio data, a phone book, etc.). In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the terminal device 100 by running the instructions stored in the internal memory 121, and / or the instructions stored in the memory provided in the processor. For example, data such as a resident model and a non-resident model obtained from the cloud server 200 can be stored in the internal memory 121.
[0129] Now refer to Figure 8 , which is a block diagram of a system 800 according to an embodiment of the present application. The system 800 can be applied to the terminal device 100 or the cloud server 200. Hereinafter, the server 200 is taken as an example for illustration. Figure 8 Schematically shown is an example system 800 according to multiple embodiments. In one embodiment, the system 800 can include one or more processors 804, a system control logic 808 connected to at least one of the processors 804, a system memory 812 connected to the system control logic 808, a non-volatile memory (NVM) 816 connected to the system control logic 808, and a network interface 820 connected to the system control logic 808.
[0130] In some embodiments, the processor 804 can include one or more single-core or multi-core processors. In some embodiments, the processor 804 can include any combination of a general-purpose processor and a dedicated processor (such as a graphics processor, an application processor, a baseband processor, etc.). In an embodiment where the system 800 employs an eNB (Evolved Node B, enhanced base station) 101 or a RAN (Radio Access Network) controller 102, the processor 804 can be configured to execute various conforming embodiments, for example, as Figure 4 shown in the embodiments.
[0131] In some embodiments, the system control logic 808 can include any suitable interface controller to provide any suitable interface to at least one of the processors 804 and / or any suitable device or component communicating with the system control logic 808.
[0132] In some embodiments, the system control logic 808 may include one or more memory controllers to provide an interface to the system memory 812. The system memory 812 may be used to load and store data and / or instructions. In some embodiments, the memory 812 of the system 800 may include any suitable volatile memory, such as a suitable dynamic random access memory (DRAM).
[0133] The network interface 820 may include a transceiver to provide a radio interface for the system 800 to communicate with any other suitable devices (such as a front-end module, an antenna, etc.) through one or more networks. For example, receiving a model application request from the terminal device 200, or sending information of a non-resident model to the terminal device 200. In some embodiments, the network interface 820 may be integrated with other components of the system 800. For example, the network interface 820 may be integrated with at least one of the processor 804, the system memory 812, the NVM / memory 816, and a firmware device (not shown) having instructions. When at least one of the processors 804 executes the instructions, the system 800 implements the method as Figure 3 shown.
[0134] The network interface 820 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, the network interface 820 may be a network adapter, a wireless network adapter, a telephone modem, and / or a wireless modem.
[0135] In one embodiment, at least one of the processors 804 may be logically packaged with one or more controllers for the system control logic 808 to form a system-in-package (SiP). In one embodiment, at least one of the processors 804 may be integrated with the logic of one or more controllers for the system control logic 808 on the same die to form a system-on-chip (SoC). For example, for selecting a non-resident model corresponding to a tag to be classified.
[0136] The system 800 may further include: an input / output (I / O) device 832. The I / O device 832 may include a user interface that enables a user to interact with the system 800; the design of the peripheral component interface enables peripheral components to also interact with the system 800. In some embodiments, the system 800 further includes sensors for determining at least one of environmental conditions and location information related to the system 800.
[0137] In some embodiments, the user interface may include, but is not limited to, a display (such as a liquid crystal display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (such as a still image camera and / or a video camera), a flashlight (such as a light-emitting diode flash), and a keyboard.
[0138] In some embodiments, the Peripheral Component Interconnect (PCI) may include, but is not limited to, non-volatile memory ports, audio jacks, and power interfaces.
[0139] Embodiments of the mechanisms disclosed in this application may be implemented in hardware, software, firmware, or any combination of these implementation methods. Embodiments of this application may be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.
[0140] The program code may be applied to the input instructions to perform the various functions described in this application and generate output information. The output information may be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a Digital Signal Processor (DSP), a microcontroller, an Application Specific Integrated Circuit (ASIC), or a microprocessor.
[0141] The program code may be implemented in a high-level procedural language or an object-oriented programming language in order to communicate with the processing system. When needed, the program code may also be implemented in assembly language or machine language. In fact, the mechanisms described in this application are not limited to the scope of any specific programming language. In any case, the language may be a compiled language or an interpreted language.
[0142] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more transient or non-transient machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or via other computer-readable media. Thus, a machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, floppy disks, optical disks, optical discs, Compact Disc Read-Only Memories (CD-ROMs), magneto-optical discs, Read-Only Memories (ROMs), Random Access Memories (RAMs), Erasable Programmable Read-Only Memories (EPROMs), Electrically Erasable Programmable Read-Only Memories (EEPROMs), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in electrical, optical, acoustic, or other forms using the Internet. Thus, a machine-readable medium includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0143] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Instead, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0144] It should be noted that each unit / module mentioned in the device embodiments of the present application is a logical unit / module. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or can be implemented as a combination of multiple physical units / module. The physical implementation manner of these logical units / module themselves is not the most important. The combination of the functions implemented by these logical units / module is the key to solving the technical problems proposed by the present application. In addition, to highlight the innovative part of the present application, the above device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application. This does not mean that there are no other units / modules in the above device embodiments.
[0145] It should be noted that in the examples and descriptions of this patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0146] Although the present application has been illustrated and described by referring to certain preferred embodiments of the present application, those of ordinary skill in the art should understand that various changes can be made in form and detail without departing from the spirit and scope of the present application.
Claims
1. A method for classifying images on a terminal device, characterized in that, it includes: The terminal device classifies multiple images on the terminal device into a first type through an image classification model in the terminal device; The terminal device determines conditions for reclassifying the multiple images; The terminal device sends a reclassification request to the server, and the reclassification request is used to request to obtain a first image sub-model corresponding to the first type from the server. Among them, the reclassification request includes a target capacity, and the target capacity represents the storage capacity of the terminal device that can be used to store the image sub-model obtained from the server. The value of the target capacity changes with the change of the free storage capacity of the terminal device, and the target capacity is less than the free storage capacity; The terminal device receives a reclassification response from the server, and the reclassification response includes the first image sub-model. Among them, the sum of the data amounts of all image sub-models in the reclassification response is less than or equal to the storage capacity represented by the target capacity; The terminal device classifies the multiple images into multiple first subtypes of the first type through the first image sub-model; The terminal device deletes the first image sub-model.
2. The method according to claim 1, characterized in that, The conditions for reclassifying the multiple images include at least one of the following: The number of the multiple images is greater than or equal to a first quantity threshold; The ratio of the number of the multiple images to all images to be classified on the terminal device is greater than or equal to a second quantity threshold, where all images to be classified are all images or part of the images stored in the terminal device.
3. The method according to claim 1, characterized in that, The terminal device classifies multiple images on the terminal device into a first type through an image classification model in the terminal device, including: The terminal device classifies the multiple images into the first type through the image classification model and classifies the multiple images into multiple second subtypes of the first type.
4. The method according to claim 3, characterized in that, The conditions for reclassifying the multiple images include: The ratio of the number of images of a target subtype in the multiple images to the number of the multiple images is greater than or equal to a third quantity threshold, where the target subtype is the subtype of the images to be reclassified in the multiple images.
5. The method according to any one of claims 1 to 4, characterized in that, The conditions for reclassifying the multiple images include at least one of the following: The terminal device is in a charging state; The battery power of the terminal device is greater than or equal to a power threshold; The terminal device is connected to a wireless network; The terminal device receives an instruction from the user allowing reclassification.
6. The method according to claim 1 or 3, characterized in that, The terminal device classifies multiple images on the terminal device into a first type through an image classification model in the terminal device, including: The terminal device labels the multiple images classified as the first type with a first main label for identifying the first type through the image classification model.
7. The method according to claim 1, wherein, the reclassification request includes a first main label, and the first main label is used to identify the first type and corresponds to the first image sub-classification model.
8. The method according to claim 7, wherein, the reclassification request further includes a second main label and a third main label, and the reclassification request is further used to request to obtain from the server a second image sub-classification model corresponding to the second main label and a third image sub-classification model corresponding to the third main label; wherein, the second main label is used to identify the images classified as the second type by the image classification model, and the third main label is used to identify the images classified as the third type by the image classification model.
9. The method according to claim 8, wherein, the reclassification response further includes the second image sub-classification model, or the reclassification response further includes the second image sub-classification model and the third image sub-classification model.
10. The method according to claim 3, wherein, the terminal device classifies the multiple images into multiple first sub-types of the first type through the first image sub-classification model, including: the terminal device re-classifies the multiple images from the multiple second sub-types into the multiple first sub-types through the first image sub-classification model.
11. The method according to claim 6, wherein, the terminal device classifies the multiple images into multiple first sub-types of the first type through the first image sub-classification model, including: the terminal device labels the images belonging to different first sub-types with different first sub-labels through the image sub-classification model.
12. A method for classifying images on a terminal device, wherein, it includes: The server receives a reclassification request from the terminal device, and the reclassification request is used to request to obtain a first image sub-classification model from the server, wherein the first image sub-classification model is used to classify multiple images classified as the first type by the image classification model of the terminal device into multiple first sub-types of the first type; The server searches for the first image sub-classification model corresponding to the first type according to the reclassification request, and the reclassification request further includes a target capacity, wherein the target capacity represents the storage capacity of the terminal device that can be used to store the image sub-classification model obtained from the server, the target capacity changes with the change of the free storage capacity of the terminal device, and the target capacity is less than the free storage capacity; The server sends a reclassification response to the terminal device, and the reclassification response includes the first image sub-classification model, wherein the sum of the data amounts of all the image sub-classification models in the reclassification response is less than or equal to the storage capacity represented by the target capacity.
13. The method according to claim 12, wherein, The reclassification request includes a first main label for identifying the first type and corresponding to the first image segmentation model.
14. The method according to claim 12, wherein, the reclassification request further includes a second main label and a third main label; and the method further includes: The server determines that the image segmentation models in the reclassification response sent to the terminal device only include the first image segmentation model and the second image segmentation model corresponding to the second main label according to the target capacity in the reclassification request, wherein the sum of the data volumes of the first image segmentation model and the second image segmentation model is less than or equal to the storage capacity represented by the target capacity.
15. The method according to claim 12, wherein, the reclassification request further includes a second main label and a third main label; and the method further includes: The server determines that the image segmentation models in the reclassification response sent to the terminal device include the first image segmentation model, the second image segmentation model corresponding to the second main label, and the third image segmentation model corresponding to the third main label according to the target capacity in the reclassification request, wherein the sum of the data volumes of the first image segmentation model, the second image segmentation model, and the third image segmentation model is less than or equal to the storage capacity represented by the target capacity.
16. A machine-readable medium, wherein, instructions are stored on the machine-readable medium, and when the instructions are executed on the machine, the machine executes the method for classifying images on a terminal device according to any one of claims 1 to 15.
17. A system for classifying images on a terminal device, comprising: a memory for storing instructions executed by one or more processors of the system, and the processor, which is one of the processors of the system, for executing the method for classifying images on a terminal device according to any one of claims 1 to 15.
Citation Information
Patent Citations
Picture label generation method and device, terminal and storage medium
CN110309339A