Image classification model training method, electronic equipment and readable storage medium
By constructing positive and negative sample image training sets and using different modal information for iterative training, the problem of cumbersome data labeling in image classification model training is solved, and higher training accuracy and classification effect are achieved.
Patent Information
- Application Number
- CN202311519477.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-11-14
AI Technical Summary
When training image classification models, the prior art requires detailed data annotation of sample images, which leads to complex and cumbersome labeling process, making it difficult to achieve better classification effects.
By constructing a positive sample image training set and a negative sample image training set, iterative training is performed using information of different modes of the same thing to obtain the target image classification model. When building a negative sample training set, this method only needs to perform data annotation once, which reduces the cumbersomeness of data annotation and improves the accuracy of training.
It reduces the cumbersomeness of sample data annotation, improves the training accuracy of image classification models, and achieves better classification effects.
Smart Images

Figure CN120047762A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and particularly to a method for training an image classification model, an electronic device, and a readable storage medium. Background Art
[0002] With the development of terminal technologies, the application of image processing technologies has become increasingly widespread. For example, in the field of intelligent transportation, it can be used for road recognition, etc. Another example is in the medical field, where it can be used for medical diagnosis, etc. Among them, the recognition and classification of images are the basis of image processing technologies. In order to achieve the recognition and classification of images, an electronic device can train an image classification model.
[0003] Currently, a user can label a sample image with text information, which is used to describe the scene category to which the sample image belongs. For example, in the field of intelligent transportation, the sample image can be labeled as belonging to a viaduct scene, an urban road scene, a rain and fog scene, etc. based on the text information. The electronic device can perform model training based on the sample images labeled with different scene categories to obtain an image classification model. The trained image classification model can be used to identify the scene category corresponding to any image.
[0004] However, in the process of training an image classification model, in order to obtain an image classification model with better classification effects, it is required that the sample images can be labeled in more detail during data labeling. However, the work of labeling data is complex and cumbersome, which is difficult to achieve. Therefore, there is an urgent need for a method for training an image classification model that can improve the classification effect. Summary of the Invention
[0005] This application provides a method for training an image classification model, an electronic device, and a readable storage medium, which can be used to reduce the cumbersome degree of sample data labeling and improve the accuracy of training an image classification model. The technical solutions are as follows:
[0006] In a first aspect, a method for training an image classification model is provided, which is applied to an electronic device. The method includes:
[0007] Obtain a positive sample image training set, which includes multiple sample images and multiple first text information. The scene category to which each sample image belongs is at least one of multiple scene categories. The multiple sample images correspond one-to-one with the multiple first text information. One first text information is used to describe the scene category to which the corresponding sample image belongs; according to the positive sample image training set, determine a negative sample image training set, which includes multiple sample images. The text information of each sample image in the negative sample image training set is second text information. The second text information of any sample image in the negative sample image training set is one of the multiple first text information other than the first text information corresponding to any sample image; based on the positive sample image training set and the negative sample image training set, perform iterative training on the initial classification model to obtain a target image classification model, which can identify images belonging to at least one of multiple scene categories.
[0008] It should be noted that each sample image belonging to at least one of multiple scene categories means that any one sample image may involve multiple scene categories at the same time or may only involve one scene category.
[0009] As an example, the image content of each sample image in the negative sample image training set does not match the scene category described by the corresponding second text information.
[0010] It should be noted that the multiple sample images in the negative sample image training set are the same as the multiple sample images in the positive sample image training set. That is, the electronic device can determine the multiple sample images in the positive sample image training set as the multiple sample images in the negative sample image training set, and for the same sample image, the electronic device can adjust the text information corresponding to the sample image in the negative sample training set according to the multiple first text information. In other words, for any one sample image in the negative sample image, the text information corresponding to the sample image in the positive sample image training set is different from the sample information corresponding to the sample image in the negative sample image training set.
[0011] Since, in the process of constructing the negative sample image training set, multiple sample images in the negative sample image training set are the same as those in the positive sample image training set, and the negative sample image training set also includes multiple pieces of first text information. However, for the same sample image, the text information corresponding to the sample image in the positive sample image training set is different from the sample information corresponding to the sample image in the negative sample image training set. Thus, in the process of constructing the positive sample image training set and the negative sample image training set, only one data annotation is required, thereby reducing the complexity of data annotation and improving the efficiency of data annotation. Additionally, since, in the process of training the image classification model, information of different modalities of the same object is used, the trained image classification model is more accurate.
[0012] As an example of the present application, the operation of the electronic device to determine the negative sample image training set based on the positive sample image training set includes:
[0013] Determine the similarity between the target sample image and the first text information corresponding to each target scene category, where the target sample image is any sample image in the positive sample image training set, and each target scene category refers to each scene category among the multiple scene categories except the scene category to which the target sample image belongs; determine the first text information corresponding to the target scene category with the maximum similarity as the second text information corresponding to the target sample image; determine the multiple sample images and the second text information corresponding to each sample image as the negative sample image training set.
[0014] As an example, the electronic device can also determine the similarity between the target sample image and the first text information of each first text information that does not describe the scene category to which the target sample image belongs, and determine the first text information of the scene category that does not describe the scene category to which the target sample image belongs with the maximum similarity as the second text information.
[0015] It should be noted that by determining the first text information corresponding to the target scene category with the maximum similarity as the second text information corresponding to the target sample image, not only can the construction of negative samples be successfully completed, but also, since the similarity between the second text information and the corresponding target sample image is the highest, thus, training the image classification model to be trained with the target sample image and the corresponding second text information can improve the accuracy of the image classification model in recognizing images with relatively high similarity.
[0016] As an example of the present application, the operation of the electronic device to determine the similarity between the target sample image and the first text information corresponding to each target scene category includes:
[0017] Process the target sample image through a pre-trained target image encoder to obtain the visual feature vector of the target sample image; process the first text information corresponding to each target scene category through a pre-trained target text encoder to obtain the text feature vector of the first text information corresponding to each target scene category; determine the similarity between the visual feature vector of the target sample image and the text feature vector of the first text information corresponding to each target scene category.
[0018] It should be noted that the target text encoder and the target image encoder are obtained through coordinated training during the pre-training process.
[0019] It is worth noting that processing each first text information through a pre-trained target text encoder and processing each sample image through a target image encoder can ensure the consistency of the feature vectors.
[0020] As an example of the present application, before the electronic device determines the similarity between the target sample image and the first text information corresponding to each target scene category, it can also obtain multiple sample text information and multiple sample training images. The multiple sample text information corresponds one-to-one with the multiple sample training images, and each sample text information in the multiple sample text information is used to describe the scene category to which the corresponding sample training image belongs; iteratively train the initial text encoder based on the multiple sample text information, and iteratively train the initial image encoder based on the multiple sample training images; during the iterative training process, determine the loss value of the first loss function between the text encoder obtained after each training and the image encoder obtained after each training; in the case where the loss value converges, determine the text encoder obtained at the time of convergence as the target text encoder, and determine the image encoder obtained at the time of convergence as the target image encoder. The target text encoder is used to determine the text feature vector of the first text information corresponding to each target scene category, and the target image encoder is used to determine the visual feature vector of the target sample image.
[0021] It should be noted that the multiple sample text information corresponds one-to-one with the multiple sample training images. The one-to-one correspondence between the multiple sample text information and the multiple sample training images means that for any sample text information D, there is a unique sample training image in the multiple sample training images whose image content is the same as the content described by the any sample text information D. In addition, during the iterative training process, the corresponding sample text information and the sample training image can be trained in pairs.
[0022] Thus, by training the text encoder and the image encoder using image-text matching data (i.e., sample text information and corresponding sample training images), the two modalities of information, namely the sample training images and the corresponding sample text information, can be mapped to the same feature space, so that text encoders and image encoders with better training effects can be obtained.
[0023] As an example of the present application, the operation of the electronic device for iteratively training the initial classification model based on the positive sample image training set and the negative sample image training set to obtain the target image classification model includes:
[0024] Concatenating the visual feature vectors of each sample image in the positive sample image training set with the text feature vectors of the corresponding first text information to obtain a plurality of positive sample mixed feature vectors; concatenating the visual feature vectors of each sample image in the negative sample image training set with the text feature vectors of the corresponding second text information respectively to obtain a plurality of negative sample mixed feature vectors; iteratively training the initial classification model according to the plurality of positive sample mixed feature vectors and the plurality of negative sample mixed feature vectors; during the iterative training process, determining the loss value of the second loss function between the classification result of the image classification model obtained after each training and the preset result; when the loss value of the second loss function converges, determining the image classification model obtained at the time of convergence as the target image classification model.
[0025] It should be noted that the preset result is the classification label corresponding to each sample image in the positive sample image training set and the classification label corresponding to each sample image in the negative sample image training set. That is, in the case of training according to the sample images in the positive sample image training set, the preset result is the classification label corresponding to each sample image in the positive sample image training set, and in the case of training according to the sample images in the negative sample image training set, the preset result is the classification label corresponding to each sample image in the negative sample image training set.
[0026] Since different modalities of information of the same thing are used in the process of training the image classification model to be trained, the trained image classification model is more accurate.
[0027] As an example of the present application, after the electronic device iteratively trains the initial classification model based on the positive sample image training set and the negative sample image training set to obtain the target image classification model, it can also obtain the image to be classified; determine the visual feature vector of the image to be classified; determine the similarity between the visual feature vector of the image to be classified and each first text information among multiple first text informations, obtaining multiple similarities; splice the visual feature vector of the image to be classified with the text feature vector corresponding to each first text information among N first text informations respectively, obtaining N hybrid feature vectors, where the N first text informations are the first text informations corresponding to each similarity among the top N similarities arranged from largest to smallest according to the multiple similarities, and N is a positive integer greater than or equal to 1; process the N hybrid feature vectors through the target image classification model to obtain the classification result of the image to be classified.
[0028] Since an image classification model can identify all scene categories to which the image to be classified belongs through the target image classification model, it is possible to obtain the scene category to which the image to be classified belongs as much as possible at one time, improving the efficiency of image classification of the target image classification model. Moreover, the electronic device can adopt different display schemes according to the scene category to which the image to be classified belongs, improving the display quality and display effect of the image to be classified.
[0029] As an example of the present application, the operation of the electronic device to obtain the image to be classified includes:
[0030] When the camera is turned on, determining that the preview image captured by the camera is the image to be classified; or,
[0031] When receiving an image selection operation, determining that the image selected by the image selection operation is the image to be classified.
[0032] It should be noted that the situation where the camera is turned on includes the situation where the camera scans QR codes, barcodes and other information after being turned on, or the situation where the camera performs text recognition, object recognition after being turned on, or the situation where the camera takes pictures after being turned on, etc.
[0033] In this way, since the image to be classified can be the preview image or the image selected by the user, the electronic device can realize the scene recognition of any image, thereby increasing the application scenarios of the image classification model and improving the practicability of the image classification model.
[0034] As an example of the present application, when the camera of the electronic device is turned on, the operation of determining the preview image captured by the camera as the image to be classified includes: when the camera is turned on, displaying a scene recognition control in the shooting interface, and performing image capture through the camera to obtain a preview image, where the scene recognition control is used to control whether to perform scene recognition; in response to an activation operation on the scene recognition control, determining the preview image as the image to be classified.
[0035] As an example, when the scene recognition control is in the off state, that is, when the user does not perform an activation operation on the scene recognition control, the electronic device will not perform subsequent scene recognition (or image recognition) operations.
[0036] In this way, by setting the scene recognition control, the user can independently select whether to perform scene recognition, thereby increasing the interactivity with the user, and saving the operating resources of the electronic device when scene recognition is not required.
[0037] As an example of the present application, the image to be classified is the preview image; thus, after the electronic device determines the scene category to which the image to be classified belongs according to the classification result output by the target image classification model, it can also receive a shooting operation; in response to the shooting operation, storing the exposed preview image in the image folder corresponding to the scene category to which the preview image belongs.
[0038] As an example, in order to save storage space, the electronic device can also store the image to be classified in any one of the multiple scene categories to which the image to be classified belongs. Alternatively, the electronic device stores an image to be classified and stores the image identifier of the image to be classified and its corresponding scene category. When it is necessary to display images according to the classification method, the electronic device can obtain the image to be classified according to the image identifier of the image to be classified and display the image to be classified in the classification display interface.
[0039] In this way, by storing the exposed preview image in the image folder corresponding to the scene category to which the preview image belongs, it is convenient for the user to search for images according to the scene category, improving the interactivity with the user and the user stickiness.
[0040] As an example, when the image to be classified is the preview image, after the electronic device determines the scene category to which the image to be classified belongs, it can also display a scene label in the preview image, where the scene label is used to describe the scene category to which the image to be classified belongs.
[0041] As an example, when the image to be classified is a preview image, after the electronic device determines the scene category to which the image to be classified belongs, it can also determine the imaging scheme of the image to be classified according to the scene category of the image to be classified, such as determining the exposure parameters, filter scheme, shooting mode, display resolution, etc. of the image to be classified.
[0042] In a second aspect, an electronic device is provided. The structure of the electronic device includes a processor and a memory. The memory is used to store a program for supporting the electronic device to execute the training method of the image classification model provided in the first aspect above, and to store data involved in implementing the training method of the image classification model described in the first aspect above. The processor is configured to execute the program stored in the memory. The electronic device may further include a communication bus, and the communication bus is used to establish a connection between the processor and the memory.
[0043] In a third aspect, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the training method of the image classification model described in the first aspect above.
[0044] In a fourth aspect, a computer program product containing instructions is provided. When it runs on a computer, it causes the computer to execute the training method of the image classification model described in the first aspect above.
[0045] The technical effects obtained in the second, third, and fourth aspects above are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0047] Figure 2 is a schematic diagram of another application scenario provided by an embodiment of the present application;
[0048] Figure 3 is a schematic diagram of another application scenario provided by an embodiment of the present application;
[0049] Figure 4 is a schematic diagram of another application scenario provided by an embodiment of the present application;
[0050] Figure 5 is a schematic diagram of another application scenario provided by an embodiment of the present application;
[0051] Figure 6 is a flowchart of a training method of an image classification model provided by an embodiment of the present application;
[0052] Figure 7It is a schematic diagram of the training process of an image classification model provided by an embodiment of the present application;
[0053] Figure 8 It is a schematic flowchart of an image classification method provided by an embodiment of the present application;
[0054] Figure 9 It is a schematic diagram of the process of an image classification method provided by an embodiment of the present application;
[0055] Figure 10 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application;
[0056] Figure 11 It is a block diagram of the software system of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0057] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the implementation manners of the present application in detail with reference to the accompanying drawings.
[0058] It should be understood that the "multiple" mentioned in the present application refers to two or more. In the description of the present application, unless otherwise specified, " / " means "or", for example, A / B may represent A or B; the "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, for the convenience of clearly describing the technical solutions of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. do not necessarily limit to be different.
[0059] Referring to "one embodiment" or "some embodiments" described in the specification of the present application means that specific features, structures, or characteristics described in conjunction with the embodiment are included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. appearing in different parts of this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0060] With the development of terminal technology, the application of image processing technology has become increasingly widespread. Among them, image recognition and classification are the basis of image processing technology. Usually, electronic devices can perform image recognition or scene recognition through an image classification model. In some scenarios, in order to enable users to quickly find the images they need, a mobile phone can classify the images in the gallery according to the shown scene categories, such as Figure 1 as shown. To facilitate image classification, a pre-trained image classification model can be built into the mobile phone. That is to say, before leaving the factory, a pre-trained image classification model can be built into the mobile phone. Among them, the image classification model can be obtained by other electronic devices through model training based on sample images of different annotated scene categories. The sample images of different annotated scene categories refer to the staff annotating different sample images through text information, and the text information is used to describe the scene category to which the sample image belongs. For example, the sample image can be annotated as belonging to a viaduct scene, an urban road scene, a rain and fog scene, a building scene, a food scene, a night scene, a document scene, a text scene, etc. through text information.
[0061] However, in the process of training an image classification model, in order to obtain an image classification model with better classification effects, it is required that the sample images can be more detailedly annotated during data annotation, and the work of annotating data is complex and cumbersome, which is difficult to achieve. Of course, there are also some image classification models that only use a single piece of information (such as only using image information) as training data, resulting in the trained image classification model being unable to recognize some complex scenes.
[0062] In order to improve the classification accuracy of an image classification model and reduce the difficulty of annotating data. An embodiment of the present application provides a method for training an image classification model. In this method, an electronic device can obtain a positive sample image training set, which includes multiple sample images and multiple first text information. The scene category to which each sample image belongs is at least one of multiple scene categories, and each first text information can describe the scene category to which the corresponding sample image in the multiple sample images belongs; determine a negative sample image training set according to the positive sample image training set. The negative sample image training set includes multiple sample images, and the text information corresponding to each sample image in the multiple sample images in the negative sample image training set is second text information, and the second text information is one of the first text information except the first text information corresponding to any one sample image among the multiple first text information; based on the positive sample image training set and the negative sample image training set, perform iterative training on the initial image classification model to obtain a target image classification model that can identify images belonging to at least one of multiple scene categories. Since in the process of constructing the negative sample image training set, the multiple sample images in the negative sample image training set are the multiple sample images in the positive sample image training set, and the negative sample image training set also includes multiple first text information. However, for the same sample image, the text information corresponding to the sample image in the positive sample image training set is different from the sample information corresponding to the sample image in the negative sample image training set. In this way, in the process of constructing the positive sample image training set and the negative sample image training set, only one data annotation is required, thereby reducing the complexity of data annotation and improving the efficiency of data annotation. In addition, since in the process of training the image classification model to be trained, information of different modalities of the same thing is used, the trained image classification model is more accurate.
[0063] For ease of understanding, before introducing the method provided by the embodiments of the present application in detail, the application scenarios involved in the embodiments of the present application will be introduced next.
[0064] Please refer to Figure 2 , Figure 2 which is a schematic diagram of an application scenario provided by the embodiments of the present application. In one application scenario, when a user is using a mobile phone, the user may take a photo through the camera of the mobile phone. After the mobile phone starts the camera, it can display on the display screen such as Figure 2The shooting interface shown in Figure (a) therein. A preview image A is displayed in this shooting interface, and the preview image A includes a blue sky, grassland, and trees. When the mobile phone captures the preview image A, it can input the preview image A into the target image classification model. This target image classification model can identify images belonging to at least one scene category among multiple scene categories. For example, this target image classification model can identify images corresponding to the building scene category, the tree scene category, the grassland scene category, the blue sky scene category, the flower scene category, the night scene category, the beach scene category, the ocean scene category, the playground scene category, the school scene category, etc. The target image classification model can perform scene recognition (or image recognition, image classification) on the preview image A, and the target image classification model can output the classification result of the preview image A. See Figure 2 Figure (b) therein. The mobile phone can, according to the classification result, display the scene label of the scene category to which the preview image A belongs on the preview image A. For example, if it is determined that the scene categories to which the preview image A belongs include the sky scene category, the grassland scene category, and the tree scene category, then the mobile phone can display the scene labels "sky", "grassland", and "tree" at the positions of the blue sky, grassland, and trees in the preview image A.
[0065] In another application scenario, the mobile phone can also adjust the image display scheme according to the scene category to which the preview image A belongs. For example, adjust the shooting mode, adjust the exposure parameter, adjust the display resolution, etc. Exemplarily, if the scene category to which the preview image A belongs is the overexposed scene category, such as this overexposed scene category includes the moon scene category, the flame scene category, the sun scene category, the electric lamp scene category, etc., that is, when the preview image A includes light sources such as the moon, a light source, or a flame, the scene category to which the preview image A belongs usually includes the overexposed scene category. In this case, the mobile phone can adjust the shooting mode or reduce the exposure parameter of the preview image A. As shown in Figure 3 Figure (a) therein. If the preview image A includes the moon, then the mobile phone, through the target image classification model, identifies that the scene category to which the preview image A belongs includes the moon scene category, and then the mobile phone can adjust the shooting mode to the full moon mode. In the full moon mode, the exposure parameter of the preview image A changes. That is, the mobile phone can reduce the exposure degree of the preview image A and shorten the exposure duration of the preview image A, and display the preview image B shown in Figure 3 Figure (b) therein.
[0066] In yet another application scenario, see Figure 4In Figure (a), if the user is satisfied with the current preview image A, the user can click on the shooting control P1 in the shooting interface; in response to the click operation on the shooting control P1, the mobile phone can perform image exposure based on the preview image A to obtain the preview image A after exposure, and the mobile phone can store the preview image A after exposure in the image folder corresponding to the scene category to which the preview image A belongs, that is, the mobile phone can store the preview image A after exposure in at least one of the image folders corresponding to the sky scene, the grassland scene, and the tree scene, such as storing it in the image folder corresponding to the sky scene category. After that, if the user needs to search for images of a certain scene category taken, refer to Figure 4 In Figure (b), the user can click on the application icon of the gallery application in the desktop; in response to the click operation on the application icon of the gallery application, the mobile phone can display as Figure 4 shown in the image preview interface P2 in Figure (c). The user can click on the "Discover" control P3 displayed in the image preview interface P2; in response to the click operation on the "Discover" control P3, the mobile phone can display as Figure 4 shown in the classification view interface P4 in Figure (d). In this classification view interface P4, multiple image folders can be displayed. For example, the "People" folder, the "Places" folder, and the "Things" folder, etc. can be displayed. In the "Things" folder, folders such as "Buildings", "Sky", "Portraits", "Grassland", "Trees", and "Night Scenes" can be included. In this way, the user can continue to search for more detailed scene categories in the "Things" folder.
[0067] Please refer to Figure 5 , Figure 5 which is a schematic diagram of an application scenario provided by an embodiment of the present application. In another application scenario, after the mobile phone starts the camera, in the shooting interface shown in Figure (a) as Figure 5 not only can the preview image A collected by the camera be displayed, but also a scene recognition control can be displayed in the shooting interface. If the user needs to identify the scene category of the taken image, then the user can click on the scene recognition control P5; in response to the click operation on the scene recognition control P5, refer to Figure 5 In Figure (b), the mobile phone can change the display mode of the scene recognition control P5 (in this embodiment of the present application, it is described by taking the display color of the scene recognition control P5 changing from white background and black characters to black background and white characters as an example), and input the preview image A into the target image classification model. The target image classification model processes the preview image A respectively, and the target image classification model can output the classification result of the preview image A.
[0068] It should be noted that only the above is used in the embodiments of the present applicationFigures 2 - 5 The scenario shown is used as an example for illustration and does not constitute a limitation on the embodiments of the present application.
[0069] Based on the application scenarios provided in the above embodiments, the training method of the image classification model provided in the embodiments of the present application is introduced below. Figure 6 , Figure 6 The present invention is a flowchart of a training method for an image classification model according to an exemplary embodiment. As an example but not a limitation, the method is described by taking application in an electronic device as an example. The method may include some or all of the following contents:
[0070] Step 601: Obtain a positive sample image training set.
[0071] It should be noted that the positive sample image training set includes multiple sample images and multiple first text information. The scene category to which each sample image belongs is at least one of the multiple scene categories. The multiple sample images correspond one-to-one to the multiple first text information. One first text information is used to describe the scene category to which the corresponding sample image belongs.
[0072] As an example, a one-to-one correspondence between multiple sample images and multiple first text information means that for any one first text information C among the multiple first text information, there is only one sample image among the multiple sample images whose image content is the same as the content described by the any one first text information C.
[0073] For example, one of the first text information among the multiple first text information may be "flower, a photo of flower sea" or "flower, an image of a flower sea", then there is a sample image whose image content is a flower or a flower sea among the multiple sample images. And the first text information describes that the scene category to which the sample image belongs is a flower scene category.
[0074] It should be noted that the scene category to which each sample image belongs is at least one of the multiple scene categories, which means that any sample image may involve multiple scene categories at the same time, or may involve only one scene category.
[0075] Exemplarily, the multiple scene categories may include a blue sky scene category, a tree scene category, a flower scene category, a cat scene category, a dog scene category, a sea scene category, a building scene category, a portrait scene category, a night scene category, a snow scene category, etc. If the image content of a sample image includes objects such as a blue sky, flowers, and buildings, then the scene categories to which the sample image belongs include the blue sky scene category, the flower scene category, and the building scene category, and the first text information corresponding to the sample image is "An image of a blue sky, flowers, and buildings", and this first text information describes three scene categories. If the image content of a sample image includes a cat, then the scene category to which the sample image belongs is the cat scene category, and the first text information corresponding to the sample image is "An image of a cat".
[0076] As an example, for any one of the multiple sample images, the electronic device may receive an input operation for the any one sample image and determine the text information carried by the input operation as the first text information corresponding to the any one sample image. Alternatively, the electronic device may perform a text scraping operation on a specific page and determine the scraped text information as the first text information for the any one sample image, where the specific page is a page describing the scene category to which the any one sample image belongs, and the scene category to which the any one sample image belongs is described by text information in the specific page.
[0077] Step 602: Determine a negative sample image training set according to the positive sample image training set.
[0078] It should be noted that the negative sample image training set includes multiple sample images. The text information of each sample image in the negative sample image training set is the second text information, and the second text information of any one sample image in the negative sample image training set is one of the first text information other than the first text information corresponding to the any one sample image.
[0079] As an example, the image content of each sample image in the negative sample image training set does not match the scene category described by the corresponding second text information.
[0080] In some embodiments, the operation of the electronic device to determine the negative sample image training set according to the positive sample image training set includes: determining the similarity between the target sample image and the first text information corresponding to each target scene category, where the target sample image is any sample image in the positive sample image training set, and each target scene category refers to each scene category in the multiple scene categories except the scene category to which the target sample image belongs; determining the first text information corresponding to the target scene category with the largest similarity as the second text information corresponding to the target sample image; and determining the multiple sample images and the second text information corresponding to each sample image as the negative sample image training set.
[0081] It should be noted that the multiple sample images in the negative sample image training set are the same as the multiple sample images in the positive sample image training set. That is, the electronic device can determine the multiple sample images in the positive sample image training set as the multiple sample images in the negative sample image training set, and for the same sample image, the electronic device can adjust the text information corresponding to the sample image in the negative sample training set according to the multiple first text information. In other words, for any sample image in the negative sample images, the text information corresponding to the sample image in the positive sample image training set is different from the sample information corresponding to the sample image in the negative sample image training set.
[0082] In some embodiments, for the target sample image (any sample image in the negative sample image set, or any sample image in the multiple sample images), the electronic device can determine any other first text information other than the first text information corresponding to the target sample image (the first text information corresponding to the target sample image in the positive sample image training set) as the second text information corresponding to the target sample image in the negative sample image training set. Of course, in order to improve the accuracy of the trained image classification model in classifying similar images, the electronic device can determine the first text information corresponding to the target scene category with the largest similarity as the second text information corresponding to the target sample image.
[0083] Exemplarily, the image content of the target sample image is a dog, and the first text information corresponding to the target sample image in the positive sample image training set is "an image of a dog", that is, the scene category corresponding to the target sample image is the dog scene category. Then, the electronic device can determine the first text information corresponding to other scene categories in the multiple scene categories except the dog scene category (wherein, if the first text information describes both the dog scene category and the grassland scene category, then the electronic device can also obtain the text information describing the grassland scene category in the first text information, but will not obtain the text information describing the dog scene category in the first text information). After that, the electronic device can determine the similarity between the target sample image and the first text information corresponding to the determined other scene categories, and determine the first text information corresponding to the target scene category with the highest similarity as the second text information corresponding to the target sample image.
[0084] As an example, the electronic device can also determine the similarity between the target sample image and each first text information that does not describe the scene category to which the target sample image belongs, and determine the first text information that does not describe the scene category to which the target sample image belongs and has the highest similarity as the second text information.
[0085] Exemplarily, the image content of the target sample image is a dog, and the first text information corresponding to the target sample image in the positive sample image training set is "an image of a dog", that is, the scene category corresponding to the target sample image is the dog scene category. Then, the electronic device can determine the first text information that does not include the description of the dog scene category, determine the similarity between the determined target sample image and each first text information that does not describe the dog scene category, and determine the first text information that does not describe the dog scene category and has the highest similarity as the second text information.
[0086] It should be noted that by determining the first text information corresponding to the target scene category with the highest similarity as the second text information corresponding to the target sample image, not only can the construction of negative samples be successfully completed, but also, since the similarity between the second text information and the corresponding target sample image is the highest, thus, by using the target sample image and the corresponding second text information to train the image classification model to be trained, the accuracy of the image classification model in recognizing images with higher similarity can be improved.
[0087] In some embodiments, the operation of the electronic device to determine the similarity between the target sample image and the first text information corresponding to each target scene category in the target scene category includes: processing the target sample image through a pre-trained target image encoder to obtain a visual feature vector of the target sample image; processing the first text information corresponding to each target scene category through a pre-trained target text encoder to obtain a text feature vector of the first text information corresponding to each target scene category; and determining the similarity between the visual feature vector of the target sample image and the text feature vector corresponding to each target scene category.
[0088] As an example, the electronic device may determine at least one of the Euclidean distance, cosine distance, and Jaccard distance, etc. between the visual feature vector of the target sample image and the text feature vector corresponding to each target scene category. In the case of determining one of the distances, the obtained distance is determined as the similarity between the visual feature vector of the target sample image and the text feature vector corresponding to each target scene category; in the case of determining multiple distances, the average value of the multiple distances is determined as the similarity between the visual feature vector of the target sample image and the text feature vector corresponding to each target scene category; or, in the case of determining multiple distances, different weights are assigned to each distance and then added together, and the obtained sum is determined as the similarity between the visual feature vector of the target sample image and the text feature vector corresponding to each target scene category. The embodiments of the present application do not make specific limitations on this.
[0089] It should be noted that the target text encoder and the target image encoder are obtained through mutual cooperation training in the pre-training process.
[0090] It is worth noting that by processing each first text information through the pre-trained target text encoder and processing each sample image through the target image encoder respectively, the consistency of the feature vectors can be ensured.
[0091] In some embodiments, before the electronic device determines the similarity between the target sample image and the first text information corresponding to each target scene category in the target scene category, the target text encoder and the target image encoder may also be pre-trained.
[0092] As an example, an electronic device can obtain multiple sample text information and multiple sample training images, the multiple sample text information correspond one-to-one to the multiple sample training images, and each sample text information in the multiple sample text information is used to describe the scene category to which the corresponding sample training image belongs; iteratively train the initial text encoder based on the multiple sample text information, and iteratively train the initial image encoder based on the multiple sample training images; during the iterative training process, determine the loss value of the first loss function between the text encoder obtained after each training and the image encoder obtained after each training; when the loss value converges, determine the text encoder obtained at the time of convergence as the target text encoder, and determine the image encoder obtained at the time of convergence as the target image encoder, the target text encoder is used to determine the text feature vector of the first text information corresponding to each target scene category, and the target image encoder is used to determine the visual feature vector of the target sample image.
[0093] It should be noted that the one-to-one correspondence between the multiple sample text information and the multiple sample training images means that for any one sample text information D among the multiple sample text information, there is only one sample training image among the multiple sample training images whose image content is the same as the content described by the any one sample text information D.
[0094] It should also be noted that, during the iterative training process, the corresponding sample text information and the sample training image can be paired for training. For example, the sample text information is "a photo of a cat" or "an image of a cat", and the corresponding sample training image is an image of a cat. The sample text information is input into the text encoder to train the text encoder, and the sample training image is input into the image encoder to train the image encoder.
[0095] It is worth noting that, since the text encoder and the image encoder are trained using image-text matching data (i.e., sample text information and corresponding sample training images), the two modal information of sample training images and corresponding sample text information can be mapped to the same feature space. By calculating the similarity between the sample training images and the sample text information, the similarity between the sample training images and their corresponding sample text information is constrained to be the highest, and the similarity with other sample text information is low, thereby obtaining a text encoder and an image encoder with better training effect.
[0096] In some embodiments, during the iterative training process, the operation of determining the loss value of the first loss function between the text encoder obtained after each training and the image encoder obtained after each training includes: during the iterative training process, determining the sample text feature vector output by the text encoder obtained after each training and the sample visual feature output by the image encoder obtained after each training, and then determining the loss value of the first loss function according to the sample text feature vector and the sample visual feature vector.
[0097] Exemplarily, referring to Figure 7 , since during the iterative training of the text encoder and the image encoder by the electronic device, multiple sample text information can be input into the text encoder, and multiple sample training images can be input into the image encoder. The text encoder can determine the sample text feature vector of the input sample text information, the image encoder can determine the sample visual feature vector of the input sample training image, and then the electronic device determines the loss value of the first loss function according to the sample text feature vector and the sample visual feature vector.
[0098] As an example, the first loss function can be an Image-Text Contrastive (ITC) loss function, and the first loss function can be shown as the following first formula (1).
[0099]
[0100] It should be noted that in the above first formula (1), D w represents the Euclidean distance between two sample feature vectors (X 1 and X 2 ), P is the dimension of the feature, Y is the label indicating whether the two sample feature vectors match, where Y being 1 represents that the two samples are similar or match, Y being 0 indicates that the two samples are not similar or do not match, m is a set threshold, and L(W, (Y, X 1 , X 2 )) is the loss value.
[0101] As an example, the first loss function can be represented as the above first formula (1), and of course, it can also be represented in other ways. Exemplarily, the first loss function can also be represented as the following second formula (2).
[0102]
[0103] It should be noted that in the above second formula (2), L itc is the loss value, s(I, T) and s(T, I) are the similarities between vectors, ω and v are network parameters, g v (v cls ) is the text feature vector, g’w (w’ cls ) is a visual feature vector, M is the number of scene categories, is the matching similarity between any text feature vector and the corresponding visual feature vector. is the matching similarity between any visual feature vector and the corresponding text feature vector.
[0104] As an example, the case where the loss value converges is the case where the loss value is less than or equal to the first preset value, and / or the change in the loss value is less than or equal to the second preset value. Both the first preset value and the second preset value can be set in advance according to requirements.
[0105] It should be noted that by training the text encoder and the image encoder, the target text encoder and the image encoder are obtained, making the training accuracy higher.
[0106] In some embodiments, the electronic device can not only, when the loss value converges, determine the text encoder obtained at the time of convergence as the target text encoder and determine the image encoder obtained at the time of convergence as the target image encoder. The electronic device can also determine the number of training iterations. When the number of training iterations is greater than or equal to the number threshold, the electronic device can determine the text encoder obtained when the number of training iterations is greater than or equal to the number threshold as the target text encoder, and determine the image encoder obtained when the number of training iterations is greater than or equal to the number threshold as the target image encoder.
[0107] It should be noted that the number threshold can also be set in advance according to requirements. For example, the number threshold can be 100 times, 150 times, etc.
[0108] In some embodiments, the electronic device can not only determine the text feature vector of each first text information and the visual feature of each sample image in the above manner, but also determine the text feature vector of each first text information and the visual feature vector of each sample image in other ways. Exemplarily, the electronic device can determine the visual feature vector of each sample image through a visual base network model (such as a Resent (Residual Neural Network) model, etc.), and determine the text feature vector of each first text information through a multilingual text model. Or, the electronic device performs vectorization processing on each first text information through a text vector model to obtain the text feature vector of each first text information, and the text vector model can be determined based on a pre-trained language representation model (Bidirectional Encoder Representation from Transformers, BETR). The embodiments of the present application do not make specific limitations on this.
[0109] Since the negative sample image training set is determined based on the positive sample image training set, generally, to ensure the balance of model training, the ratio between the number of sample images in the positive sample image training set and the number of sample images in the negative sample image training set is usually 1:1. Of course, it can also be other ratios. For example, if some sample images are selected from multiple sample images as the sample images in the negative sample training set, then the ratio between the number of sample images in the positive sample image training set and the number of sample images in the negative sample image training set can also be 2:1, 1.5:1, etc. The embodiments of the present invention do not make specific limitations on this.
[0110] In some embodiments, each sample image in the positive sample image training set corresponds to a classification label, and each sample image in the negative sample training set also corresponds to a classification label, where the classification label is used to indicate the relationship between the sample image and the corresponding text information. The classification label can be [0, 1] and [1, 0], and the first element in the classification label represents the probability that the sample image does not match the corresponding text information, and the second element represents the probability that the sample image matches the corresponding text information. That is, the classification label of each sample image in the positive sample image training set can be [0, 1]. The first element 0 represents that the probability that the sample image does not match the corresponding first text information is 0, and the second element 1 represents that the probability that the sample image matches the corresponding first text information is 1. The classification label of each sample image in the negative sample image training set can be [1, 0]. The first element 1 represents that the probability that the sample image does not match the corresponding second text information is 1, and the second element 0 represents that the probability that the sample image matches the corresponding second text information is 0.
[0111] It should be noted that only the above classification labels are taken as examples in the embodiments of the present application, which do not constitute a limitation to the embodiments of the present application. The classification label can also be other styles of labels.
[0112] Step 603: Based on the positive sample image training set and the negative sample image training set, perform iterative training on the initial classification model to obtain a target image classification model.
[0113] It should be noted that the target image classification model can identify images belonging to at least one scene category among multiple scene categories. The initial classification model can be an Image Text Matching (ITM) module.
[0114] In some embodiments, the operation of the electronic device to iteratively train the initial classification model based on the positive sample image training set and the negative sample image training set to obtain the target image classification model includes: concatenating the visual feature vectors of each sample image in the positive sample image training set with the text feature vectors of the corresponding first text information to obtain a plurality of positive sample mixed feature vectors; concatenating the visual feature vectors of each sample image in the negative sample image training set with the text feature vectors of the corresponding second text information respectively to obtain a plurality of negative sample mixed feature vectors; iteratively training the initial classification model according to the plurality of positive sample mixed feature vectors and the plurality of negative sample mixed feature vectors; during the iterative training process, determining the loss value of the second loss function between the classification result of the image classification model obtained after each training and the preset result; when the loss value of the second loss function converges, determining the image classification model obtained at the time of convergence as the target image classification model. This process can refer to Figure 7 the training schematic diagram shown.
[0115] It should be noted that the preset result is the classification label corresponding to each sample image in the positive sample image training set and the classification label corresponding to each sample image in the negative sample image training set. That is, in the case of training according to the sample images in the positive sample image training set, the preset result is the classification label corresponding to each sample image in the positive sample image training set, and in the case of training according to the sample images in the negative sample image training set, the preset result is the classification label corresponding to each sample image in the negative sample image training set.
[0116] In some embodiments, during the process of the electronic device iteratively training the initial classification model according to the plurality of positive sample mixed feature vectors and the plurality of negative sample mixed feature vectors, the electronic device can also perform dimension elevation or dimension reduction processing on each mixed feature vector (including each positive sample mixed feature vector in the plurality of positive sample mixed feature vectors and each negative sample mixed feature vector in the plurality of negative sample mixed feature vectors) according to the dimension of the model parameters of the initial classification model, so that the dimension of each mixed feature is the same as the dimension of the network parameters of the target image classification model. Then, according to the processed plurality of mixed feature vectors, the initial classification model is iteratively trained.
[0117] As an example, the second loss function can be an ITM loss function. Of course, it can also be other loss functions, and the embodiments of the present application do not make specific limitations in this regard.
[0118] It should be noted that the second loss function can be represented by the following third formula (3).
[0119] L itm =E (I,T′) ~DH(y itm ,p itm(I, M)) (3)
[0120] In the embodiments of the present application, during the process of constructing the negative sample image training set, multiple sample images in the negative sample image training set are the same as those in the positive sample image training set, and the negative sample image training set also includes multiple pieces of first text information. However, for the same sample image, the text information corresponding to the sample image in the positive sample image training set is different from the sample information corresponding to the sample image in the negative sample image training set. In this way, during the process of constructing the positive sample image training set and the negative sample image training set, only one data annotation is required, thereby reducing the complexity of data annotation and improving the efficiency of data annotation. In addition, since different modality information of the same thing is used during the process of training the image classification model, the trained image classification model is more accurate.
[0121] It should be noted that when the electronic device obtains the target image classification model, the electronic device can perform scene recognition (or image recognition, or image classification) on images belonging to different scene categories through the target image classification model. To understand the embodiments of the present application, the manner in which the electronic device recognizes the scene category of an image through the target image classification model will be explained below. In addition, the inference process of the target image classification model is the same as the process of performing image recognition using the target image classification model, and the embodiments of the present application will not explain the inference process of the target image classification model again.
[0122] Please refer to Image 8, Figure 8 is a schematic flowchart of a method for image classification shown according to an exemplary embodiment. By way of example and not limitation, this method is described by taking an application to an electronic device as an example, and this method may include the following parts or all of the content:
[0123] Step 801: Obtain an image to be classified.
[0124] As an example, the operations for the electronic device to obtain an image to be classified include: when the camera is turned on, determining the preview image captured by the camera as the image to be classified; or, when an image selection operation is received, determining the image selected by the image selection operation as the image to be classified. It can be seen from this that the image to be classified can be any image. For example, the image to be classified can be an image downloaded from the network, or a preview image captured by the camera of the electronic device, or any image stored in the electronic device.
[0125] It should be noted that the situation where the camera is turned on includes the situation where the camera scans information such as QR codes and barcodes after being turned on, or the situation where the camera performs text recognition and object recognition after being turned on, or the situation where the camera performs shooting after being turned on, etc.
[0126] As an example, the image selection operation may refer to the selection operation of any image in the gallery of the electronic device (constituted by stored images), or the download operation, save operation, etc. of network images. The embodiments of the present application do not make specific limitations on this.
[0127] It should be noted that since the image to be classified can be a preview image or an image selected by the user, the electronic device can implement scene recognition for any image, thereby increasing the application scenarios of the target image classification model and improving the practicability of the target image classification model.
[0128] In some embodiments, when the camera is turned on, the operation for the electronic device to determine that the preview image collected by the camera is the image to be classified includes: when the camera is turned on, a scene recognition control is displayed in the shooting interface, and image acquisition is performed through the camera to obtain a preview image. The scene recognition control is used to control whether to perform scene recognition; in response to the opening operation of the scene recognition control, it is determined that the preview image is the image to be classified. Exemplarily, the scene can refer to the Figure 5 application scenarios shown above.
[0129] Since scene recognition is not required for all scenes, in order to enable the user to selectively perform scene recognition, the electronic device can also display a scene recognition control in the shooting interface.
[0130] As an example, when the scene recognition control is in the closed state, that is, when the user does not perform the opening operation on the scene recognition control, the electronic device will not perform the subsequent scene recognition (or image recognition) operation.
[0131] It should be noted that by setting the scene recognition control, the user can independently select whether to perform scene recognition, thereby increasing the interactivity with the user and saving the operating resources of the electronic device when scene recognition is not required.
[0132] Step 802: Determine the visual feature vector of the image to be classified.
[0133] As an example, the electronic device can process the visual feature vector of the image to be classified through the above-mentioned target image encoder to obtain the visual feature vector of the image to be classified, as Figure 9 shown. Alternatively, the electronic device can also determine it in other ways, such as through the above-described visual basic network model (such as the Resent (Residual Neural Network) model, etc.) to determine the visual feature vector of the image to be classified. The embodiments of the present application do not make specific limitations on this.
[0134] Step 803: Determine the similarity between the visual feature vector of the image to be classified and each of the multiple first text messages, obtaining multiple similarities.
[0135] It should be noted that the electronic device may store the first text feature information corresponding to each scene category in multiple scene categories and / or the text feature vectors of the first text feature information corresponding to each scene category. Thus, as Figure 7 shown, the electronic device may determine the similarity between the visual feature vector of the image to be classified and each of the multiple first text messages, obtaining multiple similarities.
[0136] In some embodiments, the electronic device may determine at least one of the Euclidean distance, cosine distance, and Jaccard distance, etc. between the visual feature vector of the image to be classified and each text feature vector. When determining one of the distances, the obtained distance is determined as the similarity between the visual feature vector of the image to be classified and each text feature vector; when determining multiple distances, the average of the multiple distances is determined as the similarity between the visual feature vector of the image to be classified and each text feature vector; or, when determining multiple distances, different weights are assigned to each distance and then added, and the obtained sum is determined as the similarity between the visual feature vector of the image to be classified and each text feature vector. The embodiments of the present application do not make specific limitations on this.
[0137] Step 804: Concatenate the visual feature vector of the image to be classified with the text feature vectors corresponding to each of the N first text messages respectively, obtaining N hybrid feature vectors.
[0138] It should be noted that the N first text messages are the first text messages corresponding to each of the top N similarities after arranging the multiple similarities from largest to smallest, and N is a positive integer greater than or equal to 1.
[0139] As an example, when the electronic device obtains multiple similarities, it may sort the multiple similarities in descending order to obtain a first sorting result. Obtain the text feature vectors of the first text messages corresponding to each of the top N similarities in the first sorting result; concatenate each of the obtained N text feature vectors with the visual feature vector to be classified, obtaining N hybrid feature vectors.
[0140] As an example, when the electronic device obtains multiple similarities, it can also sort the multiple similarities in ascending order to obtain a second sorting result. Obtain the text feature vectors corresponding to each of the similarities among the last N similarities in the second training result. Concatenate each of the obtained N text feature vectors with the visual feature vector to be classified to obtain N mixed feature vectors.
[0141] As an example, when the electronic device obtains multiple similarities, it can also traverse the multiple similarities. And each time it traverses, it obtains the maximum similarity among the similarities being traversed, and then continues to traverse the remaining similarities except the maximum similarity being traversed to continue obtaining the maximum similarity among the remaining similarities; repeat the traversal operation until the electronic device obtains N similarities; then, the electronic device obtains the text feature vectors corresponding to each of the N similarities. Concatenate each of the obtained N text feature vectors with the visual feature vector to be classified to obtain N mixed feature vectors.
[0142] In some embodiments, the operations in steps 803 and 804 above can be implemented by the electronic device through the target image classification model or can be implemented without the target image classification model. That is, after the electronic device obtains the visual feature vector of the image to be classified, it can input the visual feature vector of the image to be classified into the target image classification model. Multiple text feature vectors of the first text information can be stored in the target image classification model. The electronic device can determine the similarities between the visual feature vector of the image to be classified and each first text feature vector through the target image classification model to obtain multiple similarities; then concatenate the visual feature vector of the image to be classified with each of the N text feature vectors respectively to obtain N mixed feature vectors. Or, see Figure 9 , after the electronic device obtains N mixed feature vectors without passing through the target image classification model, it inputs the N mixed feature vectors into the target image classification model, and then performs the operation in step 805 below.
[0143] Step 805: Process the N mixed feature vectors through the target image classification model to obtain the classification result of the image to be classified.
[0144] In some embodiments, the electronic device can perform dimension increasing or dimension decreasing processing on each of the N mixed feature vectors according to the dimension of the network parameters of the target image classification model, so that the dimension of each of the N mixed feature vectors is the same as the dimension of the network parameters of the target image classification model. Then perform relevant classification processing on the N mixed features after dimension increasing or dimension decreasing to obtain the classification result of the image to be classified.
[0145] Since the target image classification model can identify images belonging to at least one scene category among multiple scene categories, the electronic device can obtain multiple classification results for the image to be classified. The electronic device can determine the scene category to which the image to be classified belongs according to the multiple classification results.
[0146] It should be noted that the electronic device can represent the classification result through the classification label mentioned in step 601 above. Of course, it can also be represented in other ways. For example, it can be represented by at least one of information such as letters, numbers, patterns, identifiers, etc. The embodiments of the present application do not make specific limitations in this regard.
[0147] Exemplarily, each classification result in the multiple classification results output by the target image classification model can be represented by a letter. When the letter output by the target image classification model for a scene category E that can be recognized is "yes", it indicates that the scene category to which the image to be classified belongs is the scene category E. When the letter output by the target image classification model for the scene category E is "no", it indicates that the scene category to which the image to be classified belongs is not the scene category E.
[0148] Since there may be things corresponding to multiple scene categories in one image, the image to be classified can belong to at least one scene category among multiple scene categories at the same time. In this case, the electronic device can determine the scene category to which the image to be classified belongs as the at least one scene category.
[0149] Exemplarily, the image to be classified includes a dog, grassland, trees, and a blue sky. The target image classification model can recognize images belonging to the sky scene category, grassland scene category, night scene category, cute pet scene category, portrait scene category, snow scene category, etc. The electronic device can input N mixed feature vectors corresponding to the image to be classified into the target image classification model. The target image classification model can perform scene recognition on the image to be classified and output multiple classification results for each scene category. Each classification result in the multiple classification results can be represented by the above classification labels. Among them, for the sky scene category, the classification label output by the target image classification model is [0, 1]. For the grassland scene category, the classification label output by the target image classification model is [0, 1]. For the night scene category, the classification label output by the target image classification model is [1, 0]. For the cute pet scene category, the classification label output by the target image classification model is [0, 1]. For the portrait scene category, the classification label output by the target image classification model is [1, 0], and for the snow scene category, the classification label output by the target image classification model is [1, 0]. From the classification labels output by the target image classification model for each scene category, the electronic device can determine that the scene category to which the image to be classified belongs can be the cute pet scene category, grassland scene category, and sky scene category.
[0150] As an example, when the image to be classified is a preview image, if the user is satisfied with the preview image, then the user can trigger a shooting operation, and the electronic device can receive the shooting operation; in response to the shooting operation, the electronic device can store the exposed preview image in the image folder corresponding to the scene category to which the preview image belongs. Exemplarily, the scene can refer to the Figure 4 scene shown above.
[0151] As can be seen from the above, the scene category to which the image to be classified belongs may be multiple. Then, when it is necessary to store the image to be classified, the electronic device can store the image to be classified in the image folders corresponding to each scene category among the multiple scene categories to which the image to be classified belongs.
[0152] Of course, in order to save storage space, the electronic device can also store the image to be classified in any one of the multiple scene categories to which the image to be classified belongs. Or, the electronic device stores an image to be classified and stores the image identifier of the image to be classified and its corresponding scene category. When it is necessary to display the image according to the classification method, the electronic device can obtain the image to be classified according to the image identifier of the image to be classified and display the image to be classified in the classification display interface.
[0153] It should be noted that by storing the preview image after exposure in the image folder corresponding to the scene category to which the preview image belongs, it is convenient for users to search for images according to the scene category, improving the interaction with users and user stickiness.
[0154] As an example, when the image to be classified is a preview image, after the electronic device determines the scene category to which the image to be classified belongs, it can also display a scene label in the preview image, and this scene label is used to describe the scene category to which the image to be classified belongs. Exemplarily, if the scene categories to which the image to be classified belongs include the cute pet scene category, the grassland scene category, and the sky scene category, then the electronic device can display "cute pet", "grassland", and "sky" in the preview image, and this scene can refer to the Figure 2 application scenarios shown above.
[0155] As an example, when the image to be classified is a preview image, after the electronic device determines the scene category to which the image to be classified belongs, it can also determine the imaging scheme of the image to be classified according to the scene category to which the image to be classified belongs, such as determining the exposure parameter, filter scheme, shooting mode, display resolution, etc. of the image to be classified. Exemplarily, this scene can refer to the Figure 3 scenarios shown above.
[0156] Exemplarily, when the scene category to which the image to be classified belongs is the moon scene, the electronic device can reduce the exposure parameter of the image to be classified to obtain a clearer moon image. Or, the electronic device can control the camera to enter the full moon mode (also known as the moon shooting mode) to change the imaging scheme of the image to be classified.
[0157] Exemplarily, when the scene category to which the image to be classified belongs is the text scene category or the scene category with rich textures, the electronic device can perform super-resolution processing on the image to be classified, that is, enlarge the image display resolution of the image to be displayed.
[0158] In the embodiments of the present application, since a single image classification model, the target image classification model, can identify all the scene categories to which the image to be classified belongs, it is possible to obtain the scene categories to which the image to be classified belongs as much as possible at one time, improving the efficiency of the target image classification model for image classification, and the electronic device can adopt different display schemes according to the scene categories to which the image to be classified belongs, improving the display quality and display effect of the image to be classified.
[0159] After explaining the training method of the image classification model provided in the embodiments of the present application in detail, the electronic device involved in the embodiments of the present application will be described.
[0160] As an example, this method can be applied to an electronic device capable of model training. By way of example and not limitation, the electronic device can be, but is not limited to, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, a vehicle-mounted device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a mobile phone, etc., and the embodiments of the present application do not make any limitations in this regard.
[0161] In addition, the electronic device can also apply the trained target image classification model, and the electronic device for training the target image classification model and the electronic device for applying the target image classification model can be the same electronic device or different electronic devices, and the embodiments of the present application do not make specific limitations in this regard.
[0162] Figure 10 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Refer to Figure 10 , the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0163] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0164] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0165] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0166] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0167] In some embodiments, the processor 110 may include one or more interfaces, such as an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0168] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection manners in the above embodiments, or a combination of multiple interface connection manners.
[0169] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor. The GPU is used to execute mathematical and geometric calculations and is used for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0170] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel may adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is an integer greater than 1.
[0171] The electronic device 100 can realize the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0172] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP may be provided in the camera 193.
[0173] The camera 193 is used to capture static images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transfers the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is an integer greater than 1.
[0174] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0175] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0176] The NPU is a neural-network (NN) computing processor. By referring to the structure of a biological neural network, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.
[0177] As an example, the NPU may include the target image classification model provided in the embodiments of the present application.
[0178] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0179] The internal memory 121 can be used to store computer-executable program codes, and the computer-executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.). In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0180] The electronic device 100 can implement audio functions, such as music playback, recording, etc., through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc.
[0181] The pressure sensor 180A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can include at least two parallel plates with conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 194, the electronic device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 180A. In some embodiments, touch operations with the same touch position but different touch operation intensities can correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.
[0182] The gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 180B detects the shaking angle of the electronic device 100, calculates the distance that the lens module needs to compensate according to the angle, and enables the lens to offset the shaking of the electronic device 100 through reverse movement to achieve anti-shake. The gyroscope sensor 180B can also be used for navigation and somatosensory game scenarios.
[0183] The distance sensor 180F is used to measure distance. The electronic device 100 can measure distance through infrared or laser. In some embodiments, in a shooting scenario, the electronic device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.
[0184] The ambient light sensor 180L is used to sense the ambient light brightness. The electronic device 100 can adaptively adjust the brightness of the display screen 194 according to the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance during photography. The ambient light sensor 180L can also cooperate with the proximity light sensor 180G to detect whether the electronic device 100 is in a pocket to prevent accidental touch.
[0185] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint photography, fingerprint answering calls, etc.
[0186] The touch sensor 180K, also known as the "touch panel". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as a "touch screen". The touch sensor 180K is used to detect touch operations acting on or near it. The touch sensor 180K can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, in a different position from the display screen 194.
[0187] Next, the software system of the electronic device 100 will be described.
[0188] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In this application embodiment, the Android system with a layered architecture is taken as an example to exemplarily describe the software system of the electronic device 100.
[0189] Figure 11It is a block diagram of a software system of an electronic device 100 provided by an embodiment of the present application. Refer to Figure 11 , the layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom are the application layer, the application framework layer, the Android runtime and the system layer, and the kernel layer.
[0190] The application layer may include a series of application packages. Such as Figure 11 shown, the application packages may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0191] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. Such as Figure 11 shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc. The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc. The content provider is used to store and obtain data, and make these data accessible to applications. These data may include videos, images, audio, incoming and outgoing calls, browsing history and bookmarks, phone book, etc. The view system includes visible controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build the display interface of an application. The display interface can be composed of one or more views. For example, it includes a view for displaying a short message notification icon, a view for displaying text, and a view for displaying pictures. The phone manager is used to provide the communication function of the electronic device 100, such as the management of call status (including answering, hanging up, etc.). The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc. The notification manager enables applications to display notification information in the status bar, can be used to convey notification-type messages, and can automatically disappear after a short stay without user interaction. For example, the notification manager is used to notify that the download is completed, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background running application. The notification manager can also be a notification that appears on the screen in the form of a dialog window, such as prompting text information in the status bar, emitting a prompt sound, the electronic device vibrating, the indicator light flashing, etc.
[0192] The Android Runtime includes core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system. The core libraries consist of two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android. The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0193] The system libraries can include multiple functional modules, such as: Surface Manager, Media Libraries, 3D graphics processing libraries (such as OpenGL ES), 2D graphics engines (such as SGL), etc. The Surface Manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications. The Media Libraries support the playback and recording of various common audio and video formats, as well as static image files, etc. The Media Libraries can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc. The 2D graphics engine is the drawing engine for 2D drawing.
[0194] The kernel layer is the layer between the hardware and the software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.
[0195] The following takes the capture and photo-taking scenario as an example to exemplarily illustrate the working processes of the software and hardware of the electronic device 100.
[0196] When the touch sensor 180K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the raw input event. Taking the touch operation as a click operation and the control corresponding to the click operation being the control of the camera application icon as an example, the camera application calls the interface of the application framework layer to start the camera application, then calls the kernel layer to start the camera driver, and captures a static image or video through the camera 193.
[0197] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a Digital Versatile Disc (DVD)), or a semiconductor medium (such as a Solid State Disk (SSD)), etc.
[0198] The above are the optional embodiments provided by the present application, which are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the technical scope disclosed in the present application shall be included in the protection scope of the present application.
Claims
1. A training method for an image classification model, characterized in that, applied to an electronic device, the method includes: Obtain a positive sample image training set, the positive sample image training set includes a plurality of sample images and a plurality of first text information, the scene category to which each sample image belongs is at least one of a plurality of scene categories, the plurality of sample images correspond to the plurality of first text information one by one, and one first text information is used to describe the scene category to which the corresponding sample image belongs; According to the positive sample image training set, determine a negative sample image training set, the negative sample image training set includes the plurality of sample images, the text information of each sample image in the negative sample image training set is second text information, and the second text information of any one sample image in the negative sample image training set is one of the plurality of first text information other than the first text information corresponding to the any one sample image; Based on the positive sample image training set and the negative sample image training set, perform iterative training on the initial classification model to obtain a target image classification model, and the target image classification model can identify images belonging to at least one of the plurality of scene categories.
2. The method according to claim 1, characterized in that, The determining the negative sample image training set according to the positive sample image training set includes: Determine the similarity between the target sample image and the first text information corresponding to each target scene category, the target sample image is any one sample image in the positive sample image training set, and each target scene category refers to each scene category other than the scene category to which the target sample image belongs among the plurality of scene categories; Determine the first text information corresponding to the target scene category with the largest similarity as the second text information corresponding to the target sample image; Determine the plurality of sample images and the second text information corresponding to each sample image as the negative sample image training set.
3. The method according to claim 2, characterized in that, The determining the similarity between the target sample image and the first text information corresponding to each target scene category includes: Process the target sample image through a pre-trained target image encoder to obtain a visual feature vector of the target sample image; Process the first text information corresponding to each target scene category through a pre-trained target text encoder to obtain a text feature vector of the first text information corresponding to each target scene category; Determine the similarity between the visual feature vector of the target sample image and the text feature vector corresponding to each target scene category.
4. The method according to claim 2 or 3, characterized in that, Before determining the similarity between the target sample image and the first text information corresponding to each target scene category, it further includes: Obtain a plurality of sample text information and a plurality of sample training images, the plurality of sample text information corresponds to the plurality of sample training images one by one, and each sample text information in the plurality of sample text information is used to describe the scene category to which the corresponding sample training image belongs; Iteratively train the initial text encoder based on the multiple sample text information, and iteratively train the initial image encoder based on the multiple sample training images; During the iterative training process, determine the loss value of the first loss function between the text encoder obtained after each training and the image encoder obtained after each training; When the loss value converges, determine the text encoder obtained at the time of convergence as the target text encoder, and determine the image encoder obtained at the time of convergence as the target image encoder. The target text encoder is used to determine the text feature vector of the first text information corresponding to each target scene category, and the target image encoder is used to determine the visual feature vector of the target sample image.
5. The method according to any one of claims 1-4, wherein, the iterative training of the initial classification model based on the positive sample image training set and the negative sample image training set to obtain a target image classification model includes: Concatenate the visual feature vector of each sample image in the positive sample image training set with the text feature vector of the corresponding first text information to obtain a plurality of positive sample mixed feature vectors; Concatenate the visual feature vector of each sample image in the negative sample image training set with the text feature vector of the corresponding second text information respectively to obtain a plurality of negative sample mixed feature vectors; Iteratively train the initial classification model according to the plurality of positive sample mixed feature vectors and the plurality of negative sample mixed feature vectors; During the iterative training process, determine the loss value of the second loss function between the classification result of the image classification model obtained after each training and the preset result; When the loss value of the second loss function converges, determine the image classification model obtained at the time of convergence as the target image classification model.
6. The method according to any one of claims 1-5, wherein, after the iterative training of the initial classification model based on the positive sample image training set and the negative sample image training set to obtain a target image classification model, further includes: Obtain an image to be classified; Determine the visual feature vector of the image to be classified; Determine the similarity between the visual feature vector of the image to be classified and each first text information among the plurality of first text information to obtain a plurality of similarities; Concatenate the visual feature vector of the image to be classified with the text feature vector corresponding to each first text information among N first text information respectively to obtain N mixed feature vectors. The N first text information are the first text information corresponding to each of the N similarities ranked from largest to smallest among the plurality of similarities, and N is a positive integer greater than or equal to 1; Process the N mixed feature vectors through the target image classification model to obtain the classification result of the image to be classified.
7. The method according to claim 6, wherein, the obtaining of the image to be classified includes: When the camera is turned on, determine the preview image collected by the camera as the image to be classified; or, In the case of receiving an image selection operation, determine that the image selected by the image selection operation is the image to be classified.
8. The method according to claim 7, wherein, the determining, when the camera is turned on, that the preview image captured by the camera is the image to be classified includes: when the camera is turned on, display a scene recognition control in the shooting interface, and perform image capture through the camera to obtain the preview image, where the scene recognition control is used to control whether to perform scene recognition; in response to an enabling operation on the scene recognition control, determine that the preview image is the image to be classified.
9. An electronic device, wherein, the structure of the electronic device includes a processor and a memory; the memory is used to store a program that supports the electronic device to execute the method according to any one of claims 1-8.
10. A computer-readable storage medium, wherein, instructions are stored in the computer-readable storage medium, and when they run on a computer, cause the computer to execute the method according to any one of claims 1-8.
Citation Information
Patent Citations
Sample set acquisition method and device, computer equipment and storage medium
CN112819023A
Image recognition method and device, computer equipment and storage medium
CN113705596A
Character recognition model training method and device
CN113947773A
Text recognition system training method in self-supervised contrast learning natural scene
CN114973226A
Image processing method and processor
CN115019218A