Image auditing method and device, electronic equipment and storage medium
By calculating object recognition and area ratio of the image to be reviewed, it is automatically determined that the image audit has been passed or failed, which solves the problems of resource waste and accuracy reduction caused by manual audit in the prior art, and achieves efficient and accurate image audit.
Patent Information
- Application Number
- CN202510364294.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-27
AI Technical Summary
Existing image auditing solutions rely on manual audits, resulting in waste of resources and reduced accuracy, and are prone to errors in judgment due to subjective factors.
By identifying the object of the image to be reviewed, determining the candidate object and its area, and calculating the area ratio of the target object area. If the ratio is greater than the preset threshold and there are no violations, it is determined that the image review is passed.
It has realized the automation of image audits, reduce human resource waste, improve audit accuracy, and avoid the influence of subjective factors.
Smart Images

Figure CN120220126A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent analysis technology, and particularly to an image review method, apparatus, electronic device, and storage medium. Background Art
[0002] In the context of current digital media and social networks, in order to meet regulatory requirements, images uploaded by users often need to be reviewed to ensure that their content complies with the regulations. However, most existing image review solutions rely on manual review, which consumes a large amount of human resources. And it can be understood that a category often contains multiple different objects, and some of the objects included in the category are compliant non-violating objects, while some are non-compliant violating objects. In the manual review method, it is usually the subjective judgment of the reviewer whether the objects belonging to the category in the image comply with the regulations. Then, due to the influence of the subjective factors of the reviewer, in the process of manual review, there may be cases of incorrect judgment, resulting in a decrease in the accuracy of image review. Summary of the Invention
[0003] The purpose of the embodiments of the present invention is to provide an image review method, apparatus, electronic device, and storage medium to improve the accuracy of image review. The specific technical solutions are as follows:
[0004] The embodiments of the present invention provide an image review method, the method comprising:
[0005] Obtain an image to be reviewed;
[0006] Perform object recognition on the image to be reviewed to obtain all candidate objects included in the image to be reviewed and the regions where each of the candidate objects is located;
[0007] Determine all objects of a first preset category from each of the candidate objects as target objects; wherein, the first preset category is a category including objects that are all non-violating objects;
[0008] Calculate the ratio of the total area of the regions where all the target objects are located to the area of the image to be reviewed;
[0009] If the ratio is greater than a preset threshold and there is no object of a second preset category among each of the candidate objects, it is determined that the image to be reviewed passes the review; wherein, the second preset category is a category including objects that are all violating objects.
[0010] In a possible embodiment, the method further comprises:
[0011] For each first preset category, obtain an image including the objects of the first preset category as the target image corresponding to the first preset category;
[0012] For each of the first preset categories, extract the image features of the target image corresponding to the first preset category to obtain the target image features corresponding to the first preset category;
[0013] For each of the first preset categories, extract the text features of the text input for the first preset category to obtain the target text features corresponding to the first preset category;
[0014] The step of determining all objects with the category of the first preset category from each of the candidate objects as the target objects includes:
[0015] For each of the candidate objects, extract the image features of the candidate object from the image to be audited as the first image features corresponding to the candidate object;
[0016] From each of the candidate objects, determine all objects whose corresponding first image features match the first preset category features corresponding to any of the first preset categories as the target objects; wherein, the first preset category features corresponding to the first preset category include: the target image features corresponding to the first preset category, and the target text features corresponding to the first preset category.
[0017] In a possible embodiment, the method further includes:
[0018] For each of the second preset categories, obtain the image of the object including the second preset category as the target image corresponding to the second preset category;
[0019] For each of the second preset categories, extract the image features of the target image corresponding to the second preset category to obtain the target image features corresponding to the second preset category;
[0020] For each of the second preset categories, extract the text features of the text input for the second preset category to obtain the target text features corresponding to the second preset category;
[0021] Determine whether there are objects with the category of the second preset category among each of the candidate objects in the following manner:
[0022] For each of the candidate objects, extract the image features of the candidate object from the image to be audited as the first image features corresponding to the candidate object;
[0023] If there is an object among the candidate objects whose corresponding first image feature matches the second preset category feature corresponding to any one of the second preset categories, then there is an object of the second preset category among the candidate objects; wherein, the second preset category feature corresponding to the second preset category includes: the target image feature corresponding to the second preset category and the target text feature corresponding to the second preset category;
[0024] If there is no object among the candidate objects whose corresponding first image feature matches the second preset category feature corresponding to any one of the second preset categories, then there is no object of the second preset category among the candidate objects.
[0025] In a possible embodiment, the number of target images corresponding to the preset category is multiple;
[0026] The target image feature corresponding to the preset category is obtained by the following method, including:
[0027] Extract the image features of each target image corresponding to the preset category respectively as the second image features corresponding to each of the target images;
[0028] Based on the second image features corresponding to each of the target images, obtain the target image feature corresponding to the preset category; wherein, the similarity between the target image feature corresponding to the preset category and the second image features corresponding to each of the target images is greater than the first preset similarity threshold;
[0029] The target text feature corresponding to the preset category is obtained by the following method, including:
[0030] Obtain the target text with the same text semantics as the text input for the preset category;
[0031] Extract the first text feature of the text input for the preset category, and extract the text features of each of the target texts respectively as the second text features corresponding to each of the target texts;
[0032] Based on the first text feature and the second text features corresponding to each of the target texts, obtain the target text feature corresponding to the preset category; wherein, the similarity between the target text feature corresponding to the preset category and the second text features corresponding to each of the target texts is greater than the second preset similarity threshold, and the similarity with the first text feature is greater than the second preset similarity threshold.
[0033] In a possible embodiment, the obtaining the target image feature corresponding to the preset category based on the second image features corresponding to each of the target images includes:
[0034] Calculate the mean of the second image features corresponding to each of the target images as the target image feature corresponding to the preset category;
[0035] The obtaining of the target text feature corresponding to the preset category according to the first text feature and the second text features corresponding to the respective target texts includes:
[0036] Calculate the mean of the first text feature and the second text features corresponding to the respective target texts as the target text feature corresponding to the preset category.
[0037] In a possible embodiment, the obtaining of the target image feature corresponding to the preset category according to the second image features corresponding to the respective target images includes:
[0038] Cluster the second image features corresponding to the respective target images to obtain a first cluster center feature as the target image feature corresponding to the preset category;
[0039] The obtaining of the target text feature corresponding to the preset category according to the first text feature and the second text features corresponding to the respective target texts includes:
[0040] Cluster the first text feature and the second text features corresponding to the respective target texts to obtain a second cluster center feature as the target text feature corresponding to the preset category.
[0041] In a possible embodiment, the method further includes:
[0042] If the ratio is not greater than a preset threshold, or there is an object with a second preset category among the candidate objects, send the image to be reviewed to the manual end for review.
[0043] In a possible embodiment, the performing of object recognition on the image to be reviewed to obtain all candidate objects included in the image to be reviewed and the regions where the respective candidate objects are located includes:
[0044] For each pixel point in the image to be reviewed, identify the category label and instance label of the pixel point;
[0045] For multiple pixel points with the same category label and the same instance label, determine the object formed by the multiple pixel points as a candidate object, and determine the region formed by the multiple pixel points as the region where the candidate object is located.
[0046] An embodiment of the present invention further provides an image review device, and the device includes:
[0047] An image acquisition module, configured to acquire an image to be reviewed;
[0048] A candidate object recognition module, configured to perform object recognition on the image to be reviewed to obtain all candidate objects included in the image to be reviewed and the regions where each of the candidate objects is located;
[0049] A target object determination module, configured to determine all objects of a first preset category from each of the candidate objects as target objects; wherein, the first preset category is a category that includes objects that are all non-violating objects;
[0050] A ratio calculation module, configured to calculate the ratio of the total area of the regions where all the target objects are located to the area of the image to be reviewed;
[0051] An audit determination module, configured to determine that the image to be reviewed passes the audit if the ratio is greater than a preset threshold and there are no objects of a second preset category among each of the candidate objects; wherein, the second preset category is a category that includes objects that are all violating objects.
[0052] In a possible embodiment, the apparatus further includes:
[0053] A first target image acquisition module, configured to, for each first preset category, acquire an image including the objects of the first preset category as the target image corresponding to the first preset category;
[0054] A first image feature extraction module, configured to, for each first preset category, extract the image features of the target image corresponding to the first preset category to obtain the target image features corresponding to the first preset category;
[0055] A first text feature extraction module, configured to, for each first preset category, extract the text features of the text input for the first preset category to obtain the target text features corresponding to the first preset category;
[0056] The determining all objects of a first preset category from each of the candidate objects as target objects includes:
[0057] For each of the candidate objects, extracting the image features of the candidate object from the image to be reviewed as the first image features corresponding to the candidate object;
[0058] Determining, from each of the candidate objects, all objects whose corresponding first image features match the first preset category features corresponding to any one of the first preset categories as target objects; wherein, the first preset category features corresponding to the first preset category include: the target image features corresponding to the first preset category and the target text features corresponding to the first preset category.
[0059] In a possible embodiment, the apparatus further comprises:
[0060] A second target image acquisition module, configured to acquire, for each second preset category, an image including an object of the second preset category as a target image corresponding to the second preset category;
[0061] A second image feature extraction module, configured to extract, for each second preset category, an image feature of the target image corresponding to the second preset category to obtain a target image feature corresponding to the second preset category;
[0062] A second text feature extraction module, configured to extract, for each second preset category, a text feature of the text input for the second preset category to obtain a target text feature corresponding to the second preset category;
[0063] Determine whether there is an object of the second preset category among the candidate objects in the following manner:
[0064] For each candidate object, extract an image feature of the candidate object from the image to be audited as a first image feature corresponding to the candidate object;
[0065] If there is an object among the candidate objects whose corresponding first image feature matches a second preset category feature corresponding to any second preset category, then there is an object of the second preset category among the candidate objects; wherein, the second preset category feature corresponding to the second preset category includes: the target image feature corresponding to the second preset category and the target text feature corresponding to the second preset category;
[0066] If there is no object among the candidate objects whose corresponding first image feature matches a second preset category feature corresponding to any second preset category, then there is no object of the second preset category among the candidate objects.
[0067] In a possible embodiment, the number of target images corresponding to the preset category is multiple;
[0068] The method for obtaining the target image feature corresponding to the preset category includes:
[0069] Extract the image features of each target image corresponding to the preset category respectively as second image features corresponding to each target image;
[0070] Obtain the target image features corresponding to the preset category according to the second image features corresponding to each of the target images; wherein, the similarity between the target image features corresponding to the preset category and the second image features corresponding to each of the target images is greater than the first preset similarity threshold;
[0071] Obtain the target text features corresponding to the preset category in the following manner, including:
[0072] Obtain a target text with the same text semantics as the text input for the preset category;
[0073] Extract the first text features of the text input for the preset category, and respectively extract the text features of each of the target texts as the second text features corresponding to each of the target texts;
[0074] Obtain the target text features corresponding to the preset category according to the first text features and the second text features corresponding to each of the target texts; wherein, the similarity between the target text features corresponding to the preset category and the second text features corresponding to each of the target texts is greater than the second preset similarity threshold, and the similarity with the first text features is greater than the second preset similarity threshold.
[0075] In a possible embodiment, the obtaining the target image features corresponding to the preset category according to the second image features corresponding to each of the target images includes:
[0076] Calculate the mean value of the second image features corresponding to each of the target images as the target image features corresponding to the preset category;
[0077] The obtaining the target text features corresponding to the preset category according to the first text features and the second text features corresponding to each of the target texts includes:
[0078] Calculate the mean value of the first text features and the second text features corresponding to each of the target texts as the target text features corresponding to the preset category.
[0079] In a possible embodiment, the obtaining the target image features corresponding to the preset category according to the second image features corresponding to each of the target images includes:
[0080] Cluster the second image features corresponding to each of the target images to obtain the first cluster center feature as the target image features corresponding to the preset category;
[0081] The obtaining the target text features corresponding to the preset category according to the first text features and the second text features corresponding to each of the target texts includes:
[0082] Cluster the first text feature and the second text features corresponding to each of the target texts to obtain a second cluster center feature as the target text feature corresponding to the preset category.
[0083] In a possible embodiment, the apparatus further includes:
[0084] An artificial review module, configured to send the image to be reviewed to the artificial side for review if the ratio is not greater than a preset threshold, or if there is an object with a second preset category among the candidate objects.
[0085] In a possible embodiment, the object recognition of the image to be reviewed to obtain all candidate objects included in the image to be reviewed and the regions where the candidate objects are located respectively includes:
[0086] For each pixel point in the image to be reviewed, identify the category label and instance label of the pixel point;
[0087] For multiple pixel points with the same category label and the same instance label, determine the object formed by the multiple pixel points as a candidate object, and determine the region formed by the multiple pixel points as the region where the candidate object is located.
[0088] An embodiment of the present invention further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0089] The memory is used to store a computer program;
[0090] The processor is configured to implement any of the above image review methods when executing the program stored on the memory.
[0091] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored, and the computer program implements any of the above image review methods when executed by a processor.
[0092] An image review method, device, electronic device, and storage medium provided by an embodiment of the present invention obtain all candidate objects included in the image to be reviewed and the regions where each candidate object is located by performing object recognition on the acquired image to be reviewed. Since the second preset category includes objects that are all illegal objects, and each candidate object is all the objects included in the image to be reviewed, if there is an object with the second preset category among the candidate objects, it means that the image to be reviewed includes illegal objects, that is, it is considered that the image to be reviewed does not meet the regulations; if there is no object with the second preset category among the candidate objects, it means that the image to be reviewed does not include illegal objects. Since there is a situation where a certain category other than the second preset category includes both illegal objects and non-illegal objects, even if there is no object with the second preset category among the candidate objects, the image to be reviewed may still include illegal objects. And since the first preset category includes objects that are all non-illegal objects, and each candidate object is all the objects included in the image to be reviewed, the non-illegal objects included in the image to be reviewed, that is, all target objects, can be determined by determining all objects with the first preset category from the candidate objects. If the ratio of the total area of the regions where all target objects are located to the area of the image to be reviewed is greater than a preset threshold, it means that the objects in most regions of the image to be reviewed belong to the first preset category, that is, the objects in most regions of the image to be reviewed are all non-illegal objects. Therefore, if the ratio is greater than the preset threshold and there is no object with the second preset category among the candidate objects, it is considered that the objects in most regions of the image to be reviewed are all non-illegal objects, and the image to be reviewed does not include illegal objects in the second preset category. At this time, it can be approximately considered that the image to be reviewed only includes non-illegal objects that meet the regulations, that is, it is considered that the image to be reviewed passes the review. It can be seen that the embodiment of the present invention can achieve automatic review of the image to be reviewed, reduce waste of human resources, avoid reduction of image review efficiency caused by manual review methods, and avoid reduction of image review accuracy caused by subjective factors of reviewers, thereby improving the accuracy of image review. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art.
[0094] Figure 1 It is a schematic flowchart of an image review method provided by an embodiment of the present invention;
[0095] Figure 2 It is a schematic flowchart of a method for obtaining the first preset category features provided by an embodiment of the present invention;
[0096] Figure 3 It is a schematic flowchart of a method for determining a target object provided by an embodiment of the present invention;
[0097] Figure 4 It is a schematic flowchart of a method for obtaining a second preset category feature provided by an embodiment of the present invention;
[0098] Figure 5 It is another schematic flowchart of an image review method provided by an embodiment of the present invention;
[0099] Figure 6a It is a schematic flowchart of a method for obtaining target image features provided by an embodiment of the present invention;
[0100] Figure 6b It is a schematic flowchart of a method for obtaining target text features provided by an embodiment of the present invention;
[0101] Figure 7 It is yet another schematic flowchart of an image review method provided by an embodiment of the present invention;
[0102] Figure 8 It is still another schematic flowchart of an image review method provided by an embodiment of the present invention;
[0103] Figure 9 It is a schematic structural diagram of an image review device provided by an embodiment of the present invention;
[0104] Figure 10 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0105] Next, the technical solutions in the embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention.
[0106] To more clearly illustrate the image review method provided by this application, the possible application scenarios of the image review method provided by this application will be exemplarily described below using the scenario of social avatar review as an example. It can be understood that the scenario of social avatar review in the following examples is only a possible application scenario of the image review method provided by this application. In other possible embodiments, the image review method provided by this application can be applied to other possible application scenarios, and the following examples do not impose any restrictions on this.
[0107] In the context of digital media and social networks, to meet regulatory requirements, social avatars (hereinafter referred to as images) uploaded by users often need to be reviewed to ensure that their content complies with the regulations. Most existing image review solutions rely on manual review, which consumes a large amount of human resources. Moreover, the number of images to be reviewed is large, and the manual review method takes a long time to complete the review of a large number of images, resulting in low efficiency of image review.
[0108] It can be understood that when conducting image review, a category often contains multiple different objects. Some of the objects included in this category comply with the regulations, while some do not. In the manual review method, it is usually the subjective judgment of the reviewer whether the objects belonging to this category in the image comply with the regulations. Then, due to the influence of the subjective factors of the reviewer, in the process of manual review, there may be cases of misjudgment, resulting in a decrease in the accuracy of image review.
[0109] Based on this, in order to avoid the reduction in image review efficiency caused by the manual review method and improve the accuracy of image review, the present invention provides an image review method, as Figure 1 shown, the method includes:
[0110] S101, obtaining the image to be reviewed.
[0111] S102, performing object recognition on the image to be reviewed to obtain all candidate objects included in the image to be reviewed and the regions where each candidate object is located.
[0112] S103, determining all objects with the first preset category from each candidate object as target objects.
[0113] Wherein, the first preset category is a category including objects that all belong to non-violating objects.
[0114] S104, calculating the ratio of the total area of the regions where all target objects are located to the area of the image to be reviewed.
[0115] S105, if the ratio is greater than the preset threshold and there is no object with the second preset category among each candidate object, it is determined that the image to be reviewed passes the review.
[0116] Wherein, the second preset category is a category including objects that all belong to violating objects.
[0117] Applying the embodiments of the present invention, by identifying objects in the to-be-reviewed image obtained, all candidate objects included in the to-be-reviewed image and the regions where each candidate object is located are obtained. Since the second preset category is a category that includes objects all belonging to the category of illegal objects, and each candidate object is all objects included in the to-be-reviewed image, therefore, if there is an object with the category of the second preset category among the candidate objects, it indicates that the to-be-reviewed image includes illegal objects, that is, it is considered that the to-be-reviewed image does not meet the regulations. If there is no object with the category of the second preset category among the candidate objects, it indicates that the to-be-reviewed image does not include illegal objects. Since there is a situation where a certain category other than the second preset category includes both illegal objects and non-illegal objects, therefore, even if there is no object with the category of the second preset category among the candidate objects, the to-be-reviewed image may still include illegal objects. And since the first preset category is a category that includes objects all belonging to the category of non-illegal objects, and each candidate object is all objects included in the to-be-reviewed image, therefore, the non-illegal objects included in the to-be-reviewed image can be determined by determining all objects with the category of the first preset category from each candidate object, that is, all target objects are determined. If the ratio of the total area of the regions where all target objects are located to the area of the to-be-reviewed image is greater than a preset threshold, it indicates that the objects in most regions of the to-be-reviewed image belong to the first preset category, that is, the objects in most regions of the to-be-reviewed image are all non-illegal objects. Therefore, if the ratio is greater than the preset threshold and there is no object with the category of the second preset category among the candidate objects, it is considered that the objects in most regions of the to-be-reviewed image are all non-illegal objects, and the to-be-reviewed image does not include illegal objects in the second preset category. At this time, it can be approximately considered that the to-be-reviewed image only includes non-illegal objects that meet the regulations, that is, it is considered that the to-be-reviewed image passes the review. It can be seen that the embodiments of the present invention can realize the automatic review of the to-be-reviewed image, reduce the waste of human resources, avoid the reduction of the image review efficiency caused by the manual review method, and avoid the reduction of the image review accuracy caused by the subjective factors of the reviewers, thereby improving the accuracy of the image review.
[0118] The foregoing S101 - S105 will be exemplarily described below:
[0119] In S101, the execution subject can obtain the to-be-reviewed image by the user inputting an image on the execution subject and using the image input by the user as the to-be-reviewed image. The execution subject refers to any device that can be used to execute the image review method provided by the present invention. For example, a mobile phone, a tablet computer, a laptop computer, etc.
[0120] The executing entity can also obtain the image to be reviewed by receiving the images sent by other devices and using the images sent by other devices as the images to be reviewed. Specifically, in the scenario of social avatar review, the user uploads an avatar for a certain social account on other devices, and the other devices send the avatar uploaded by the user to the executing entity, so that the executing entity can perform image review on the avatar uploaded by the user to determine whether the content in the avatar uploaded by the user complies with the regulations. The avatar uploaded by the user received by the executing entity is the image to be reviewed, and the acquisition of the image to be reviewed can be achieved through the above process. Other devices refer to any device that can implement the above process, such as mobile phones, tablets, laptops, etc.
[0121] In S102, the number of candidate objects included in the image to be reviewed can be multiple or one. For example, if the image to be reviewed is a solid-color background image, the candidate object in the image to be reviewed is one; if the image to be reviewed is a group photo of multiple people, the candidate objects in the image to be reviewed are multiple.
[0122] In one possible embodiment, object recognition can be performed on the image to be reviewed by means of object detection to obtain all the candidate objects included in the image to be reviewed and the regions where each candidate object is located.
[0123] In another possible embodiment, object recognition can be performed on the image to be reviewed by means of segmentation to obtain all the candidate objects included in the image to be reviewed and the regions where each candidate object is located.
[0124] In this embodiment, specifically, instance segmentation can be used to distinguish and accurately locate all the candidate objects included in the image to be reviewed, and based on the accurate location of each candidate object, all the candidate objects included in the image to be reviewed are segmented to obtain all the candidate objects included in the image to be reviewed and the regions where each candidate object is located.
[0125] It is also possible to identify the class labels of each pixel point in the image to be reviewed by means of semantic segmentation, determine multiple pixel points with the same class label and adjacent positions, use the object formed by the multiple pixel points as the candidate object, and use the region formed by the multiple pixel points as the region where the candidate object is located.
[0126] It can be understood that the object detection method can only reflect the approximate area where the detected candidate objects are located in the image to be reviewed through the detection frame, while the instance segmentation or semantic segmentation method can segment different candidate objects in the image to be reviewed to obtain the specific areas where each candidate object is located in the image to be reviewed. Therefore, compared with the object detection method, the instance segmentation or semantic segmentation method can more accurately obtain the areas where each candidate object is located, more accurately determine the proportion of each candidate object in the image to be reviewed, and improve the accuracy of image review.
[0127] In S103, the non-violating objects are set according to user requirements. Suppose all the objects included in a certain category are non-violating objects, then this category can be preset as the first preset category, and the first preset category can be one category or multiple categories. Exemplarily, suppose all plants such as pine trees, banyan trees, etc. are non-violating objects, then plants can be preset as the first preset category; or, suppose all plants and all animals such as cats, dogs, etc. are non-violating objects, then plants and animals can be preset as the first preset category.
[0128] In a possible embodiment, for each candidate object, the category of the candidate object can be determined to be the first preset category by respectively performing feature matching between the features of the candidate object and the first preset category features corresponding to each first preset category. When the category of the candidate object is the first preset category, the candidate object is used as the target object. Exemplarily, suppose there are three first preset categories in total, denoted as the first preset category 1, the first preset category 2, and the first preset category 3 respectively. For the convenience of description, the features of the candidate object are denoted as feature 1, the first preset category feature corresponding to the first preset category 1 is denoted as feature 2, the first preset category feature corresponding to the first preset category 2 is denoted as feature 3, and the first preset category feature corresponding to the first preset category 3 is denoted as feature 4. Then, feature 1 can be matched with feature 2, feature 1 can be matched with feature 3, and feature 1 can be matched with feature 4. If feature 1 and feature 2 are matched, it is determined that the category of the candidate object is the first preset category 1. The acquisition method of the first preset category features corresponding to each first preset category and the feature matching method will be exemplarily described below and will not be elaborated here.
[0129] In another possible embodiment, when performing object recognition on the image to be reviewed, the categories of each candidate object can be obtained. For example, when all the candidate objects included in the image to be reviewed and the areas where each candidate object is located are obtained by semantic segmentation, multiple pixel points with the same category label and adjacent positions are determined, the pixel area formed by the multiple pixel points is regarded as the area where the candidate object is located, and the category label of the multiple pixel points is used as the category of the candidate object.
[0130] In this embodiment, for each candidate object, it is possible to determine whether the category of the candidate object is a first preset category by comparing the category of the candidate object with each first preset category. When the category of the candidate object is the first preset category, the candidate object is used as the target object.
[0131] In S104, the area of a certain region can be expressed in the form of the number of pixel points included in the region or the physical area occupied by the region, etc. Exemplarily, assume that the area of a certain region is expressed by the number of pixel points included in the region. The number of pixel points included in the image to be reviewed is 100 pixels × 100 pixels = 10,000 pixels. The number of pixel points included in the region where target object 1 is located is 4,000 pixels, the number of pixel points included in the region where target object 2 is located is 550 pixels, and the number of pixel points included in the region where target object 3 is located is 4,050 pixels. Then the total number of pixel points included in the regions where all target objects are located respectively is 4,000 + 550 + 4,050 = 8,600 pixels. That is, the total area of the regions where all target objects are located respectively is 8,600, and the area of the image to be reviewed is 10,000. Then the ratio of the total area of the regions where all target objects are located respectively to the area of the image to be reviewed is 8,600 / 10,000 = 0.86.
[0132] Or, assume that the area of a certain region is expressed by the physical area occupied by the region. The physical area of the image to be reviewed is 100 cm 2 , the physical area occupied by the region where target object 1 is located is 40 cm 2 , the physical area occupied by the region where target object 2 is located is 10 cm 2 , the physical area occupied by the region where target object 3 is located is 15 cm 2 , then the total physical area occupied by the regions where all target objects are located respectively is (40 + 10 + 15) = 65 cm 2 . That is, the total area of the regions where all target objects are located respectively is 65 cm 2 , then the ratio of the total area of the regions where all target objects are located respectively to the area of the image to be reviewed is 65 / 100 = 0.65.
[0133] In S105, the non-compliant objects are set according to user requirements. Suppose all the objects included in a certain category are non-compliant objects, then the category can be preset as the second preset category, and the second preset category can be one category or multiple categories. Exemplarily, suppose all controlled knives such as daggers, switchblades, etc. are non-compliant objects, then the second preset category is controlled knives; or, suppose all controlled knives and all firearms such as rifles, pistols, etc. are non-compliant objects, then the second preset category is controlled knives and firearms.
[0134] In a possible embodiment, for each candidate object, the category of the candidate object can be determined to be the second preset category by respectively performing feature matching between the features of the candidate object and the second preset category features corresponding to each second preset category. Exemplarily, suppose there are three second preset categories in total, denoted as the second preset category 1, the second preset category 2, and the second preset category 3 respectively. For the convenience of description, the features of the candidate object are denoted as feature 5, the second preset category features corresponding to the second preset category 1 are denoted as feature 6, the second preset category features corresponding to the second preset category 2 are denoted as feature 7, and the second preset category features corresponding to the second preset category 3 are denoted as feature 8. Then, feature 5 can be matched with feature 6, feature 5 can be matched with feature 7, and feature 5 can be matched with feature 8. If feature 5 and feature 6 match, it is determined that the category of the candidate object is the second preset category 1.
[0135] In the case where the categories of all candidate objects are not the second preset category, it can be considered that there are no objects in the candidate objects whose category is the second preset category. The acquisition method of the second preset category features corresponding to each second preset category and the method of feature matching will be exemplarily described below, and will not be elaborated here.
[0136] In another possible embodiment, when performing object recognition on the image to be reviewed, the categories of each candidate object can be obtained. For example, when all candidate objects included in the image to be reviewed and the regions where each candidate object is located are obtained by means of semantic segmentation, multiple pixel points with the same category label and adjacent positions are determined, and the object composed of the multiple pixel points is used as the candidate object, and the category label of the multiple pixel points is used as the category of the candidate object.
[0137] In this embodiment, for each candidate object, it is possible to determine whether the category of the candidate object is a second preset category by comparing the category of the candidate object with each second preset category. If the category of the candidate object is the same as a certain second preset category, it is considered that the category of the candidate object is the second preset category; if the category of the candidate object is different from all second preset categories, it is considered that the category of the candidate object is not the second preset category. In the case where the categories of all candidate objects are not the second preset category, it can be considered that there is no object in the candidate objects whose category is the second preset category.
[0138] If there is no object in the candidate objects whose category is the second preset category, it indicates that the to-be-reviewed image does not include the illegal objects in the second preset category. And if the ratio of the total area of the regions where all target objects are located to the area of the to-be-reviewed image is greater than the preset threshold, it indicates that the objects in most regions of the to-be-reviewed image belong to the first preset category, that is, most of the objects in the to-be-reviewed image are non-illegal objects.
[0139] Therefore, if the ratio is greater than the preset threshold and there is no object in the candidate objects whose category is the second preset category, it is considered that most of the objects in the to-be-reviewed image are non-illegal objects. At this time, it can be approximately considered that the to-be-reviewed image only includes non-illegal objects that meet the regulations, that is, it is considered that the to-be-reviewed image passes the review.
[0140] If there is an object in the candidate objects whose category is the second preset category, regardless of whether the ratio is greater than the threshold, it indicates that the to-be-reviewed image includes the illegal objects in the second preset category, that is, the to-be-reviewed image must be illegal. Therefore, in the case where there is an object in the candidate objects whose category is the second preset category, it can be directly determined that the to-be-reviewed image fails the review.
[0141] If the ratio is not greater than the threshold, it is considered that the proportion of non-illegal objects in the to-be-reviewed image is small. At this time, regardless of whether there is an object in the candidate objects whose category is the second preset category, it is impossible to approximately consider that the to-be-reviewed image only includes non-illegal objects that meet the regulations. Therefore, in the case where the ratio is not greater than the threshold, it can be directly determined that the to-be-reviewed image fails the review.
[0142] That is to say, in the case where the ratio is not greater than the threshold, or there is an object in the candidate objects whose category is the second preset category, it can be directly determined that the to-be-reviewed image fails the review. In another possible embodiment, in the case where the ratio is not greater than the threshold, or there is an object in the candidate objects whose category is the second preset category, the to-be-reviewed image can also be sent to the manual end for review, and the manual end further determines whether the to-be-reviewed image meets the regulations to determine whether the to-be-reviewed image passes the review.
[0143] The preset threshold can be set according to user requirements or past experience. It can be understood that the higher the preset threshold is set, the larger the area ratio of the objects of the first preset category in the image to be reviewed when the ratio is greater than the preset threshold and there are no objects of the second preset category, that is, the larger the area ratio of the objects that are definitely not in violation in the image to be reviewed. Correspondingly, the area ratio of the objects that are neither of the first preset category nor of the second preset category in the image to be reviewed is smaller.
[0144] Moreover, it can be understood that if the category of an object is neither of the first preset category nor of the second preset category, then the object may or may not be in violation. Exemplarily, in the category of cartoon images, due to copyright reasons, a certain cartoon image cannot be used. Then, it can be considered that this cartoon image is an object in violation, and other cartoon images are objects not in violation. Thus, the category of cartoon images is neither the first preset category nor the second preset category. In this example, if the category of an object is a cartoon image, then the object may be the above-mentioned object in violation or may be an object not in violation.
[0145] Therefore, if the area ratio of the objects that are neither of the first preset category nor of the second preset category in the image to be reviewed is smaller, then the area ratio of the objects whose compliance cannot be determined in the image to be reviewed is smaller, and it can be more considered that the image to be reviewed does not include objects in violation. At this time, if it is determined that the image to be reviewed passes the review, the accuracy of image review is higher. That is, it can be considered that the higher the preset threshold is set, the higher the accuracy of image review. Based on this, if the user hopes that the accuracy of image review is higher, the preset threshold should be set higher. Exemplarily, the preset threshold can be set to 0.8. In other possible embodiments, the preset threshold can also be set to 0.85, 0.9, etc.
[0146] The foregoing S101 - S105 have been exemplarily described above. As described in the foregoing S103, the target object can be determined by feature matching. The feature matching method depends on the first preset category features corresponding to each first preset category obtained. Therefore, the method for obtaining the first preset category features corresponding to each first preset category will be exemplarily described below. Refer to Figure 2 The first preset category features corresponding to the first preset category include: the target image features corresponding to the first preset category and the target text features corresponding to the first preset category. The method for obtaining the first preset category features corresponding to the first preset category includes:
[0147] S201, for each first preset category, obtain an image including the objects of the first preset category as the target image corresponding to the first preset category.
[0148] The methods for acquiring the first preset category features corresponding to different first preset categories are the same. Therefore, the following only takes one first preset category as an example to illustrate the method for acquiring the first preset category features corresponding to the first preset category.
[0149] If an image includes an object of the first preset category, the image is a target image corresponding to the first preset category, and the target images corresponding to the first preset category may be one or more. Exemplarily, assuming that the first preset category is plants, if image 1 includes pine trees, then image 1 is a target image corresponding to the first preset category. Assuming that the first preset category is animals, if image 2 includes a cat and image 3 includes a dog, then both image 2 and image 3 are target images corresponding to the first preset category.
[0150] S202: For each first preset category, extract image features of a target image corresponding to the first preset category to obtain target image features corresponding to the first preset category.
[0151] The target image may be subjected to image feature extraction by any image feature extraction method to obtain target image features corresponding to the first preset category. For example, the image feature extraction method may be corner point detection, an image feature extraction model based on deep learning, a wavelet transform method, etc. If there are multiple target images corresponding to the first preset category, the image features corresponding to different target images are extracted by the same image feature extraction method.
[0152] Specifically, the target image corresponding to the first preset category can be input into the corner detection algorithm to obtain the image features of the target image corresponding to the first preset category output by the corner detection algorithm, that is, the target image features corresponding to the first preset category are obtained. Alternatively, the target image corresponding to the first preset category can be input into an image feature extraction model based on deep learning to obtain the image features of the target image corresponding to the first preset category output by the feature extraction model based on deep learning, that is, the target image features corresponding to the first preset category are obtained.
[0153] S203: For each first preset category, extract text features of text input for the first preset category to obtain target text features corresponding to the first preset category.
[0154] The text entered for the first preset category is the category name of the first preset category. For example, assuming that the first preset category is plants, the text entered for the first preset category should be "plants" or text with the same or similar semantics as "plants", such as "plants (植物)"; assuming that the first preset category is animals, the text entered for the first preset category should be "animals" or text with the same or similar semantics as "animals".
[0155] The text features of the text input for the first preset category can be extracted by any text feature extraction method to obtain the target text features corresponding to the first preset category. For example, the text feature extraction method can be WordEmbedding (word vector model), a text feature extraction model based on deep learning, etc.
[0156] Specifically, the text input for the first preset category can be input into Word Embedding to obtain the text features output by Word Embedding, that is, the target text features corresponding to the first preset category are obtained. Alternatively, the text input for the first preset category can be input into a text feature extraction model based on deep learning to obtain the text features output by the text feature extraction model based on deep learning, that is, the target text features corresponding to the first preset category are obtained.
[0157] The first preset category features corresponding to the first preset category are the target image features corresponding to the first preset category obtained in S202 and the target text features corresponding to the first preset category obtained in S203. The target image features corresponding to the first preset category and the target text features corresponding to the first preset category can be feature-fused by any feature fusion method to obtain the first preset category features corresponding to the first preset category. For example, the feature fusion method can be cross-modal fusion, splicing, etc.
[0158] Based on Figure 2 the acquisition method of the first preset category features corresponding to the first preset category shown in the embodiment, all objects with the category of the first preset category can be determined from each candidate object as target objects through the following feature matching method. See Figure 3 , the method includes:
[0159] S1031. For each candidate object, extract the image features of the candidate object from the image to be audited as the first image features corresponding to the candidate object.
[0160] For each candidate object, the image features of the candidate object can be extracted by any image feature extraction method as the first image features corresponding to the candidate object. The image feature extraction methods include corner detection, an image feature extraction model based on deep learning, wavelet transform method, etc. The first image features corresponding to different candidate objects are extracted by the same image feature extraction method. Moreover, the first image features corresponding to each candidate object and the target image features corresponding to the first preset category are also extracted by the same image feature extraction method.
[0161] Specifically, for a certain candidate object, the area where the candidate object is located can be cropped from the image to be audited to obtain the candidate image corresponding to the candidate object. Then, the candidate image corresponding to the candidate object is input into the corner detection algorithm to obtain the image features of the candidate image output by the corner detection algorithm, that is, the first image features corresponding to the candidate object.
[0162] Alternatively, the candidate image corresponding to the candidate object can also be input into the deep learning-based image feature extraction model to obtain the image features of the candidate image output by the deep learning-based feature extraction model, that is, the first image features corresponding to the candidate object.
[0163] S1032. From all candidate objects, determine all objects whose corresponding first image features match the first preset category features corresponding to any first preset category as the target objects.
[0164] Among them, the first preset category features corresponding to the first preset category include: the target image features corresponding to the first preset category and the target text features corresponding to the first preset category.
[0165] For each candidate object, calculate the similarity between the first image features corresponding to the candidate object and the first preset category features corresponding to each first preset category respectively. If there is a first preset category feature whose similarity with the first image features corresponding to the candidate object is greater than the third preset similarity threshold, it is considered that the candidate object matches the first preset category features corresponding to any first preset category, and the candidate object is used as the target object; if there is no first preset category feature whose similarity with the first image features corresponding to the candidate object is greater than the third preset similarity threshold, it is considered that the candidate object does not match the first preset category features corresponding to all first preset categories, and the candidate object is not used as the target object.
[0166] The similarity between the first image features and the first preset category features can be calculated by means such as cosine similarity, Euclidean distance, and Mahalanobis distance. The third preset similarity threshold can be set according to the user's needs. For example, the third preset similarity threshold can be set to 0.75, 0.8, 0.85, etc.
[0167] By selecting this embodiment, an image including an object of a first preset category can be obtained as the target image corresponding to the first preset category. By extracting the image features of the target image corresponding to the first preset category, the target image features corresponding to the first preset category are obtained, and by extracting the text features of the text input for the first preset category, the target text features corresponding to the first preset category are obtained. Based on the target image features corresponding to the first preset category and the feature matching between the target text features corresponding to the first preset category and the first image features corresponding to the candidate object, among all the candidate objects, the objects whose corresponding first image features match the first preset category features corresponding to any first preset category are determined as the target objects. Compared with the method of feature matching only through one dimension of features, i.e., text features or image features, this embodiment can integrate two different dimensions of features, namely target text features and target image features, improve the accuracy of feature matching, improve the accuracy of the determined target objects, and thus improve the accuracy of image review.
[0168] The method for obtaining the first preset category features and determining the target objects has been exemplarily described above. As described in S105 above, it can be determined through feature matching that there is no object of the second preset category among the candidate objects. The feature matching method depends on the obtained second preset category features corresponding to each second preset category. Therefore, the method for obtaining the second preset category features corresponding to each second preset category will be exemplarily described below. Refer to Figure 4 , the second preset category features corresponding to the second preset category include: the target image features corresponding to the second preset category and the target text features corresponding to the second preset category. In this example, the method for obtaining the second preset category features corresponding to the second preset category includes:
[0169] S401. For each second preset category, obtain an image including an object of the second preset category as the target image corresponding to the second preset category.
[0170] The methods for obtaining the second preset category features corresponding to different second preset categories are the same. Therefore, below, only one second preset category will be taken as an example to illustrate the method for obtaining the second preset category features corresponding to the second preset category.
[0171] If an image includes an object of the second preset category, then the image is the target image corresponding to the second preset category. The target image corresponding to the second preset category can be one or more. Exemplarily, assume that the second preset category is a firearm and Image 4 includes a rifle, then Image 4 is the target image corresponding to the second preset category. Assume that the second preset category is a controlled knife, Image 5 includes a switchblade, and Image 6 includes a dagger, then both Image 5 and Image 6 are target images corresponding to the second preset category.
[0172] S402. For each second preset category, extract the image features of the target image corresponding to the second preset category to obtain the target image features corresponding to the second preset category.
[0173] Any image feature extraction method can be used to extract the image features of the target image to obtain the target image features corresponding to the second preset category. For example, the image feature extraction method can be corner detection, an image feature extraction model based on deep learning, wavelet transform method, etc. If there are multiple target images corresponding to the second preset category, the image features corresponding to different target images are extracted by the same image feature extraction method.
[0174] Specifically, the target image corresponding to the second preset category can be input into a corner detection algorithm to obtain the image features of the target image corresponding to the second preset category output by the corner detection algorithm, that is, the target image features corresponding to the second preset category are obtained. Alternatively, the target image corresponding to the second preset category can be input into an image feature extraction model based on deep learning to obtain the image features of the target image corresponding to the second preset category output by the feature extraction model based on deep learning, that is, the target image features corresponding to the second preset category are obtained.
[0175] S403. For each second preset category, extract the text features of the text input for the second preset category to obtain the target text features corresponding to the second preset category.
[0176] The text input for the second preset category is the category name of the second preset category. Exemplarily, assuming the second preset category is a firearm, the text input for the second preset category should be "firearm" or text with the same or similar semantics as "firearm", such as "gun"; assuming the second preset category is a controlled knife, the text input for the second preset category should be "controlled knife" or text with the same or similar semantics as "controlled knife", such as "controlled knives".
[0177] Any text feature extraction method can be used to extract the text features of the text input for the second preset category to obtain the target text features corresponding to the second preset category. For example, the text feature extraction method can be Word Embedding (word vector model), a text feature extraction model based on deep learning, etc.
[0178] Specifically, the text input for the second preset category can be input into Word Embedding to obtain the text features output by Word Embedding, that is, the target text features corresponding to the second preset category are obtained. Alternatively, the text input for the second preset category can be input into a text feature extraction model based on deep learning to obtain the text features output by the text feature extraction model based on deep learning, that is, the target text features corresponding to the second preset category are obtained.
[0179] The second preset category features corresponding to the second preset category are the target image features corresponding to the second preset category obtained in S402 and the target text features corresponding to the second preset category obtained in S403. Any feature fusion method can be used to fuse the target image features corresponding to the second preset category and the target text features corresponding to the second preset category to obtain the second preset category features corresponding to the second preset category. For example, the feature fusion method can be cross-modal fusion, splicing, etc.
[0180] Based on Figure 4 the acquisition method of the second preset category features corresponding to the second preset category shown in the embodiment, it can be determined whether there is an object of the second preset category among the candidate objects through the following feature matching method. See Figure 5 , the method includes:
[0181] S501, for each candidate object, extract the image features of the candidate object from the image to be reviewed as the first image features corresponding to the candidate object.
[0182] S501 is the same as the aforementioned S1031. For the relevant description in the aforementioned S1031, reference can be made and will not be repeated here. The first image features corresponding to different candidate objects are extracted by the same image feature extraction method. Moreover, the first image features corresponding to each candidate object and the target image features corresponding to the second preset category are also extracted by the same image feature extraction method.
[0183] S502, if there is an object among the candidate objects whose corresponding first image features match the second preset category features corresponding to any second preset category, then there is an object of the second preset category among the candidate objects.
[0184] Among them, the second preset category features corresponding to the second preset category include: the target image features corresponding to the second preset category and the target text features corresponding to the second preset category.
[0185] S503, if there is no object among the candidate objects whose corresponding first image features match the second preset category features corresponding to any second preset category, then there is no object of the second preset category among the candidate objects.
[0186] In S502 - S503, for each candidate object, the similarity between the first image feature corresponding to the candidate object and the second preset category features corresponding to each second preset category is calculated respectively. If, for a certain candidate object, there exists a second preset category feature whose similarity with the first image feature corresponding to the candidate object is greater than the fourth preset similarity threshold, it indicates that the category of the candidate object is a certain second preset category, and it also indicates that among all candidate objects, there exists an object whose corresponding first image feature matches the second preset category feature corresponding to any second preset category, that is, among all candidate objects, there exists an object whose category is the second preset category.
[0187] Exemplarily, assume that the candidate objects are object 1 - object 3, and the first image features corresponding to object 1 - object 3 are feature 1, feature 2, and feature 3 respectively; the second preset categories are category 1 - category 2, and the second preset category features corresponding to category 1 - category 2 are feature 4 and feature 5 respectively. Among them, the similarity between feature 1 and feature 4 is greater than the fourth similarity threshold, then it indicates that the category of object 1 is category 1, and among all candidate objects, there exists an object whose category is the second preset category.
[0188] If, for all candidate objects, there does not exist a second preset category feature whose similarity with the first image feature corresponding to the candidate object is greater than the fourth preset similarity threshold, it indicates that the categories of all candidate objects are not a certain second preset category, and it also indicates that among all candidate objects, there does not exist an object whose corresponding first image feature matches the second preset category feature corresponding to any second preset category, that is, among all candidate objects, there does not exist an object whose category is the second preset category.
[0189] Exemplarily, assume that the candidate objects are object 1 - object 3, and the first image features corresponding to object 1 - object 3 are feature 1, feature 2, and feature 3 respectively; the second preset categories are category 1 - category 2, and the second preset category features corresponding to category 1 - category 2 are feature 4 and feature 5 respectively. When the similarities between feature 1 and feature 4, feature 1 and feature 5, feature 2 and feature 4, feature 2 and feature 5, feature 3 and feature 4, and feature 3 and feature 5 are all not greater than the fourth similarity threshold, it is considered that for all candidate objects, there does not exist a second preset category feature whose similarity with the first image feature corresponding to the candidate object is greater than the fourth preset similarity threshold, that is, among all candidate objects, there does not exist an object whose category is the second preset category.
[0190] The similarity between the first image feature and the second preset category feature can be calculated by means such as cosine similarity, Euclidean distance, Mahalanobis distance, etc. The fourth preset similarity threshold can be set according to the user's needs. For example, the fourth preset similarity threshold can be set to: 0.75, 0.8, 0.85, etc. The third preset similarity threshold and the fourth preset similarity threshold can be the same or different.
[0191] By selecting this embodiment, an image including an object of the second preset category can be obtained as the target image corresponding to the second preset category. By extracting the image features of the target image corresponding to the second preset category, the target image features corresponding to the second preset category can be obtained, and by extracting the text features of the text input for the second preset category, the target text features corresponding to the second preset category can be obtained. Based on the target image features corresponding to the second preset category and the target text features corresponding to the second preset category and the second image features corresponding to the candidate object, feature matching is performed to determine whether there is an object of the second preset category among the candidate objects. Compared with the method of feature matching only through one-dimensional features such as text features or image features, this embodiment can comprehensively use two different-dimensional features of target text features and target image features, improve the accuracy of feature matching, improve the accuracy of determining whether there is an object of the second preset category among the candidate objects, and thus improve the accuracy of image review.
[0192] For the acquisition method of the first preset category feature corresponding to the first preset category and the second preset category feature corresponding to the second preset category, in a possible embodiment, the first preset category feature corresponding to the first preset category can be obtained by the method in the foregoing S201-S203, and the second preset category feature corresponding to the second preset category can be obtained by the method in the foregoing S401-S402 or the method in the foregoing S403. That is, in this embodiment, the first preset category feature corresponding to the first preset category includes: the target image feature corresponding to the first preset category and the target text feature corresponding to the first preset category, and the second preset category feature corresponding to the second preset category includes: the target image feature corresponding to the second preset category or the target text feature corresponding to the second preset category.
[0193] In another possible embodiment, the first preset category features corresponding to the first preset category can be obtained in the manner of the foregoing S201-S202 or the manner of the foregoing S203, and the second preset category features corresponding to the second preset category can be obtained in the manner of the foregoing S401-S403. That is, in this embodiment, the first preset category features corresponding to the first preset category include: the target image features corresponding to the first preset category or the target text features corresponding to the first preset category, and the second preset category features corresponding to the second preset category include: the target image features corresponding to the second preset category and the target text features corresponding to the second preset category.
[0194] In yet another possible embodiment, the first preset category features corresponding to the first preset category can be obtained in the manner of the foregoing S201-S203, and the second preset category features corresponding to the second preset category can be obtained in the manner of the foregoing S401-S403. That is, in this embodiment, the first preset category features corresponding to the first preset category include: the target image features corresponding to the first preset category and the target text features corresponding to the first preset category, and the second preset category features corresponding to the second preset category include: the target image features corresponding to the second preset category and the target text features corresponding to the second preset category.
[0195] The method for obtaining the second preset category features and the method for determining whether there is an object of the second preset category among the candidate objects have been exemplarily described above. It can be understood that a preset category often includes multiple objects, and the image features of different objects have certain differences. Therefore, if the target image features corresponding to the preset category are obtained only based on a target image of an object including the preset category, the obtained target image features corresponding to the preset category may not be able to reflect the image features of different objects in the preset category, resulting in the target image features corresponding to the preset category being incomplete, thereby reducing the accuracy of feature matching and the accuracy of image review.
[0196] Based on this, in order to improve the comprehensiveness of the target image features corresponding to the obtained preset category, improve the accuracy of feature matching, and improve the accuracy of image review. In one possible embodiment, refer to Figure 6a , the target image features corresponding to the preset category are obtained in the following manner, including:
[0197] S601, for each preset category, obtain the images of the objects including the preset category as the target images corresponding to the preset category.
[0198] The number of target images corresponding to the preset category is multiple.
[0199] The methods for obtaining the target image features corresponding to different preset categories are the same. Therefore, in the following, only one preset category will be taken as an example to illustrate the method for obtaining the target image features corresponding to the preset category. Among them, the preset category can refer to the aforementioned first preset category, or the aforementioned second preset category, or both the aforementioned first preset category and the second preset category.
[0200] If the preset category is the first preset category, then S601 is similar to the aforementioned S201, with the only difference being that the number of target images corresponding to the first preset category in S201 can be one or more, while the number of target images corresponding to the first preset category in S601 is multiple, and the rest are the same. Therefore, reference can be made to the relevant description in S201 above, and details will not be repeated here.
[0201] If the preset category is the second preset category, then S601 is similar to the aforementioned S401, with the only difference being that the number of target images corresponding to the second preset category in S401 can be one or more, while the number of target images corresponding to the second preset category in S601 is multiple, and the rest are the same. Therefore, reference can be made to the relevant description in S401 above, and details will not be repeated here.
[0202] S602. For each preset category, extract the image features of each target image corresponding to the preset category respectively as the second image features corresponding to each target image.
[0203] Any image feature extraction method can be used to extract the image features of each target image respectively to obtain the second image features corresponding to each target image. For example, the image feature extraction method can be corner detection, an image feature extraction model based on deep learning, wavelet transform method, etc. The image features corresponding to different target images are extracted by the same image feature extraction method.
[0204] Specifically, each target image corresponding to the preset category can be input into the corner detection algorithm respectively to obtain the second image features corresponding to each target image output by the corner detection algorithm. Or, each target image corresponding to the preset category can be input into the image feature extraction model based on deep learning respectively to obtain the second image features corresponding to each target image output by the feature extraction model based on deep learning.
[0205] S603. For each preset category, obtain the target image features corresponding to the preset category according to the second image features corresponding to each target image.
[0206] Among them, the similarity between the target image features corresponding to the preset category and the second image features corresponding to each target image is greater than the first preset similarity threshold.
[0207] Statistically analyze the second image features corresponding to each target image to obtain the target image features corresponding to the preset category, so that the similarity between the target image features corresponding to the preset category and the second image features corresponding to each target image is greater than the first preset similarity threshold. The first preset similarity threshold can be set according to the user's needs. For example, the first preset similarity threshold can be set to: 0.75, 0.8, 0.85, etc.
[0208] In a possible embodiment, the mean value of the second image features corresponding to each target image can be calculated as the target image features corresponding to the preset category.
[0209] In another possible embodiment, the second image features corresponding to each target image can be clustered to obtain the first cluster center feature as the target image features corresponding to the preset category.
[0210] Specifically, input the second image features corresponding to each target image into the clustering algorithm, and set the number of clusters in the clustering algorithm to 1 to obtain a cluster generated by the clustering algorithm and the cluster center feature of this cluster. This cluster corresponds to a preset category, and the cluster center feature of this cluster is the feature of the preset category corresponding to this cluster. The cluster center feature of this cluster is the first cluster center feature, and the cluster center feature of this cluster is used as the target image features corresponding to the preset category.
[0211] Selecting this embodiment, for each preset category, multiple images including objects of the preset category can be obtained as the target images corresponding to the preset category, and the image features of each target image corresponding to the preset category are extracted as the second image features corresponding to each target image, and the target image features corresponding to the preset category are obtained according to the second image features corresponding to each target image, so that the similarity between the target image features corresponding to the preset category and the second image features corresponding to each target image is greater than the first preset similarity threshold. In this embodiment, the target image features corresponding to the preset category can be statistically obtained through the second image features corresponding to multiple target images, so that the comprehensiveness of the obtained target image features can be improved, the accuracy of feature matching can be improved, and the accuracy of image review can be improved.
[0212] Moreover, a preset category may have multiple different text representations. For example, the category "gun" can also be called "firearm", and the category "person" can also be called "human", etc. Therefore, if the target text features corresponding to the preset category are obtained only according to one category name of the preset category, the obtained target text features corresponding to the preset category may not be able to reflect the text features of different names of the preset category, making the target text features corresponding to the preset category not comprehensive enough, thus reducing the accuracy of feature matching and the accuracy of image review.
[0213] Based on this, in order to improve the comprehensiveness of the target text features corresponding to the preset categories obtained, improve the accuracy of feature matching, and improve the accuracy of image review. In a possible embodiment, referring to Figure 6b , the target text features corresponding to the preset categories are obtained in the following manner, including:
[0214] S611. For each preset category, obtain a target text with the same text semantics as the text input for the preset category.
[0215] The methods for obtaining the target text features corresponding to different preset categories are the same. Therefore, in the following, only one preset category is taken as an example to illustrate the method for obtaining the target text features corresponding to the preset category. Among them, the preset category may refer to the aforementioned first preset category, or the aforementioned second preset category, or both the aforementioned first preset category and the second preset category.
[0216] If the text input for the preset category is the category name of the preset category, then the target text with the same text semantics as the text input for the preset category refers to: the text with the same semantics as the category name of the preset category. Exemplarily, assuming the preset category is "firearm", and the text input for the preset category is "firearm", then the target text may be: gun, firearm, etc.
[0217] The execution subject can obtain the target text by means of the user inputting a text with the same text semantics as the text input for the preset category; it can also input the text input for the preset category into a deep text matching model to obtain the target text output by the deep text matching model with the same text semantics as the text input for the preset category. The execution subject can also obtain the target text with the same text semantics as the text input for the preset category by means other than the above two methods, and the present application does not make any limitation thereto.
[0218] S612. For each preset category, extract the first text feature of the text input for the preset category, and respectively extract the text features of each target text as the second text features corresponding to each target text.
[0219] Any text feature extraction method can be used to extract the text features of the text input for the preset category and each target text to obtain the first text feature and the second text features corresponding to each target text. For example, the text feature extraction method can be Word Embedding (word vector model), a text feature extraction model based on deep learning, etc. The first text feature of the text input for the preset category and the second text features corresponding to different target texts are extracted by the same text feature extraction method.
[0220] Specifically, the text input for the preset category and each target text can be respectively input into WordEmbedding to obtain the first text feature output by WordEmbedding and the second text features corresponding to each target text. Alternatively, the text input for the preset category and each target text can be respectively input into a deep learning-based text feature extraction model to obtain the first text feature output by the deep learning-based text feature extraction model and the second text features corresponding to each target text.
[0221] S613. For each preset category, according to the first text feature and the second text features corresponding to each target text, obtain the target text feature corresponding to the preset category.
[0222] Among them, the similarity between the target text feature corresponding to the preset category and the second text features corresponding to each target text is greater than the second preset similarity threshold, and the similarity between the target text feature corresponding to the preset category and the first text feature is greater than the second preset similarity threshold.
[0223] Perform statistics on the first text feature and the second text features corresponding to each target text to obtain the target text feature corresponding to the preset category, so that the similarity between the target text feature corresponding to the preset category and the second text features corresponding to each target text is greater than the second preset similarity threshold, and the similarity between the target text feature corresponding to the preset category and the first text feature is greater than the second preset similarity threshold. The second preset similarity threshold can be set according to the user's needs. For example, the second preset similarity threshold can be set to: 0.75, 0.8, 0.85, etc. The first preset similarity threshold and the second preset similarity threshold can be the same or different.
[0224] In a possible embodiment, the mean value of the first text feature and the second text features corresponding to each target text can be calculated as the target text feature corresponding to the preset category.
[0225] In another possible embodiment, the first text feature and the second text features corresponding to each target text can be clustered to obtain the second cluster center feature as the target text feature corresponding to the preset category.
[0226] Specifically, input the first text feature and the second text features corresponding to each target text into a clustering algorithm, and set the number of clusters in the clustering algorithm to 1 to obtain a cluster generated by the clustering algorithm and the cluster center feature of this cluster. This cluster corresponds to a preset category, and the cluster center feature of this cluster is the feature of the preset category corresponding to this cluster. The cluster center feature of this cluster is the second cluster center feature, and the cluster center feature of this cluster is used as the target text feature corresponding to the preset category.
[0227] Selecting this embodiment, for each preset category, the target text with the same text semantics as the text input for the preset category can be obtained, the first text feature of the text input for the preset category can be extracted, and the text features of each target text can be extracted respectively as the second text features corresponding to each target text. According to the first text feature and the second text features corresponding to each target text, the target text feature corresponding to the preset category is obtained, so that the similarity between the target text feature corresponding to the preset category and the second text features corresponding to each target text is greater than the second preset similarity threshold, and the similarity with the first text feature is greater than the second preset similarity threshold. In this embodiment, the target text feature corresponding to the preset category can be obtained by statistically analyzing the second text features corresponding to each of the multiple target texts with the same text semantics as the text input for the preset category and the first text feature of the text input for the preset category, so that the comprehensiveness of the obtained target text feature can be improved, the accuracy of feature matching can be improved, and the accuracy of image review can be improved.
[0228] In a possible embodiment, for the first preset category, the target image features corresponding to each first preset category can be obtained in the following manner: for each first preset category, an image including an object of the first preset category is obtained as the target image 1 corresponding to the first preset category; the number of the target images 1 corresponding to the first preset category is multiple; for each first preset category, the image features of each target image 1 corresponding to the first preset category are extracted respectively as the second image features 1 corresponding to each target image 1; for each first preset category, according to the second image features 1 corresponding to each target image 1, the target image feature corresponding to the first preset category is obtained, wherein the similarity between the target image feature corresponding to the first preset category and the second image features 1 corresponding to each target image 1 is greater than the first preset similarity threshold.
[0229] This embodiment is similar to the Figure 6a embodiment shown above, the only difference being that the preset category in the Figure 6a embodiment shown above is replaced with the first preset category, the target image is replaced with the target image 1, and the second image feature is replaced with the second image feature 1, and the rest are the same. Therefore, the relevant descriptions in the Figure 6a embodiment shown above can be referred to and will not be elaborated here.
[0230] For the first preset category, the target text features corresponding to each first preset category can be obtained in the following manner: For each first preset category, obtain target text 1 with the same text semantics as the text input for the first preset category; For each first preset category, extract the first text feature 1 of the text input for the first preset category, and respectively extract the text features of each target text 1 as the second text feature 1 corresponding to each target text 1; For each first preset category, based on the first text feature 1 and the second text feature 1 corresponding to each target text 1, obtain the target text feature corresponding to the first preset category; wherein, the similarity between the target text feature corresponding to the first preset category and the second text feature 1 corresponding to each target text 1 is greater than the second preset similarity threshold, and the similarity with the first text feature 1 is greater than the second preset similarity threshold.
[0231] This embodiment is similar to the Figure 6b embodiment shown above, with the only difference being that the preset category in the Figure 6b embodiment shown above is replaced by the first preset category, the target text is replaced by target text 1, the first text feature is replaced by the first text feature 1, and the second text feature is replaced by the second text feature 1. The rest are the same, so reference can be made to the relevant descriptions in the Figure 6b embodiment shown above and will not be elaborated here.
[0232] For the second preset category, the target image features corresponding to each second preset category can be obtained in the following manner: For each second preset category, obtain an image including an object of the second preset category as the target image 2 corresponding to the second preset category; The number of target images 2 corresponding to the second preset category is multiple; For each second preset category, respectively extract the image features of each target image 2 corresponding to the second preset category as the second image feature 2 corresponding to each target image 2; For each second preset category, based on the second image feature 2 corresponding to each target image 2, obtain the target image feature corresponding to the second preset category, wherein the similarity between the target image feature corresponding to the second preset category and the second image feature 2 corresponding to each target image 2 is greater than the first preset similarity threshold.
[0233] This embodiment is similar to the Figure 6a embodiment shown above, with the only difference being that the preset category in the Figure 6a embodiment shown above is replaced by the second preset category, the target image is replaced by target image 2, and the second image feature is replaced by the second image feature 2. The rest are the same, so reference can be made to the relevant descriptions in the Figure 6a embodiment shown above and will not be elaborated here.
[0234] For the second preset category, the target text features corresponding to each second preset category can be obtained in the following manner: for each second preset category, obtain the target text 2 with the same text semantics as the text input for the second preset category; for each second preset category, extract the first text feature 2 of the text input for the second preset category, and respectively extract the text features of each target text 2 as the second text feature 2 corresponding to each target text 2; for each second preset category, obtain the target text feature corresponding to the second preset category according to the first text feature 2 and the second text feature 2 corresponding to each target text 2; wherein, the similarity between the target text feature corresponding to the second preset category and the second text feature 2 corresponding to each target text 2 is greater than the second preset similarity threshold, and the similarity with the first text feature 2 is greater than the second preset similarity threshold.
[0235] This embodiment is similar to the Figure 6b embodiment shown above, the only difference being that the preset category in the Figure 6b embodiment shown above is replaced with the second preset category, the target text is replaced with the target text 2, the first text feature is replaced with the first text feature 2, and the second text feature is replaced with the second text feature 2, and the rest are the same. Therefore, reference can be made to the relevant descriptions in the Figure 6b embodiment shown above, and details will not be elaborated here.
[0236] The acquisition methods of the target image feature and the target text feature have been exemplarily described above. Referring to the description in S102 above, it can be understood that if only semantic segmentation is used to perform object recognition on the image to be audited, different objects in the same category cannot be distinguished; if only instance segmentation is used to perform object recognition on the image to be audited, the categories of each candidate object cannot be directly determined. Based on this, in order to more accurately distinguish each candidate object in the image to be audited and be able to determine the category of each candidate object, the present invention provides a method for performing object recognition on the image to be audited. Refer to Figure 7 , the method includes:
[0237] S1021, for each pixel point in the image to be audited, identify the category label and instance label of the pixel point.
[0238] The class label is used to indicate the class to which a pixel belongs, and the instance label is used to indicate the specific object to which a pixel belongs. Exemplarily, assume that pixel 1 belongs to the pixels that make up object 1. Then, the class to which pixel 1 belongs is a person, and the specific object to which it belongs is person 1. That is, the class label of pixel 1 is a person, and the instance label is person 1. Assume that pixel 2 belongs to the pixels that make up object 2. Then, the class to which pixel 2 belongs is a person, and the specific object to which it belongs is person 2. That is, the class label of pixel 2 is a person, and the instance label is person 2.
[0239] S1022. For multiple pixels with the same class label and the same instance label, determine the object formed by the multiple pixels as a candidate object, and determine the region formed by the multiple pixels as the region where the candidate object is located.
[0240] Group the pixels according to their respective class labels and instance labels. If two pixels have the same class label and the same instance label, then these two pixels are the pixels included in the same object, and these two pixels can be grouped into the same group. Group the pixels in the above manner, that is, group the pixels with the same class label and the same instance label into the same group. Each group includes multiple pixels with the same class label and the same instance label.
[0241] For each group of pixels, use the object formed by the group of pixels as the candidate object, and use the region formed by the group of pixels as the region where the candidate object is located, so as to obtain all the candidate objects included in the image to be reviewed and the regions where each candidate object is located.
[0242] By selecting this embodiment, the class to which a pixel belongs can be reflected by the class label of the pixel, the specific object to which a pixel belongs can be reflected by the instance label of the pixel, and for multiple pixels with the same class label and the same instance label, determine the object formed by the multiple pixels as a candidate object, and determine the region formed by the multiple pixels as the region where the candidate object is located. Through the above method, the class of each candidate object can be determined, and the accurate distinction of different candidate objects in the image to be reviewed can be realized, the accuracy of the determined candidate objects can be improved, and thus the accuracy of image review can be improved.
[0243] Figure 7The illustrated embodiment is a method for object recognition of an image to be reviewed through panoramic segmentation. Specifically, the image to be reviewed can be preprocessed to obtain the preprocessed image to be reviewed, and the preprocessed image to be reviewed is input into an open-set panoramic segmentation algorithm based on CLIP (Contrastive Language-Image Pretraining) features, so as to obtain all candidate objects included in the image to be reviewed output by the open-set panoramic segmentation algorithm based on CLIP features, the category of each candidate object, and the mask of each candidate object. The mask of each candidate object is the area where each candidate object is located. Among them, the preprocessing can include operations such as adjusting the image size and standardization, so that the preprocessed image to be reviewed meets the requirements of the open-set panoramic segmentation algorithm based on CLIP features for the input image.
[0244] The open-set panoramic segmentation algorithm based on CLIP features obtains the category of candidate objects by means of feature matching between the image CLIP features of the candidate objects after segmentation and the CLIP features of preset categories in the database. Since the CLIP features of the preset categories are obtained by extracting the text features of the text corresponding to any preset category and the image features of the image corresponding to the preset category, and are not limited to fixed text features or fixed image features, the CLIP features can be regarded as an open-set feature, and the panoramic segmentation algorithm such as the open-set panoramic segmentation algorithm based on CLIP features is an open-set panoramic segmentation algorithm. Among them, the text corresponding to the preset category is the aforementioned Figure 6b For the text input for the preset category and the target text in the embodiment, the image corresponding to the preset category is the aforementioned Figure 6a target image corresponding to the preset category in the embodiment.
[0245] The flowchart of the image review method provided by the present invention can be as Figure 8 shown. The method includes:
[0246] S801, import a picture.
[0247] The picture is the aforementioned image to be reviewed. S801 is equivalent to the aforementioned S101, and the relevant description in the aforementioned S101 can be referred to and will not be elaborated here.
[0248] S802, perform picture preprocessing.
[0249] S802 is equivalent to preprocessing the image to be reviewed to obtain the preprocessed image to be reviewed. The preprocessing can include operations such as adjusting the image size and standardization, so that the preprocessed image to be reviewed meets the requirements of the open-set panoramic segmentation algorithm based on CLIP features for the input image when performing panoramic segmentation in S804.
[0250] S803 extracts the CLIP features of the blacklist and whitelist targets.
[0251] The blacklist is the aforementioned second preset category, and the whitelist is the aforementioned first preset category. S803 is equivalent to obtaining the first preset category features corresponding to the first preset category and the second preset category features corresponding to the second preset category. Refer to the relevant descriptions in the foregoing Figure 2 , Figure 4 , Figure 6a and Figure 6b the relevant descriptions in the embodiments, and will not be elaborated here.
[0252] S804 performs panoramic segmentation.
[0253] S804 is equivalent to S1021 - S1022 in the foregoing Figure 7 embodiments. Refer to the relevant descriptions of S1021 - S1022 in the foregoing Figure 7 embodiments, and will not be elaborated here.
[0254] S805 performs feature comparison and logical processing.
[0255] S805 is equivalent to S103 in the foregoing. Refer to the relevant descriptions of S103, and will not be elaborated here.
[0256] S806 calculates the proportion of the whitelist target area.
[0257] The whitelist target is the aforementioned target object. S806 is equivalent to S104 in the foregoing. Refer to the relevant descriptions of S104, and will not be elaborated here.
[0258] S807 determines whether the area proportion is greater than the threshold T.
[0259] The area proportion is the aforementioned ratio, and the threshold T is the aforementioned preset threshold. S807 is equivalent to determining whether the ratio is greater than the preset threshold. If so, execute S808; if not, execute S809. In S105 above, an exemplary description of how to determine whether the ratio is greater than the preset threshold has been given. Refer to the relevant descriptions in S105, and will not be elaborated here.
[0260] S808 determines the existence of the blacklist target.
[0261] The blacklist target is the candidate object whose category is the second preset category mentioned above. S808 is equivalent to determining whether there is an object whose category is the second preset category among the candidate objects. If so, S810 is executed; if not, S811 is executed. In the previous S105, an exemplary description of how to determine whether there is an object whose category is the second preset category among the candidate objects has been given. You can refer to the relevant description in the previous S105 and will not be elaborated here.
[0262] S809, submit for further review.
[0263] S809 is equivalent to if the ratio is not greater than the preset threshold, then send the image to be reviewed to the manual end for review.
[0264] S810, mark as a whitelist image.
[0265] S810 is equivalent to if the ratio is greater than the preset threshold and there is no object whose category is the second preset category among the candidate objects, then determine that the image to be reviewed passes the review. After S810, S812 is executed.
[0266] S811, submit for further review.
[0267] S811 is equivalent to if there is an object whose category is the second preset category among the candidate objects, then send the image to be reviewed to the manual end for review.
[0268] S812, output the result.
[0269] S812 is equivalent to after determining that the image to be reviewed passes the review, output the review result of passing the review.
[0270] Corresponding to the foregoing image review method, an embodiment of the present invention further provides an image review device. See Figure 9 , the device includes:
[0271] An image acquisition module 901, configured to acquire an image to be reviewed;
[0272] A candidate object recognition module 902, configured to perform object recognition on the image to be reviewed to obtain all candidate objects included in the image to be reviewed and the regions where each candidate object is located;
[0273] A target object determination module 903, configured to determine all objects whose category is the first preset category from the candidate objects as target objects; where the first preset category is a category that includes objects that are all non-violating objects;
[0274] A ratio calculation module 904, configured to calculate the ratio of the total area of the regions where all target objects are located to the area of the image to be reviewed;
[0275] An audit determination module 905 is configured to determine that the image to be audited passes the audit if the ratio is greater than a preset threshold and there is no object of a second preset category among the candidate objects; wherein, the second preset category is a category that includes objects all belonging to the category of illegal objects.
[0276] In a possible embodiment, the apparatus further includes:
[0277] A first target image acquisition module, configured to, for each first preset category, acquire an image including an object of the first preset category as the target image corresponding to the first preset category;
[0278] A first image feature extraction module, configured to, for each first preset category, extract the image features of the target image corresponding to the first preset category to obtain the target image features corresponding to the first preset category;
[0279] A first text feature extraction module, configured to, for each first preset category, extract the text features of the text input for the first preset category to obtain the target text features corresponding to the first preset category;
[0280] Determining all objects of the first preset category from among the candidate objects as target objects includes:
[0281] For each candidate object, extract the image features of the candidate object from the image to be audited as the first image features corresponding to the candidate object;
[0282] Determine, from among the candidate objects, all objects whose corresponding first image features match the first preset category features corresponding to any first preset category as target objects; wherein, the first preset category features corresponding to the first preset category include: the target image features corresponding to the first preset category and the target text features corresponding to the first preset category.
[0283] In a possible embodiment, the apparatus further includes:
[0284] A second target image acquisition module, configured to, for each second preset category, acquire an image including an object of the second preset category as the target image corresponding to the second preset category;
[0285] A second image feature extraction module, configured to, for each second preset category, extract the image features of the target image corresponding to the second preset category to obtain the target image features corresponding to the second preset category;
[0286] A second text feature extraction module, configured to, for each second preset category, extract the text features of the text input for the second preset category to obtain the target text features corresponding to the second preset category;
[0287] Determine whether there is an object of the second preset category among the candidate objects in the following manner:
[0288] For each candidate object, extract the image features of the candidate object from the image to be audited as the first image features corresponding to the candidate object;
[0289] If there is an object among the candidate objects whose corresponding first image features match the second preset category features corresponding to any second preset category, then there is an object of the second preset category among the candidate objects; wherein, the second preset category features corresponding to the second preset category include: the target image features corresponding to the second preset category and the target text features corresponding to the second preset category;
[0290] If there is no object among the candidate objects whose corresponding first image features match the second preset category features corresponding to any second preset category, then there is no object of the second preset category among the candidate objects.
[0291] In a possible embodiment, the number of target images corresponding to the preset category is multiple;
[0292] Obtain the target image features corresponding to the preset category in the following manner, including:
[0293] Extract the image features of each target image corresponding to the preset category as the second image features corresponding to each target image respectively;
[0294] Obtain the target image features corresponding to the preset category according to the second image features corresponding to each target image respectively; wherein, the similarity between the target image features corresponding to the preset category and the second image features corresponding to each target image respectively is greater than the first preset similarity threshold;
[0295] Obtain the target text features corresponding to the preset category in the following manner, including:
[0296] Obtain the target text with the same text semantics as the text input for the preset category;
[0297] Extract the first text features of the text input for the preset category, and extract the text features of each target text as the second text features corresponding to each target text respectively;
[0298] Obtain the target text features corresponding to the preset category according to the first text features and the second text features corresponding to each target text respectively; wherein, the similarity between the target text features corresponding to the preset category and the second text features corresponding to each target text respectively is greater than the second preset similarity threshold, and the similarity with the first text features is greater than the second preset similarity threshold.
[0299] In a possible embodiment, obtaining the target image features corresponding to a preset category according to the second image features corresponding to each target image includes:
[0300] Calculating the mean value of the second image features corresponding to each target image as the target image features corresponding to the preset category;
[0301] Obtaining the target text features corresponding to a preset category according to the first text features and the second text features corresponding to each target text includes:
[0302] Calculating the mean value of the first text features and the second text features corresponding to each target text as the target text features corresponding to the preset category.
[0303] In a possible embodiment, obtaining the target image features corresponding to a preset category according to the second image features corresponding to each target image includes:
[0304] Clustering the second image features corresponding to each target image to obtain the first clustering center feature as the target image features corresponding to the preset category;
[0305] Obtaining the target text features corresponding to a preset category according to the first text features and the second text features corresponding to each target text includes:
[0306] Clustering the first text features and the second text features corresponding to each target text to obtain the second clustering center feature as the target text features corresponding to the preset category.
[0307] In a possible embodiment, the apparatus further includes:
[0308] An artificial review module, configured to send the image to be reviewed to the artificial terminal for review if the ratio is not greater than a preset threshold, or if there is an object with a second preset category among the candidate objects.
[0309] In a possible embodiment, performing object recognition on the image to be reviewed to obtain all candidate objects included in the image to be reviewed and the regions where each candidate object is located, includes:
[0310] For each pixel point in the image to be reviewed, identifying the category label and instance label of the pixel point;
[0311] For multiple pixel points with the same category label and the same instance label, determining the object formed by the multiple pixel points as a candidate object, and determining the region formed by the multiple pixel points as the region where the candidate object is located.
[0312] The embodiment of the present invention also provides an electronic device, such as Figure 10As shown in the figure, it includes a processor 1001, a communication interface 1002, a memory 1003, and a communication bus 1004. Among them, the processor 1001, the communication interface 1002, and the memory 1003 complete communication with each other through the communication bus 1004.
[0313] The memory 1003 is used to store computer programs.
[0314] When the processor 1001 is used to execute the program stored on the memory 1003, the following steps are implemented:
[0315] Obtain the image to be reviewed.
[0316] Perform object recognition on the image to be reviewed to obtain all candidate objects included in the image to be reviewed and the regions where each candidate object is located.
[0317] Determine all objects with the first preset category from each candidate object as target objects; among them, the first preset category is the category of objects that all included objects belong to non-violating objects.
[0318] Calculate the ratio of the total area of the regions where all target objects are located to the area of the image to be reviewed.
[0319] If the ratio is greater than the preset threshold and there is no object with the second preset category among each candidate object, it is determined that the image to be reviewed passes the review; among them, the second preset category is the category of objects that all included objects belong to violating objects.
[0320] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0321] The communication interface is used for communication between the above terminal and other devices.
[0322] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0323] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU for short), a Network Processor (NP for short), etc.; it may also be a Digital Signal Processor (DSP for short), an Application Specific Integrated Circuit (ASIC for short), a Field-Programmable Gate Array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0324] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the image review method described in any one of the above embodiments is implemented.
[0325] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, and when it runs on a computer, it causes the computer to execute the image review method described in any one of the above embodiments.
[0326] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that the computer can access, or a data storage device such as a server, a data center, etc. that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a Solid State Disk (SSD)).
[0327] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0328] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the apparatus, electronic device, storage medium and computer program product, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the partial description of the method embodiments for the relevant parts.
[0329] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. An image review method, characterized in that: The method comprises: Get the image to be reviewed; Performing object recognition on the image to be reviewed to obtain all candidate objects included in the image to be reviewed and the regions where each candidate object is located; Determine all objects of a first preset category from the candidate objects as target objects; wherein the first preset category includes objects that are not in violation of the rules; Calculating the ratio of the total area of the regions where all the target objects are located to the area of the image to be reviewed; If the ratio is greater than a preset threshold, and there is no object of the second preset category among the candidate objects, then it is determined that the image to be reviewed has passed the review; wherein the second preset category is a category in which all the objects included belong to illegal objects.
2. The method according to claim 1, characterized in that The method further comprises: For each first preset category, acquiring an image including an object of the first preset category as a target image corresponding to the first preset category; For each of the first preset categories, extracting image features of a target image corresponding to the first preset category to obtain target image features corresponding to the first preset category; For each of the first preset categories, extract text features of text input for the first preset category to obtain target text features corresponding to the first preset category; The step of determining all objects of the first preset category from the candidate objects as target objects includes: For each candidate object, extracting an image feature of the candidate object from the image to be reviewed as a first image feature corresponding to the candidate object; From each of the candidate objects, determine an object whose corresponding first image features all match the first preset category features corresponding to any of the first preset categories as a target object; wherein the first preset category features corresponding to the first preset category include: a target image feature corresponding to the first preset category, and a target text feature corresponding to the first preset category.
3. The method according to claim 1, characterized in that The method further comprises: For each second preset category, acquiring an image including an object of the second preset category as a target image corresponding to the second preset category; For each of the second preset categories, extracting image features of the target image corresponding to the second preset category to obtain target image features corresponding to the second preset category; For each of the second preset categories, extract text features of text input for the second preset category to obtain target text features corresponding to the second preset category; Determine whether there is an object of the second preset category among the candidate objects by: For each candidate object, extracting an image feature of the candidate object from the image to be reviewed as a first image feature corresponding to the candidate object; If there is an object in each of the candidate objects whose corresponding first image feature matches the second preset category feature corresponding to any of the second preset categories, then there is an object of the second preset category in each of the candidate objects; wherein the second preset category feature corresponding to the second preset category includes: a target image feature corresponding to the second preset category, and a target text feature corresponding to the second preset category; If there is no object among the candidate objects whose corresponding first image feature matches the second preset category feature corresponding to any second preset category, then there is no object of the second preset category among the candidate objects.
4. The method according to claim 2 or 3, characterized in that: The number of target images corresponding to the preset category is multiple; The target image features corresponding to the preset category are obtained by: Respectively extracting image features of each target image corresponding to the preset category as second image features corresponding to each target image; According to the second image features corresponding to each of the target images, the target image features corresponding to the preset category are obtained; wherein the similarity between the target image features corresponding to the preset category and the second image features corresponding to each of the target images is greater than a first preset similarity threshold; The target text features corresponding to the preset category are obtained by: Acquire a target text having the same semantics as the text input for the preset category; Extracting first text features of the text input for the preset category, and respectively extracting text features of each of the target texts as second text features corresponding to each of the target texts; According to the first text feature and the second text feature corresponding to each of the target texts, the target text feature corresponding to the preset category is obtained; wherein the similarity between the target text feature corresponding to the preset category and the second text feature corresponding to each of the target texts is greater than a second preset similarity threshold, and the similarity between the target text feature and the first text feature is greater than the second preset similarity threshold.
5. The method according to claim 4, characterized in that The step of obtaining the target image features corresponding to the preset category according to the second image features corresponding to each of the target images includes: Calculating the average of the second image features corresponding to each of the target images as the target image feature corresponding to the preset category; The step of obtaining the target text feature corresponding to the preset category according to the first text feature and the second text feature corresponding to each of the target texts includes: The first text feature and the average of the second text features corresponding to each of the target texts are calculated as the target text feature corresponding to the preset category.
6. The method according to claim 4, characterized in that The obtaining the target image feature corresponding to the preset category according to the second image feature corresponding to each of the target images includes: Clustering the second image features corresponding to each of the target images to obtain a first cluster center feature as a target image feature corresponding to the preset category; The step of obtaining the target text feature corresponding to the preset category according to the first text feature and the second text feature corresponding to each of the target texts includes: The first text feature and the second text feature corresponding to each of the target texts are clustered to obtain a second cluster center feature as the target text feature corresponding to the preset category.
7. The method according to claim 1, characterized in that The method further comprises: If the ratio is not greater than a preset threshold, or if there is an object of the second preset category among the candidate objects, the image to be reviewed is sent to a manual end for review.
8. The method according to claim 1, characterized in that The performing object recognition on the image to be reviewed to obtain all candidate objects included in the image to be reviewed and the regions where each candidate object is located includes: For each pixel in the image to be reviewed, identifying and obtaining a category label and an instance label of the pixel; For a plurality of pixels having the same category label and the same instance label, an object formed by the plurality of pixels is determined as a candidate object, and an area formed by the plurality of pixels is determined as an area where the candidate object is located.
9. An image review device, characterized in that: The device comprises: An image acquisition module, used to acquire images to be reviewed; A candidate object recognition module, used to perform object recognition on the image to be reviewed, and obtain all candidate objects included in the image to be reviewed and the regions where each candidate object is located; A target object determination module, used to determine all objects of a first preset category from the candidate objects as target objects; wherein the first preset category includes objects that are all non-violating objects; A ratio calculation module, used for calculating the ratio of the total area of the regions where all the target objects are located to the area of the image to be reviewed; The review and determination module is used to determine that the image to be reviewed has passed the review if the ratio is greater than a preset threshold and there is no object of the second preset category among the candidate objects; wherein the second preset category is a category in which all the objects included belong to illegal objects.
10. The device according to claim 9, characterized in that The device also includes: A first target image acquisition module, configured to acquire, for each first preset category, an image including an object of the first preset category as a target image corresponding to the first preset category; A first image feature extraction module, configured to extract, for each of the first preset categories, image features of a target image corresponding to the first preset category, to obtain target image features corresponding to the first preset category; A first text feature extraction module, configured to extract text features of text input for each of the first preset categories, and obtain target text features corresponding to the first preset categories; The step of determining all objects of the first preset category from the candidate objects as target objects includes: For each candidate object, extracting an image feature of the candidate object from the image to be reviewed as a first image feature corresponding to the candidate object; From the candidate objects, determine the objects whose corresponding first image features all match the first preset category features corresponding to any of the first preset categories as target objects; wherein the first preset category features corresponding to the first preset category include: target image features corresponding to the first preset category, and target text features corresponding to the first preset category; The device also includes: A second target image acquisition module, configured to acquire, for each second preset category, an image including an object of the second preset category as a target image corresponding to the second preset category; A second image feature extraction module, configured to extract, for each of the second preset categories, image features of a target image corresponding to the second preset category, to obtain target image features corresponding to the second preset category; A second text feature extraction module, for extracting text features of text input for each of the second preset categories, to obtain target text features corresponding to the second preset categories; Determine whether there is an object of the second preset category among the candidate objects by: For each candidate object, extracting an image feature of the candidate object from the image to be reviewed as a first image feature corresponding to the candidate object; If there is an object in each of the candidate objects whose corresponding first image feature matches the second preset category feature corresponding to any of the second preset categories, then there is an object of the second preset category in each of the candidate objects; wherein the second preset category feature corresponding to the second preset category includes: a target image feature corresponding to the second preset category, and a target text feature corresponding to the second preset category; If there is no object whose corresponding first image feature matches the second preset category feature corresponding to any second preset category among the candidate objects, then there is no object of the second preset category among the candidate objects; The number of target images corresponding to the preset category is multiple; The target image features corresponding to the preset category are obtained by: Respectively extracting image features of each target image corresponding to the preset category as second image features corresponding to each target image; According to the second image features corresponding to each of the target images, the target image features corresponding to the preset category are obtained; wherein the similarity between the target image features corresponding to the preset category and the second image features corresponding to each of the target images is greater than a first preset similarity threshold; The target text features corresponding to the preset category are obtained by: Acquire a target text having the same semantics as the text input for the preset category; Extracting first text features of the text input for the preset category, and respectively extracting text features of each of the target texts as second text features corresponding to each of the target texts; According to the first text feature and the second text feature corresponding to each of the target texts, a target text feature corresponding to the preset category is obtained; wherein the similarity between the target text feature corresponding to the preset category and the second text feature corresponding to each of the target texts is greater than a second preset similarity threshold, and the similarity between the target text feature and the first text feature is greater than the second preset similarity threshold; The obtaining the target image feature corresponding to the preset category according to the second image feature corresponding to each of the target images includes: Calculating the average of the second image features corresponding to each of the target images as the target image feature corresponding to the preset category; The step of obtaining the target text feature corresponding to the preset category according to the first text feature and the second text feature corresponding to each of the target texts includes: Calculating the average of the first text feature and the second text feature corresponding to each of the target texts as the target text feature corresponding to the preset category; The obtaining the target image feature corresponding to the preset category according to the second image feature corresponding to each of the target images includes: Clustering the second image features corresponding to each of the target images to obtain a first cluster center feature as a target image feature corresponding to the preset category; The step of obtaining the target text feature corresponding to the preset category according to the first text feature and the second text feature corresponding to each of the target texts includes: Clustering the first text features and the second text features corresponding to each of the target texts to obtain a second cluster center feature as a target text feature corresponding to the preset category; The device also includes: A manual review module, configured to send the image to be reviewed to a manual end for review if the ratio is not greater than a preset threshold, or if there is an object of a second preset category among the candidate objects; The performing object recognition on the image to be reviewed to obtain all candidate objects included in the image to be reviewed and the regions where each candidate object is located includes: For each pixel in the image to be reviewed, identifying and obtaining a category label and an instance label of the pixel; For a plurality of pixels having the same category label and the same instance label, an object formed by the plurality of pixels is determined as a candidate object, and an area formed by the plurality of pixels is determined as an area where the candidate object is located.
11. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 8 when executing a program stored in a memory.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 8 are implemented.