A method, apparatus, electronic device and storage medium for image recognition
By pre-identifying and classifying pictures of published content, the problem of repulsive images affecting user experience is solved, and the repulsive images are recognized and processed before distribution is realized, improving user experience and content quality.
Patent Information
- Application Number
- CN201910546203.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-03-26
AI Technical Summary
In the prior art, when a repulsive picture is used as a cover picture, it will seriously affect the user's reading experience, and existing remedial measures need to wait for user feedback before they can be implemented, resulting in the continued existence of adverse experiences.
By obtaining the attribute information of the content to be published and the mark of the picture, the non-human and manual image recognition terminals are used to classify and identify the pictures easily, identify the easily disgusting pictures in advance, and sort out their category and disgusting information into the attribute information before distribution.
Identifying and processing of repulsive images before distribution reduces the user's bad experience when reading, improves the quality and user experience of content release, and saves manual review costs.
Smart Images

Figure CN112115958B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition technology, and in particular to a picture recognition method, device, electronic device and storage medium. Background Art
[0002] With the rapid development of computer technology, informatization plays an increasingly important role in human social life. Users can browse some articles or videos through certain applications on terminal devices (mobile phones, IPADs, laptops). Pictures are a very important part of the content of articles or videos, especially the cover picture, which is directly related to the first impression of the user's content, and directly affects the user's click-through conversion effect and the user's perception of the content. The sources of content on the Internet are very wide and very many, including articles, pictures and videos, all of which have corresponding cover pictures. However, among the massive amount of pictures, there are a small number of offensive pictures such as teeth, acne, claustrophobia, snakes, blood, violence and horror. If these offensive pictures are used as cover pictures, it will seriously affect the user's experience and effect, and ultimately affect the number of clicks on the content.
[0003] In the prior art, the platform for publishing content publishes the content to the user end, and only after receiving the user's feedback information will it manually review the feedback pictures and take corresponding remedial measures. This remedial method makes the bad reading experience caused by offensive pictures for users continue. Summary of the invention
[0004] The embodiments of the present application provide a method, device, electronic device and storage medium for image recognition, which can provide a basis for reducing the bad reading experience caused by offensive images to users.
[0005] On the one hand, an embodiment of the present application provides a method for image recognition, the method comprising:
[0006] Acquire the content to be published and the attribute information of the content to be published sent by the content production end; the attribute information includes the first picture in the content to be published and the mark of the first picture;
[0007] Sending a first picture to the picture recognition end; the first picture carries a mark;
[0008] Receiving identification information sent by the image recognition terminal; the identification information includes a tag, a category of the first image, and offensive information of the first image, wherein the offensive information is used to identify whether the first image is an offensive image;
[0009] If the objectionable information matches the preset information, the category and the objectionable information are sorted into the attribute information according to the tag.
[0010] On the other hand, a picture recognition device is provided, the device comprising:
[0011] A receiving module, configured to obtain the content to be published sent by the content production end and the attribute information of the content to be published; the attribute information includes the first picture in the content to be published and the mark of the first picture;
[0012] A sending module, configured to send the first picture to the picture recognition end; the first picture carries a mark;
[0013] The receiving module is configured to receive the recognition information sent by the picture recognition end; the recognition information includes the mark, the category of the first picture, and the offensive information of the first picture, where the offensive information is used to identify whether the first picture is an offensive picture;
[0014] A processing module, configured to, if the offensive information matches the preset information, sort the category and the offensive information into the attribute information according to the mark.
[0015] On the other hand, an electronic device is provided, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the picture recognition method as described above.
[0016] On the other hand, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set, or an instruction set is stored, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the picture recognition method as described above.
[0017] On the other hand, an embodiment of the present application provides a picture recognition method, and the method includes:
[0018] Receiving a first picture sent by a server end, where the first picture carries a mark of the first picture;
[0019] Determining the category of the first picture and the offensive information of the first picture; the offensive information is used to identify whether the first picture is an offensive picture;
[0020] Sending recognition information to the server end; the recognition information includes the mark, the category of the first picture, and the offensive information of the first picture.
[0021] On the other hand, a picture recognition device is provided, and the device includes:
[0022] A receiving module, configured to receive a first picture sent by a server end, where the first picture carries a mark of the first picture;
[0023] A determining module, configured to determine the category of the first picture and the offensive information of the first picture; the offensive information is used to identify whether the first picture is an offensive picture;
[0024] A sending module for sending recognition information to the server side; the recognition information includes a tag, the category of the first picture, and the offensive information of the first picture.
[0025] On the other hand, an electronic device is provided, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the picture recognition method as described above.
[0026] On the other hand, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set, or an instruction set is stored, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the picture recognition method as described above.
[0027] The picture recognition method, device, and storage medium provided by the embodiments of the present application have the following technical effects:
[0028] Obtain the content to be published sent by the content production side and the attribute information of the content to be published. The attribute information includes the first picture in the content to be published and the tag of the first picture, and send the first picture to the picture recognition side; the first picture carries the tag. Receive the recognition information sent by the picture recognition side; the recognition information includes the tag, the category of the first picture, and the offensive information of the first picture. Among them, the offensive information is used to identify whether the first picture is an offensive picture. If the offensive information matches the preset information, sort the category and the offensive information into the attribute information according to the tag. Since the first picture has been recognized before distribution, the recognition information will be used as the distribution condition in the subsequent distribution process, laying a foundation for reducing the bad reading experience caused by offensive pictures to users. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0030] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application;
[0031] Figure 2 is a schematic diagram of an application environment provided by an embodiment of the present application;
[0032] Figure 3It is a schematic flowchart of a picture recognition method provided by an embodiment of the present application;
[0033] Figure 4 It is a schematic flowchart of a method for determining the category of a first picture and the objectionable information of the first picture provided by an embodiment of the present application;
[0034] Figure 5 It is a schematic flowchart of a method for determining the category of a first picture and the objectionable information of the first picture provided by an embodiment of the present application;
[0035] Figure 6 It is a schematic structural diagram of a second non-artificial picture recognition terminal provided by an embodiment of the present application;
[0036] Figure 7 It is a schematic structural diagram of a fourth recognition model provided by an embodiment of the present application;
[0037] Figure 8 It is a schematic structural diagram of a picture recognition device provided by an embodiment of the present application;
[0038] Figure 9 It is a schematic structural diagram of a picture recognition device provided by an embodiment of the present application;
[0039] Figure 10 It is a hardware structure block diagram of a server for a picture recognition method provided by an embodiment of the present application. Detailed implementation manners
[0040] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0041] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0042] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an application environment provided by an embodiment of the present application, including a server side 101 and a picture recognition side 102. Among them, the server side 101 includes a database. The server side 101 is used to obtain the content to be published and the attribute information of the content to be published sent by the content production side, and store the content to be published and the attribute information in the database. Among them, the attribute information includes the first picture of the content to be published and the mark of the first picture. The server side 101 sends the first picture to the picture recognition side 102, and the first picture carries the mark. The picture recognition side 102 determines the category of the first picture and the offensive information of the first picture according to the first picture. The offensive information is used to identify whether the first picture is an offensive picture. The server side 101 receives the recognition information sent by the picture recognition side 102, and the recognition information includes the mark, the category of the first picture and the offensive information of the first picture. If the offensive information matches the preset information, the server side 101 sorts the category and the offensive information into the attribute information of the content to be published where the first picture is located according to the mark.
[0043] Optionally, the picture recognition side 102 can be a computer terminal, a server or a similar device. The data between the server side 101 and the picture recognition side 102 can be transmitted through a wired link or a wireless link. The selection of the communication link type can be determined according to the actual application situation and application environment.
[0044] In order to complete a more accurate picture recognition process, in an optional embodiment, the above picture recognition side can be divided into a non-artificial picture recognition side and an artificial picture recognition side. Please refer to Figure 2 , Figure 2 which is a schematic diagram of an application environment provided by an embodiment of the present application, including a server side 201, a first non-artificial picture recognition side 202, a second non-artificial picture recognition side 203, a first artificial picture recognition side 204 and a second artificial picture recognition side 205. Among them, the server side 201 includes a database. The first non-artificial picture recognition side 202 and the second non-artificial picture recognition side 203 belong to the non-artificial picture recognition side and can determine the category of the first picture and the offensive information based on the picture. The first artificial picture recognition side 204 and the second artificial picture recognition side 205 belong to the artificial picture recognition side, and the offensive information of the first picture needs to be determined by the reviewer based on the first picture.
[0045] In an embodiment of the present application, the server 201 is configured to obtain the content to be published sent by the content production end and the attribute information of the content to be published, and store the content to be published and the attribute information in a database. The attribute information includes a first picture of the content to be published and a label of the first picture. The server 201 sends the first picture to the first non-artificial picture recognition end 202 and the second non-artificial picture recognition end 203 respectively, and the first picture carries the label. The first non-artificial picture recognition end 202 and the second non-artificial picture recognition end 203 determine the category of the first picture and attempt to determine the offensive information of the first picture. If the first non-artificial picture recognition end 202 or the second non-artificial picture recognition end 203 can determine the offensive information of the first picture, the recognition information is directly sent to the server 201.
[0046] If the first non-artificial picture recognition end 202 cannot determine the offensive information of the first picture, the first picture is sent to the first artificial picture recognition end 204 so that the first artificial picture recognition end 204 can determine the offensive information of the first picture. If the second non-artificial picture recognition end 203 cannot determine the offensive information of the first picture, the first picture is sent to the second artificial picture recognition end 205 so that the second artificial picture recognition end 205 can determine the offensive information of the first picture. The server 201 receives the recognition information sent by one or more of the first non-artificial picture recognition end 202, the second non-artificial picture recognition end 203, the first artificial picture recognition end 204, and the second artificial picture recognition end 205. The recognition information includes the label, the category of the first picture, and the offensive information of the first picture. If the offensive information matches the preset information, the server 201 sorts the category and offensive information of the first picture into the attribute information of the content to be published where the first picture is located according to the label.
[0047] Optionally, the first non-artificial picture recognition end 202, the second non-artificial picture recognition end 203, the first artificial picture recognition end 204, and the second artificial picture recognition end 205 can be computer terminals, servers, or similar devices. The data between the devices in the above platform (including between the server 201 and the first non-artificial picture recognition end 202, between the server 201 and the second non-artificial picture recognition end 203, between the server 201 and the first artificial picture recognition end 204, between the server 201 and the second artificial picture recognition end 205, between the first non-artificial picture recognition end 202 and the first artificial picture recognition end 204, and between the second non-artificial picture recognition end 203 and the second artificial picture recognition end 205) can be transmitted through a wired link, can also be transmitted through a wireless link, or can be transmitted in a combination of a wired link and a wireless link. The selection of the communication link type can be determined according to the actual application situation and application environment.
[0048] The following introduces a specific embodiment of a picture recognition method of the present application. Figure 3 FIG. Figure 3 is a schematic flowchart of a picture recognition method provided by an embodiment of the present application. This specification provides method operation steps such as in the embodiment or flowchart, but based on routine or non-creative labor, more or fewer operation steps may be included. The step order listed in the embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or server product executes, it may execute in the order shown in the embodiment or the drawing or execute in parallel (for example, in an environment of parallel processors or multi-threaded processing). Specifically, as Figure 3 shown, the method may include:
[0049] S301: The server side obtains the content to be published sent by the content production side and the attribute information of the content to be published. The attribute information includes the first picture in the content to be published and the mark of the first picture.
[0050] In the embodiment of the present application, the content production side may be an application program on a terminal in modes such as user-generated content (UGC), professional-generated content (PGC), professional user-generated content (PUGC), and multi-channel network (MCN). Among them, UGC refers to a mode in which users upload the content to be published to a website for other users to view and download. PGC refers to a mode in which relatively professional users or experts upload the content to be published to a website. Compared with UGC, PGC is more professional in classification and the content quality can also be guaranteed to a certain extent. PUGC refers to a content production mode that combines UGC and PGC. MCN is a mode that combines PGC content so that the content to be published can be continuously output.
[0051] In the embodiment of the present application, the content to be published may be in formats such as text, picture, video, audio, picture plus text, audio plus text, etc. Among them, some content to be published (such as picture and text) contains one or more first pictures, and the publisher can select one of the first pictures as the cover picture; some formats may not contain pictures (such as video), but in order to attract viewers to watch or listen to the content to be published, the publisher can select a first picture as the cover picture. Optionally, the publishers in each mode included in the above content production side can connect to the application programming interface through the terminal or the backend interface and send the content to be published to the server side.
[0052] In the embodiments of the present application, the attribute information of the content to be published is used to describe the content to be published, including the title of the content to be published, the label of the title of the content to be published, the identity information of the publisher, the abstract of the published content, the first picture, the label of the first picture, the publishing time, the memory required for the content to be published, the format of the content to be published, the original information, etc. The original information is used to identify whether the content to be published is original by the publisher. The label of the first picture may be the URL address of the first picture or an indication information, and the indication information is used to indicate the content to be published corresponding to the first picture, so that after the server obtains the category or the offensive information of the first picture subsequently, it can be sorted into the attribute information of the content to be published where the first picture is located according to the label. The label of the title of the content to be published is used to indicate the content to be published.
[0053] Optionally, the attribute information further includes the overall type of the content to be published. In an alternative embodiment, the server can call a classification algorithm to classify the content to be published. For example, the content to be published uploaded can be an article, and the classification result can be technology - Internet - Internet finance - micro - loan, etc. In another alternative embodiment, the overall classification can be performed manually. For example, the content to be published uploaded is a game video, so the manual classification result can be game - mobile game - multiplayer competitive game - Game A.
[0054] In the embodiments of the present application, there may be an intermediate device between the content production end and the server end, such as an uplink and downlink interface server. After establishing a communication connection with the content production end, the intermediate device is used to receive the content to be published and the attribute information of the content to be published uploaded by the content production end, and send the content to be published and the attribute information to the server end, so that the server end can store and process them.
[0055] In an alternative embodiment, the attribute information of the content to be published can be stored in a storage area of the server end, and the content to be published can be stored in another storage area of the server end. There is a mapping table between the content to be published and the attribute information of the content to be published, which is convenient for the database to determine the location of the other party in the storage area according to the known party. Further, the content to be published can be stored in different storage areas of the server end according to the format of the content to be published.
[0056] S303: The server end sends the first picture to the picture recognition end, and the first picture carries a label.
[0057] In the embodiments of the present application, the label of the first picture is used to identify the content to be published to which the first picture belongs, so that after the server end receives the category and the offensive information of the first picture sent by the picture recognition end subsequently, it can sort the category and the offensive information of the first picture into the corresponding attribute information according to the label.
[0058] In the embodiments of the present application, the image recognition terminal can be divided into a non-artificial image recognition terminal and an artificial image recognition terminal. In this way, the accuracy of image recognition can be improved. The recognition processes based on the non-artificial image recognition terminal and the artificial image recognition terminal will be described below.
[0059] S305: The image recognition terminal determines the category of the first image and the offensive information of the first image, where the offensive information is used to identify whether the first image is an offensive image.
[0060] As mentioned above, the image recognition terminal can include a non-artificial image recognition terminal and an artificial image recognition terminal. There are many methods for the embodiments of the present application to determine the category of the first image and the offensive information of the first image. Four optional implementation manners will be introduced below based on the non-artificial image recognition terminal and the artificial image recognition terminal.
[0061] In an optional implementation manner, the embodiments of the present application can determine the category of the first image and the offensive information of the first image through the first non-artificial image recognition terminal and / or the first artificial image recognition terminal. Figure 4 FIG. is a schematic flowchart of a method for determining the category of the first image and the offensive information of the first image provided by the embodiments of the present application. The method may include:
[0062] S401: The first non-artificial image recognition terminal determines the first feature information of the first image according to the first image sent by the server terminal.
[0063] In the embodiments of the present application, the first non-artificial image recognition terminal may include a first recognition model, and the first recognition model can be used to determine the first feature information of the first image. Specifically, the first non-artificial image recognition terminal takes the first image as the input of the first recognition model and obtains the first feature information of the first image from the output of the first recognition model.
[0064] In a structure of an optional first recognition model, the first recognition model may include a plurality of convolutional layers and a plurality of fully connected layers. After the first image passes through the plurality of convolutional layers and the plurality of fully connected layers, a first feature information with a preset dimension can be obtained.
[0065] For example, the structure of the first recognition model is successively: the first convolutional layer, the second convolutional layer, the third convolutional layer, and the first fully connected layer.
[0066] Optionally, before the first non-artificial image recognition terminal inputs the first image into the first recognition model, the size of the first image can be changed to obtain a first image with a preset size. For example, after size change, a first image of 640*360*3 (i.e., 640 pixels in width, 360 pixels in height, and 3 channels) can be obtained as a first image of 227*227*3, which can also be called a feature plane of 227*227*3.
[0067] The first convolutional layer may include a first sub-convolutional layer and a first pooling layer. Among them, the first sub-convolutional layer includes 96 convolutional kernels of 11*11, and the pooling window of the first pooling layer is 3*3. The first convolutional layer receives the feature plane of 227*227*3, uses the 96 convolutional kernels of 11*11 in the first sub-convolutional layer to perform a convolution operation on the feature plane with a stride of 4 pixels, and performs max-pooling or average-pooling processing with a stride of 2 pixels through the 3*3 pooling window to obtain a feature plane of 27*27*96. Optionally, before performing the convolution operation on the feature plane of 227*227*3, edge pixel padding can also be performed on the feature plane.
[0068] The second convolutional layer may include a second sub-convolutional layer and a second pooling layer. Among them, the second sub-convolutional layer includes 256 convolutional kernels of 3*3, and the pooling window of the second pooling layer is 2*2. The second convolutional layer receives the feature plane of 27*27*96, uses the 256 convolutional kernels of 3*3 in the second sub-convolutional layer to perform a convolution operation on the feature plane with a stride of 1 pixel, and performs max-pooling or average-pooling processing with a stride of 2 pixels through the 2*2 pooling window to obtain a feature plane of 13*13*256. Optionally, before performing the convolution operation on the feature plane of 27*27*96, edge pixel padding can also be performed on the feature plane.
[0069] The third convolutional layer includes 384 convolutional kernels of 5*5. The third convolutional layer receives the feature plane of 13*13*256, uses the 384 convolutional kernels of 5*5 to perform a convolution operation on the feature plane with a stride of 1 pixel, and obtains a feature plane of 13*13*384. Optionally, before performing the convolution operation on the feature plane of 13*13*256, edge pixel padding can also be performed on the feature plane.
[0070] The first fully connected layer receives the feature plane of 13*13*384, and after processing, outputs a feature plane of 1*1*1024. The feature plane of 1*1*1024 is used to represent the first feature information of the first picture, that is, the 1024-dimensional feature vector after the first picture is vectorized.
[0071] The structure of the first recognition model described above is only an optional embodiment of this application, and the structure of the first recognition model can be determined according to actual needs.
[0072] S403: The first non-artificial picture recognition end determines a plurality of distance values according to the second feature information and the first feature information of each second picture in the plurality of second pictures; wherein, the second picture is an offensive picture obtained after receiving user feedback and passing the review.
[0073] In the embodiments of the present application, the second picture may be a repulsive picture obtained after receiving user feedback and passing the review. The repulsive picture is a picture of certain preset categories, which may include pictures that are likely to cause disgust to users, such as pimples, pregnant women giving birth, cysts, teeth, trypophobia, QR codes, snakes, lizards, etc. The categories included in the preset categories can be determined according to the actual situation. In order to distinguish different types of repulsive pictures, a first label can be set to represent the category of the repulsive picture. For example, the character "11" is used to represent pimples, the character "12" is used to represent pregnant women giving birth, the character "13" is used to represent cysts, the character "14" is used to represent teeth...
[0074] In an alternative embodiment, after the previous content to be published is published, when a user browses it, the user may be disgusted with some of the pictures and provide feedback information to the platform that provides the content to be published. The feedback information includes the identity information of the publisher of the content to be published where the picture is located and the URL address of the picture. The platform may include a third artificial picture recognition terminal. The platform distributes the feedback information to the third artificial picture recognition terminal. The third artificial picture recognition terminal obtains the picture according to the URL address of the picture and displays the picture on the display screen of the third artificial picture recognition terminal. The staff of the platform can determine whether the picture is a repulsive picture. If so, the recognition result input by the staff is that the picture is a repulsive picture and the category of the picture. The third artificial picture recognition terminal stores the picture and the category of the picture in a specific storage area. If not, the recognition result input by the staff is that the picture is not a repulsive picture, and the third artificial picture recognition terminal can delete the picture. In this way, each picture in the specific storage area is a repulsive picture obtained after receiving user feedback and passing the review, that is, the second picture in the above text. In an alternative embodiment, the third artificial picture recognition terminal reports the identity information of the publisher to the server terminal, so as to determine the reported times of each publisher based on the identity information of the publisher in the future.
[0075] In the embodiments of the present application, the first non-artificial picture recognition terminal may determine the second feature information of the second picture based on the method for determining the first feature information of the first picture in the above text. That is to say, the second feature information may also be a 1024-dimensional feature vector after the second picture is vectorized. Optionally, the specific storage area not only stores the second picture and the category of the second picture, but may also store the second feature information of the second picture. Or, the specific storage area stores the URL address of the second picture, the second feature information of the second picture, and the category of the second picture. Or, the specific storage area stores the second feature information of the second picture and the category of the second picture. As time goes by, the specific storage area can store more and more second pictures. The first non-artificial picture recognition terminal can perform vectorization processing on the newly added second pictures according to a preset time period to obtain new second feature vectors.
[0076] In an implementation method for determining a distance value based on second feature information and first feature information, the first non-artificial image recognition terminal may include a second recognition model. Specifically, the first non-artificial image recognition terminal uses the second feature information and the first feature information as the input of the second recognition model, and obtains the distance value from the output of the second recognition model. Optionally, the second recognition model includes a second fully connected layer, a third fully connected layer, and a contrast loss function layer.
[0077] Continuing with the above example of the 1*1*1024 feature plane:
[0078] The first non-artificial image recognition terminal inputs the first feature information of the first image, that is, the 1*1*1024 feature plane, into the second fully connected layer. After processing, it outputs a 1*1*256 feature plane. The 1*1*256 feature plane is input into the third fully connected layer, and after processing, it outputs a 1*1*2 feature plane, that is, a 2D feature vector. Similarly, the 2D feature vector corresponding to the second image can be obtained.
[0079] Optionally, the first non-artificial image recognition terminal inputs the 2D feature vector corresponding to the first image and the 2D feature vector corresponding to the second image into the contrast loss function layer to determine the distance value between the two images. In an alternative implementation, the Euclidean distance value between the first images can be determined as the distance value. Assume that the 2D feature vector of the first image is (x1, y1), and the 2D feature vector of the second image is (x2, y2), then the distance value can be expressed by formula (1):
[0080] ……Formula (1)
[0081] According to the above method for determining the distance value, the distance values between each second feature information and the first feature information in a specific storage area can be determined.
[0082] S405: The first non-artificial image recognition terminal determines the distance value with the smallest numerical value among the multiple distance values as the first distance value.
[0083] In the embodiment of the present application, the smaller the numerical value of the distance value, the more similar the second image corresponding to the distance value and the first image. Therefore, the distance value with the smallest numerical value among the multiple distance values can be determined as the first distance value.
[0084] S407: The first non-artificial image recognition terminal determines the category of the second image corresponding to the first distance value as the category of the first image.
[0085] In the embodiment of the present application, the first non-artificial image recognition terminal can determine the category of the most similar second image as the category of the first image. Assume that the type of the second image is acne, and the type of the first image is acne.
[0086] In the embodiments of the present application, the first distance value may be a very large value. For example, if the first picture is not a picture of any of the above categories but a landscape picture, the first distance value can still be obtained based on all the second pictures. However, the value of the first distance value may be extremely large (e.g., 200). To reduce the unnecessary computing work of the first non-artificial picture recognition end, in an optional implementation, the first distance value and the third distance value can be compared before S407. The third distance value is a very large value. If the first distance value exceeds the third distance value, it indicates that the first picture is not a picture of the preset category. Otherwise, the category of the second picture corresponding to the first distance value is determined as the category of the first picture.
[0087] S409: The first non-artificial picture recognition end determines whether the first distance value is less than a preset second distance value. If so, go to S411; if not, go to S415.
[0088] In the embodiments of the present application, the second distance value can be set according to evaluation or empirical values (e.g., 50).
[0089] S411: The first non-artificial picture recognition end determines that the first picture is an offensive picture.
[0090] S413: The first non-artificial picture recognition end determines the offensive information of the first picture based on the offensive picture. End this process.
[0091] In the embodiments of the present application, the offensive information can be represented by a specific string or character. For example, any one of "0", "true", and "yes" can be used as the offensive information to indicate that the first picture is an offensive picture.
[0092] S415: The first non-artificial picture recognition end sends the first picture to the first artificial picture recognition end, and the first picture carries a mark and a category.
[0093] S417: After receiving the first picture, the first artificial picture recognition end determines whether the first picture is an offensive picture according to the received instruction; if so, go to S419; if not, go to S421.
[0094] Optionally, if the first distance value is greater than or equal to the second distance value, after receiving the first picture, the first artificial picture recognition end displays the first picture on the display screen of the first artificial picture recognition end. If the staff determines that it is an offensive picture, the first artificial picture recognition end receives the first instruction, which is used to indicate that the first picture is an offensive picture. If the staff determines that it is not an offensive picture, the first artificial picture recognition end receives the second instruction, which is used to indicate that the first picture is not an offensive picture.
[0095] Optionally, if the first distance value is greater than or equal to the second distance value and less than or equal to the third distance value, after receiving the first picture, the first artificial picture recognition terminal displays the first picture on the display screen of the first artificial picture recognition terminal. If the staff member determines that the picture is an offensive picture, the first artificial picture recognition terminal receives a first instruction for indicating that the first picture is an offensive picture. If the staff member determines that the picture is not an offensive picture, the first artificial picture recognition terminal receives a second instruction for indicating that the first picture is not an offensive picture.
[0096] S419: The first artificial picture recognition terminal determines the offensive information of the first picture based on the offensive picture, and ends this process.
[0097] Based on the above representation of the offensive information, the first artificial picture recognition terminal can determine that the offensive information of the first picture is any one of "0", "true", and "yes".
[0098] S421: The first artificial picture recognition terminal determines the offensive information of the first picture, and this offensive information indicates that the first picture is not an offensive picture. End this process.
[0099] Optionally, the first artificial picture recognition terminal can determine that the offensive information of the first picture is any one of "1", "false", and "no".
[0100] In the embodiments of the present application, the above-mentioned first artificial picture recognition terminal and the third artificial picture recognition terminal can be one artificial picture recognition terminal, or different artificial picture recognition terminals. The number of artificial picture recognition terminals can be adjusted based on actual services.
[0101] In another optional implementation manner, in the embodiments of the present application, the category of the first picture and the offensive information of the first picture can be determined by the second non-artificial picture recognition terminal and / or the second artificial picture recognition terminal. Figure 5 It is a schematic flowchart of a method for determining the category of the first picture and the offensive information of the first picture provided by the embodiments of the present application. The method may include:
[0102] S501: The second non-artificial picture recognition terminal determines the category of the first picture according to the first picture sent by the server.
[0103] In the embodiments of the present application, the second non-artificial picture recognition terminal may include a third recognition model, and the third recognition model can be used to determine the category of the first picture. Specifically, the second non-artificial picture recognition terminal uses the first picture as the input of the third recognition model, and obtains the category of the first picture from the output of the third recognition model.
[0104] In the structure of an optional third recognition model, the third recognition model may include six convolutional layers, two fully connected layers, and a fully connected softmax classification layer. After the first picture is input into the third recognition model, it is sequentially subjected to convolutional processing through six convolutional layers, and then sequentially input into two fully connected layers, and the category of the first picture is output from the last fully connected softmax classification layer. The structure of the third recognition model described above is only an optional embodiment of the present application, and the structure of the third recognition model can be determined according to actual needs.
[0105] In an optional implementation manner of determining the third recognition model, pictures of preset categories (such as pictures of acne, pregnant women giving birth, cysts, teeth, trypophobia, QR codes, snakes, lizards, etc. that are likely to cause disgust to users) and pictures of non-preset categories can be collected from an online picture library, and a first label is assigned to each collected picture to indicate what category the picture belongs to. For example, the character "11" represents acne, the character "12" represents pregnant women giving birth, the character "13" represents cysts, the character "14" represents teeth... The character "00" represents the type of pictures of non-preset categories. Each picture is used as the input of the third recognition model, and the category of each picture is used as the output of the third recognition model for model training to obtain the third recognition model.
[0106] S503: The second non-artificial picture recognition end determines whether the category of the first picture is a preset category; if so, go to S505; if not, determine that the first picture is a non-disgusting picture, and the type of the first picture is other.
[0107] S505: The second non-artificial picture recognition end determines the disgust degree value of the first picture.
[0108] In the embodiment of the present application, the second non-artificial picture recognition end may include a fourth recognition model. The fourth recognition model may include several cascaded structures, a connection separation layer, a connection layer, a fully connected layer, a convolutional layer, and a fully connected softmax classification layer, and can be used to determine the disgust degree value of the first picture. Specifically, the second non-artificial picture recognition end uses the first picture as the input of the fourth recognition model and obtains the disgust degree value of the first picture from the output of the fourth recognition model.
[0109] Figure 6It is a schematic structural diagram of a second non-artificial picture recognition terminal provided by an embodiment of the present application, including a third recognition model and a fourth recognition model. Among them, the third recognition model may include six convolutional layers 621, two fully connected layers 622 and 623, and a fully connected softmax classification layer 624. The first picture can be input into the third recognition model, and convolutional processing is performed successively through six convolutional layers 621, and then input into two fully connected layers 622 and 623 in sequence, and the category of the first picture is output from the last fully connected softmax classification layer 624. If the category of the first picture is a preset category, the first picture is continued to be input into the fourth recognition model, otherwise it is input into the above function Function.
[0110] As Figure 6 shown, the fourth recognition model may include multiple cascade structures 601, 602, and 603, multiple connection separation layers 605, 606, and 607, a connection layer 608, fully connected layers 609 and 610, a convolutional layer 604, and a fully connected softmax classification layer 611. Since the category of the first picture is a tooth, which belongs to a preset category, the first picture is continued to be input into the fourth recognition model for recognition. Each layer of the fourth recognition model processes the first picture in sequence, and the data output from the fully connected softmax classification layer 611 is input into the function Function to obtain the disgust degree value of the first picture, that is, the disgust degree value of the tooth picture.
[0111] Optionally, if the category of the first picture output by the third recognition model is not a preset category, the data output by the third recognition model is input into the function Function, and the result data indicating that the first picture is a non-disgusting picture and the type of the first picture is other can be output.
[0112] Among them, the convolutional layer 604 can perform a convolution operation on the feature plane corresponding to the first picture, and output after normalization processing and ReLU non-linear processing. Each cascaded structure 601, 602, and 603 can have multiple branches. The cascaded structure 603 represents the last cascaded structure in the fourth recognition model, and the number of cascaded structures in this model can be determined according to the actual recognition situation. Each branch of the cascaded structure receives the feature plane from the previous layer, and outputs the feature plane after pooling processing, convolution operation, normalization processing, and ReLU non-linear processing. The connection and separation layers 605, 606, and 607 receive the feature planes output by each branch of the previous layer (cascaded mechanism), perform connection and then separation processing, and input to the next cascaded structure. The connection layer 608 is used to receive the feature planes output by each branch of the previous cascaded structure, and output after connection. The fully connected layers 609 and 610 receive the feature plane from the previous layer, and output after fully connected processing. The final fully connected softmax classification layer outputs the disgust degree value of the first picture based on the input feature plane.
[0113] Figure 7 It is a schematic structural diagram of a fourth recognition model provided by an embodiment of the present application, and processes the input first picture in the direction indicated by the arrow. Among them, rectangles are used to represent convolutional layers, ellipses are used to represent average pooling layers, triangles are used to represent max pooling layers, diamonds are used to represent connection layers, pentagons are used to represent dropout layers, circles are used to represent fully connected layers, and rings are used to represent softmax classification layers.
[0114] In an alternative embodiment, the first picture input to the fourth recognition model passes through 2 convolutional layers, 1 max pooling layer, 2 convolutional layers, 1 max pooling layer, and 8 cascaded structures with 8 connection layers interspersed in sequence, and then the obtained data is input to the first output branch. It passes through 1 average pooling layer, 2 convolutional layers, 1 fully connected layer, and 1 softmax classification layer included in the first output branch in sequence, and outputs the result corresponding to the first picture from the softmax classification layer. In the embodiment of the present application, this result is the disgust degree value. The different numbers of branches included in the above 8 cascaded structures, and the different numbers of layers with different functions included in each branch can be determined according to the actual recognition situation.
[0115] In another alternative embodiment, the first picture input into the fourth recognition model sequentially passes through 2 convolutional layers, 1 max pooling layer, 2 convolutional layers, 1 max pooling layer, and 8 cascaded structures with 8 connection layers interspersed, and then continues to pass through 3 cascaded structures with 3 connection layers. The obtained data is then input into the second output branch. It sequentially passes through 1 average pooling layer, 1 dropout layer, 1 fully connected layer, and 1 softmax classification layer included in the second output branch, and the result corresponding to the first picture is output from the softmax classification layer. In the embodiments of the present application, this result is the disgust degree value. The different numbers of branches included in the subsequent 3 cascaded structures, as well as the different numbers of layers with different functions included in each branch, can be determined according to the actual recognition situation. In addition, there are also some branches in the last 2 cascaded structures that are cascaded structures. Thus, the fourth recognition model with more cascaded structures can perform more delicate and accurate recognition results on the first picture.
[0116] In an alternative implementation manner for determining the fourth recognition model, pictures of preset categories (such as pictures of pimples, pregnant women giving birth, cysts, teeth, trypophobia, QR codes, snakes, lizards, etc. that are likely to cause disgust to users) can be collected from an online picture library, and a second label is assigned to each collected picture to represent the disgust degree value of the picture. The second label can be any value in the interval (0, 1]. Each picture is used as the input of the fourth recognition model, and the disgust degree value of each picture is used as the output of the fourth recognition model for model training to obtain the fourth recognition model.
[0117] S507: The second non-artificial picture recognition end determines whether the disgust degree value is greater than a preset degree value; if so, go to S509; if not, go to S513.
[0118] In the embodiments of the present application, the preset degree value (such as 0.7) can be set according to evaluation or empirical values.
[0119] In the embodiments of the present application, the greater the disgust degree value, the more likely the first picture is to cause disgust to the user, and the preset degree value can be set according to evaluation or empirical values.
[0120] S509: The second non-artificial picture recognition end determines that the first picture is a disgusting picture.
[0121] S511: The second non-artificial picture recognition end determines the disgust information of the first picture based on the disgusting picture. End this process.
[0122] In the embodiments of the present application, the disgust information can be represented by a specific string or character. For example, any one of "0", "true", and "yes" can be used as the disgust information to indicate that the first picture is a disgusting picture.
[0123] In the embodiments of the present application, the preset degree value may be a very small value. For example, if the first picture is a picture of neat and white teeth, it can still be determined based on the third recognition model that the first picture belongs to the picture of the preset category, but the degree of disgust value may be very low (such as 0.01). In order to reduce unnecessary computing work of the second non-artificial picture recognition end, in an optional implementation manner, the degree of disgust value is compared with the second preset degree value, and the second preset degree value is a relatively small value (such as 0.05). At this time, the above-mentioned preset degree value of 0.7 can be regarded as the first preset degree value. If the degree of disgust value is less than the second preset degree value, it indicates that the first picture is a non-disgusting picture. Otherwise, the first picture may or may not be a disgusting picture. Subsequently, continue with S513, that is, send the first picture to the second artificial picture recognition end, and the first picture carries a mark and a category.
[0124] S513: The second non-artificial picture recognition end sends the first picture to the second artificial picture recognition end, and the first picture carries a mark and a category.
[0125] S515: After receiving the first picture, the second artificial picture recognition end determines whether the first picture is a disgusting picture according to the received instruction; if so, go to S517; if not, go to S519.
[0126] Optionally, if the degree of disgust value is less than or equal to the preset degree value, after receiving the first picture, the second artificial picture recognition end displays the first picture on the display screen of the second artificial picture recognition end. If the staff determines that it is a disgusting picture, the second artificial picture recognition end receives the third instruction, and the third instruction is used to indicate that the first picture is a disgusting picture. If the staff determines that it is a non-disgusting picture, the second artificial picture recognition end receives the fourth instruction, and the fourth instruction is used to indicate that the first picture is a non-disgusting picture.
[0127] Optionally, if the degree of disgust value is less than or equal to the first preset degree value and greater than or equal to the second preset degree value, after receiving the first picture, the second artificial picture recognition end displays the first picture on the display screen of the second artificial picture recognition end. If the staff determines that it is a disgusting picture, the second artificial picture recognition end receives the third instruction, and the third instruction is used to indicate that the first picture is a disgusting picture. If the staff determines that it is a non-disgusting picture, the second artificial picture recognition end receives the fourth instruction, and the fourth instruction is used to indicate that the first picture is a non-disgusting picture.
[0128] S517: The second artificial picture recognition end determines the disgust information of the first picture according to the disgusting picture. End this process.
[0129] Based on the above representation of the offensive information, the second artificial image recognition terminal can determine that the offensive information of the first image is any one of "0", "true", and "yes".
[0130] S519: The second artificial image recognition terminal determines the offensive information of the first image, and this offensive information indicates that the first image is a non-offensive image. End this process.
[0131] Optionally, the second artificial image recognition terminal can determine that the offensive information of the first image is any one of "1", "false", and "no".
[0132] In another alternative implementation manner, the embodiments of the present application can determine the category of the first image and the offensive information of the first image through the first non-artificial image recognition terminal, the second non-artificial image recognition terminal, and / or the first artificial image recognition terminal, the second artificial image recognition terminal. That is, the combination of the first implementation manner and the second implementation manner, that is Figure 4 the method for determining the category of the first image and the offensive information of the first image shown and Figure 5 the combination of the method for determining the category of the first image and the offensive information of the first image shown. In this implementation manner, both the category of the first image and the offensive information of the first image are determined based on the second image fed back by the user, so that there is a practical application sample for the recognition of the first image, and there is a close connection with the business, making it more targeted. And it can also be based on the second non-artificial image recognition terminal to perform image recognition based on a large amount of training data. The combination of the two can improve the accuracy of the recognition result.
[0133] At the same time, based on the above-mentioned implementation method of the embodiment, the processing speed of offensive images can be improved, and it is not completely necessary for manual review, saving the labor cost required for review.
[0134] Optionally, the first artificial image recognition terminal, the second artificial image recognition terminal, and the third artificial image recognition terminal in the above three alternative implementation schemes can be the same artificial image recognition terminal, or different artificial image recognition terminals. The number of artificial image recognition terminals can be adjusted based on the actual business.
[0135] In another alternative implementation manner, the image recognition terminal is an artificial image recognition terminal. The artificial image recognition terminal receives the first image sent by the server, displays the first image on the display screen of the artificial image recognition terminal, and determines the category and offensive information of the first image according to the received instruction. And feedback the category and offensive information of the first image to the server side.
[0136] S307: The server receives the recognition information sent by the image recognition terminal. The recognition information includes a tag, the category of the first image, and the offensive information of the first image, where the offensive information is used to identify whether the first image is an offensive image.
[0137] In an alternative implementation, the above-mentioned image recognition terminal (including any one or more of the first non-artificial image recognition terminal, the second non-artificial image recognition terminal, the first artificial image recognition terminal, and the second artificial image recognition terminal) will only send the recognition information to the server when it is determined that the offensive information of the first image indicates that the first image is an offensive image. The recognition information includes a tag, the category of the first image, and the offensive information of the first image. For example, the recognition information is (the URL address of the first image, "14", "yes"). The tag in this recognition information is the URL address of the first image, "14" indicates that the category of the first image is teeth, and "yes" is the offensive information, indicating that the first image is an offensive image.
[0138] In an alternative implementation, as long as the image recognition terminal determines that the category of the first image belongs to a preset category, it can send the recognition information to the server. The recognition information includes a tag, the category of the first image, and the offensive information of the first image. For example, the recognition information is (the URL address of the first image, "14", "yes"). The tag in this recognition information is the URL address of the first image, "14" indicates that the category of the first image is teeth, and "yes" is the offensive information, indicating that the first image is an offensive image. The recognition information is (the URL address of the first image, "14", "no"). The tag of this recognition information is the URL address of the first image, "14" indicates that the category of the first image is teeth, and "no" is the offensive information, indicating that the first image is a non-offensive image.
[0139] In an alternative implementation, as long as the image recognition terminal has a recognition result, it can send the recognition information to the server. For example, the recognition information is (the URL address of the first image, "14", "yes"). The tag in this recognition information is the URL address of the first image, "14" indicates that the category of the first image is teeth, and "yes" is the offensive information, indicating that the first image is an offensive image. The recognition information is (the URL address of the first image, "14", "no"). The tag of this recognition information is the URL address of the first image, "14" indicates that the category of the first image is teeth, and "no" is the offensive information, indicating that the first image is a non-offensive image. The recognition information is (the URL address of the first image, "00", "no"). The tag of this recognition information is the URL address of the first image, "00" indicates that the category of the first image is other, and "no" is the offensive information, indicating that the first image is a non-offensive image.
[0140] S309: If the information that is likely to cause aversion matches the preset information, the server sorts the category and the information that is likely to cause aversion into the attribute information according to the tag.
[0141] In an alternative implementation manner, based on the above example, assuming that the preset information is "yes", the server sorts the category and the information that is likely to cause aversion into the attribute information of the content to be published where the first picture is located according to the tag.
[0142] In an alternative implementation manner, based on the above example, assuming that the preset information is "yes" and "no", the server sorts the category and the information that is likely to cause aversion into the attribute information of the content to be published where the first picture is located according to the tag. That is, regardless of whether the first picture is a picture that is likely to cause aversion, the server will sort the attribute information.
[0143] In the embodiments of the present application, in order to simplify the acquisition of the recognition information of the first picture in the content to be published obtained by the server subsequently, the first picture of the currently uploaded content to be published can be matched with the previous first picture. For example, the URL address of the current first picture can be matched with the URL address of the previous first picture for which the recognition information has been obtained. If the URL address of the previous first picture contains the URL address of the current first picture, the category and the information that is likely to cause aversion are directly assigned to the current first picture, and the category and the information that is likely to cause aversion are sorted into the attribute information of the content to be published where the current first picture is located. If the URL address of the previous first picture does not contain the URL address of the current first picture, continue with steps 303-309.
[0144] In the embodiments of the present application, the server can screen the first picture according to the received recognition information of the first picture, and screen out the first pictures marked as pictures that are likely to cause aversion in the information that is likely to cause aversion, and store them in the picture library of pictures that are likely to cause aversion for iterative training of the third recognition model and the fourth recognition model of the second non-artificial picture recognition end, so as to improve the accuracy of the third recognition model and the fourth recognition model. In an alternative implementation manner, the third recognition model and the fourth recognition model can be trained with the pictures that are likely to cause aversion in the picture library of pictures that are likely to cause aversion every other time period. It is also possible to train the third recognition model and the fourth recognition model after a certain number of pictures that are likely to cause aversion are collected in the picture library of pictures that are likely to cause aversion.
[0145] In an alternative implementation manner, the audit start time, audit end time, and audit times of the above-mentioned first artificial picture recognition end, second artificial picture recognition end, and third artificial picture recognition end can be counted into the statistical server.
[0146] In the embodiments of the present application, the system may further include a content distribution end, which is used to determine whether the content to be published can be published based on the attribute information of each content to be published on the server side. For the convenience of elaborating various optional publishing methods of the content to be published below, the content to be published with offensive information containing offensive pictures in the attribute information is called the first type of content to be published, otherwise it is called the second type of content to be published.
[0147] In an optional implementation, the first type of content to be published will not be distributed by the content distribution end to users, that is, it will not be browsed by users and will not be provided for users to download. The second type of content to be published can be distributed by the content distribution end according to the timeline. That is to say, users can browse the published second type of content to be published according to the timeline and can download the published second type of content to be published through the content download service.
[0148] In an optional implementation, when a user provides feedback information to the platform that provides the content to be published, the user information can be recorded and stored in the sensitive user library. The content distribution end can distribute the content to be published according to the user information and the attribute information. If the user information of a certain user is stored in the sensitive user library, the first type of content to be published will not be distributed by the content distribution end to this user, while the users corresponding to the user information not stored in the sensitive user library can be distributed by the content distribution end according to the preset rules. The preset rules can be that one of every ten first type of content to be published can be published, and the publisher of this one first type of published content can be an authoritative publisher.
[0149] In an optional implementation, the sensitive database may not only contain user information, but also contain the categories of pictures reported by the user. The content distribution end can distribute the content to be published according to the user information, the category of the picture, and the attribute information. If the user information of a certain user is stored in the sensitive user library and the category of the picture reported by the user is teeth, the first type of content to be published corresponding to the teeth will not be distributed by the content distribution end to this user, and one of every ten other first type of content to be published can be published.
[0150] In an alternative embodiment, the image recognition terminal can also perform disgust grading on the first image whose disgust information is a disgusting image. In an alternative embodiment, the disgust grading can be based on the first distance value or the disgust level value. For example, high disgust level, medium disgust level, and low disgust level. And classify the disgust level into the recognition information and send it to the server terminal. In this way, the content distribution terminal can distribute the content to be published according to the user information, disgust level, category and attribute information of the image. For example, if the user information of a certain user is stored in the sensitive user library, and the image reported by this user is a tooth image with a high disgust level, then the content distribution terminal can not distribute the first type of content to be published corresponding to teeth with high disgust level and medium disgust level to this user, and appropriately distribute the first type of content to be published corresponding to teeth with low disgust level, etc.
[0151] The above distribution methods are only some alternative embodiments of the present application, and all feasible embodiments based on the present application are included in the present application.
[0152] In the embodiment of the present application, since the first image has been image-recognized before distribution, the recognition information will be used as the distribution condition in the subsequent distribution process, laying a foundation for reducing the bad reading experience caused by disgusting images to users. Reduce the use of remedial measures after the user experience.
[0153] The embodiment of the present application also provides an image recognition device. Figure 8 It is a schematic structural diagram of an image recognition device provided by the embodiment of the present application, as Figure 8 shown, the device includes:
[0154] The receiving module 801 is used to obtain the content to be published and the attribute information of the content to be published sent by the content production terminal. The attribute information includes the first image in the content to be published and the mark of the first image;
[0155] The sending module 802 is used to send the first image to the image recognition terminal; the first image carries a mark;
[0156] The receiving module 801 is used to receive the recognition information sent by the image recognition terminal; the recognition information includes the mark, the category of the first image, and the disgust information of the first image, where the disgust information is used to identify whether the first image is a disgusting image;
[0157] The processing module 803 is used to, if the disgust information matches the preset information, organize the category and disgust information into the attribute information according to the mark.
[0158] The device in the embodiment of the present application and the method embodiment are based on the same application concept.
[0159] The embodiment of the present application also provides an image recognition device. Figure 9This is a schematic structural diagram of a picture recognition device provided by an embodiment of the present application. As Figure 9 shown, the device includes:
[0160] A receiving module 901 is configured to receive a first picture sent by a server side, where the first picture carries a mark of the first picture;
[0161] A determining module 902 is configured to determine the category of the first picture and the offensive information of the first picture; the offensive information is used to identify whether the first picture is an offensive picture;
[0162] A sending module 903 is configured to send recognition information to the server side; the recognition information includes the mark, the category of the first picture, and the offensive information of the first picture.
[0163] In an optional implementation manner, the device further includes:
[0164] When the determining module 902 determines the category of the first picture, if the category of the first picture is a preset category, it determines the offensive degree value of the first picture. If the offensive degree value is greater than a preset degree value, it determines that the first picture is an offensive picture, and determines the offensive information of the first picture according to the offensive picture.
[0165] In an optional implementation manner, the device further includes:
[0166] The determining module 902 determines a plurality of distance values according to the second feature information of each second picture in a plurality of second pictures and the first feature information of the first picture; where the second pictures are offensive pictures obtained after receiving user feedback and passing the review, and the first feature information of the first picture is determined according to the pixels of the first picture. It determines the smallest distance value among the plurality of distance values as the first distance value; it determines the category of the second picture corresponding to the first distance value as the category of the first picture; if the first distance value is less than a preset second distance value, it determines that the first picture is an offensive picture, and determines the offensive information of the first picture according to the offensive picture.
[0167] In an optional implementation manner, the device further includes:
[0168] If the offensive degree value is less than the preset degree value, or the first distance value is greater than the preset second distance value, the sending module 903 sends the first picture to an artificial picture recognition end, so that the artificial picture recognition end determines the offensive information of the first picture. The first picture carries the mark of the first picture and the category of the first picture.
[0169] The device and the method embodiment in the embodiments of the present application are based on the same application concept.
[0170] The method embodiments provided by the embodiments of the present application can be executed in a computer terminal, a server, or a similar computing device. Taking running on a server as an example, Figure 10 is a hardware structure block diagram of a server for an image recognition method provided by the embodiments of the present application. As Figure 10 shown, the server 1000 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1010 (the processor 1010 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1030 for storing data, and one or more storage media 1020 for storing application programs 1023 or data 1022 (such as one or more mass storage devices). Among them, the memory 1030 and the storage media 1020 may be transient storage or persistent storage. The programs stored in the storage media 1020 may include one or more modules, and each module may include a series of instruction operations on the server. Further, the central processor 1010 may be configured to communicate with the storage media 1020 and execute a series of instruction operations in the storage media 1020 on the server 1000. The server 1000 may further include one or more power supplies 1060, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1040, and / or one or more operating systems 1021, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0171] The input / output interface 1040 can be used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by the communication provider of the server 1000. In one example, the input / output interface 1040 includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the input / output interface 1040 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0172] Those of ordinary skill in the art can understand that Figure 10 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the server 1000 may further include more or fewer components than Figure 10 shown, or have a different configuration from Figure 10 shown.
[0173] An embodiment of the present application further provides a storage medium, which can be set in a server to store at least one instruction, at least one program, a code set or an instruction set related to an image recognition method in a method embodiment. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned image recognition method.
[0174] Optionally, in this embodiment, the above storage medium may be located in at least one of multiple network servers in a computer network. Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs.
[0175] It can be seen from the embodiments of the image recognition method, device or storage medium provided by the present application that in the present application, the to-be-published content sent by the content production end and the attribute information of the to-be-published content are obtained. The attribute information includes the first image in the to-be-published content and the mark of the first image, and the first image is sent to the image recognition end; the first image carries the mark. The recognition information sent by the image recognition end is received; the recognition information includes the mark, the category of the first image and the offensive information of the first image. The offensive information is used to identify whether the first image is an offensive image. If the offensive information matches the preset information, the category and the offensive information are sorted into the attribute information according to the mark. Since the first image has been image-recognized before distribution, the recognition information will be used as the distribution condition in the subsequent distribution process, laying a foundation for reducing the bad reading experience caused by offensive images to users.
[0176] It should be noted that: the above sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of this specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be executed in a different order from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0177] Each embodiment in this specification is described in a progressive manner. The same or similar parts between each embodiment can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0178] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk, an optical disc, etc.
[0179] The above are only the preferred embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for image recognition, characterized in that, The method includes: Obtaining the content to be published sent by the content production end and the attribute information of the content to be published; the attribute information includes the first picture in the content to be published and the mark of the first picture; Sending the first picture to the picture recognition end; the first picture carries the mark; Receiving the recognition information sent by the picture recognition end; the recognition information includes the mark, the category of the first picture and the objectionable information of the first picture, wherein the objectionable information is used to identify whether the first picture is an objectionable picture; If the objectionable information matches the preset information, sorting the category and the objectionable information into the attribute information according to the mark; the content to be published is published by the content distribution end at least based on the attribute information and the user information; the user information represents a sensitive user determined based on the feedback information provided by the user; the objectionable information of the first picture is determined by sending the first picture to the manual picture recognition end or the non-manual picture recognition end based on the first distance value and the objectionable degree value; the first distance value is determined by the non-manual picture recognition end based on the smallest distance value among multiple distance values; the multiple distance values are determined according to the second feature information of each of multiple second pictures and the first feature information of the first picture; wherein, the second picture is an objectionable picture obtained after receiving user feedback and passing the review; the first feature information of the first picture is determined according to the pixels of the first picture; the objectionable degree value is determined by the non-manual picture recognition end.
2. The method according to claim 1, wherein The method further includes: The picture recognition end includes a non-manual picture recognition end and a manual picture recognition end.
3. A method for image recognition, characterized in that, The method includes: Receiving the first picture sent by the server end, the first picture carrying the mark of the first picture; Determining the category of the first picture and the objectionable information of the first picture; the objectionable information is used to identify whether the first picture is an objectionable picture; Sending the recognition information to the server end; the recognition information includes the mark, the category of the first picture and the objectionable information of the first picture; The first picture is the picture in the content to be published; if the objectionable information matches the preset information, the category and the objectionable information are used to be sorted into the attribute information of the content to be published according to the mark; The content to be published is published by the content distribution end at least based on the attribute information and the user information; the user information represents a sensitive user determined based on the feedback information provided by the user; The determining the category of the first picture and the objectionable information of the first picture includes: Determining the category of the first picture; Determining multiple distance values according to the second feature information of each of multiple second pictures and the first feature information of the first picture; wherein, the second picture is an objectionable picture obtained after receiving user feedback and passing the review, and the first feature information of the first picture is determined according to the pixels of the first picture; Determining the smallest distance value among the multiple distance values as the first distance value; Determining the objectionable degree value of the first picture; Send the first picture to a human picture recognition terminal or a non - human picture recognition terminal based on the first distance value and the offensiveness degree value, so that the human picture recognition terminal or the non - human picture recognition terminal determines the offensiveness information of the first picture.
4. The method according to claim 3, wherein Determining the category of the first picture and the offensiveness information of the first picture includes: Determine the category of the first picture; If the category of the first picture is a preset category, determine the offensiveness degree value of the first picture; If the offensiveness degree value is greater than a preset degree value, determine that the first picture is an offensive picture; Determine the offensiveness information of the first picture according to the offensive picture.
5. The method according to claim 3, characterized in that, Determining the category of the first picture and the offensiveness information of the first picture includes: Determine a plurality of distance values according to the second feature information of each second picture in a plurality of second pictures and the first feature information of the first picture; wherein, the second picture is an offensive picture obtained after receiving user feedback and passing the review, and the first feature information of the first picture is determined according to the pixels of the first picture; Determine the distance value with the smallest numerical value among the plurality of distance values as the first distance value; Determine the category of the second picture corresponding to the first distance value as the category of the first picture; If the first distance value is less than a preset second distance value, determine that the first picture is an offensive picture, and determine the offensiveness information of the first picture according to the offensive picture.
6. The method according to claim 4 or 5, characterized in that, The method further includes: If the offensiveness degree value is less than the preset degree value, or the first distance value is greater than the preset second distance value, send the first picture to a human picture recognition terminal, so that the human picture recognition terminal determines the offensiveness information of the first picture; The first picture carries the mark of the first picture and the category of the first picture.
7. An image recognition device, characterized in that, The device includes: A receiving module, configured to obtain the content to be published sent by a content production end and the attribute information of the content to be published; the attribute information includes the first picture in the content to be published and the mark of the first picture; A sending module, configured to send the first picture to a picture recognition terminal; the first picture carries the mark; The receiving module is configured to receive the recognition information sent by the picture recognition terminal; the recognition information includes the mark, the category of the first picture, and the offensiveness information of the first picture, wherein the offensiveness information is used to identify whether the first picture is an offensive picture; A processing module, configured to, if the offensive information matches the preset information, organize the category and the offensive information into the attribute information according to the tag; the content to be published is published by the content distribution end at least based on the attribute information and the user information; the user information is characterized as a sensitive user determined based on the feedback information provided by the user; the offensive information of the first picture is determined by sending the first picture to the artificial picture recognition end or the non-artificial picture recognition end based on the first distance value and the offensive degree value; the first distance value is determined by the non-artificial picture recognition end based on the smallest distance value among multiple distance values; the multiple distance values are determined according to the second feature information of each second picture among multiple second pictures and the first feature information of the first picture; wherein, the second picture is an offensive picture obtained after receiving user feedback and passing the review; the first feature information of the first picture is determined according to the pixels of the first picture; the offensive degree value is determined by the non-artificial picture recognition end.
8. An image recognition device, characterized in that, The device includes: A receiving module, configured to receive a first picture sent by the server end, where the first picture carries a tag of the first picture; A determining module, configured to determine the category of the first picture and the offensive information of the first picture; the offensive information is used to identify whether the first picture is an offensive picture; A sending module, configured to send recognition information to the server end; the recognition information includes the tag, the category of the first picture, and the offensive information of the first picture; The first picture is a picture in the content to be published; if the offensive information matches the preset information, the category and the offensive information are used to be organized into the attribute information of the content to be published according to the tag; The content to be published is published by the content distribution end at least based on the attribute information and the user information; the user information is characterized as a sensitive user determined based on the feedback information provided by the user; Determining the category of the first picture and the offensive information of the first picture includes: Determining the category of the first picture; Determining multiple distance values according to the second feature information of each second picture among multiple second pictures and the first feature information of the first picture; wherein, the second picture is an offensive picture obtained after receiving user feedback and passing the review, and the first feature information of the first picture is determined according to the pixels of the first picture; Determining the smallest distance value among the multiple distance values as the first distance value; Determining the offensive degree value of the first picture; Sending the first picture to the artificial picture recognition end or the non-artificial picture recognition end based on the first distance value and the offensive degree value, so that the artificial picture recognition end or the non-artificial picture recognition end determines the offensive information of the first picture.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the picture recognition method according to any one of claims 1-2 and the picture recognition method according to any one of claims 3-6.
10. A computer-readable storage medium, characterized in that, At least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the picture recognition method according to any one of claims 1-2 and the picture recognition method according to any one of claims 3-6.
Citation Information
Patent Citations
Picture detection method and apparatus
CN105989330A
Publishing method and device of contents
CN107943811A