Content moderation from cartoon faces
Patent Information
- Application Number
- PCT/EP2024/055549
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-04
- Publication Date
- 2025-10-02
AI Technical Summary
Existing content moderation systems fail to effectively address inappropriate content in cartoon faces, particularly concerning celebrities, and there is a need for methods to ensure compliance with guidelines and standards in the cartoon space.
A computing device configured to detect and moderate inappropriate content in cartoon images or videos by identifying celebrity faces using Generative Adversarial Networks and applying content guidelines, such as contextual information, user preferences, and platform-specific policies, with the ability to remove or block non-compliant content.
Effectively filters and moderates inappropriate content, including pornography, violence, and mockery, from cartoon faces, ensuring compliance with predefined guidelines and user safety, particularly for celebrity-related content.
Abstract
Description
[0001] CONTENT MODERATION FROM CARTOON FACES
[0002] TECHNICAL FIELD
[0003] The aspects of the disclosed embodiments relate generally to content moderation, and in particular, to visual content moderation of cartoon faces.
[0004] BACKGROUND
[0005] Content moderation generally refers to the process of monitoring, reviewing and managing user-generated content on platforms such as online platforms, websites, social media networks, forums and other digital environments. This process is typically carried out by platforms, streaming services, or other entities responsible for distributing or hosting content. One of the goals of content moderation is to prevent the dissemination of content that includes inappropriate, offensive, or harmful material.
[0006] Cartoons are a form of visual art and entertainment that typically involves illustrated characters, often accompanied by dialogue or narration. They can be in various forms, including digital images, animated TV shows, comic strips, comic books and animated films or videos. Content moderation of cartoons generally involves the review and control of the cartoon content to ensure that it complies with certain standards and guidelines.
[0007] More recent years have specifically witnessed an increasing attention in cartoon media, powered by the strong demands of industrial applications and the virtual world (e.g. Metaverse). It appears however, that most of the works on faces of persons in cartoons, referred to herein as “cartoon faces”, are focused on cartoon face generation, cartoon face detection or cartoon face recognition. It would be advantageous to provide content moderation in the cartoon space, and particularly, visual content moderation of cartoon faces to ensure compliance with standards and guidelines.
[0008] Accordingly, it would be desirable to provide methods and apparatus that addresses at least some of the problems described above.
[0009] SUMMARY
[0010] The aspects of the disclosed embodiments are generally directed to content moderation in the cartoon space.
[0011] According to a first aspect, the above and further advantages are obtained by an apparatus. In one embodiment, the apparatus includes a computing device that is configured to receive one or more images as an input; detect whether content in the one or more images complies with content guidelines; and moderate the content of the one or more images when an image does not comply with the content guidelines. The aspects of the disclosed embodiments are directed to e-cartoon content moderation and filtering inappropriate content from e-cartoons.
[0012] In a possible implementation form the one or more images comprise a cartoon or a cartoon video. The aspects of the disclosed embodiments are directed to e-cartoon content moderation and filtering inappropriate content from e-cartoons.
[0013] In a possible implementation form computing device is configured to moderate the content of the one or more images by one or more of removing the image from the one or more images or blocking access to the image. The aspects of the disclosed embodiments are directed to e-cartoon content moderation and filtering inappropriate content from e-cartoons.
[0014] In a possible implementation form the computing device is further configured to, prior to moderating the content, determine if the image includes a celebrity; and moderate the content only if the image includes the celebrity. The aspects of the disclosed embodiments perform e-cartoon moderation of celebrities in the cartoon space. In a possible implementation form the computing device is further configured to determine if the image includes a celebrity by: detecting a character face in the image; and identifying whether the character face corresponds to a face of celebrity from a list of celebrities. The aspects of the disclosed embodiments perform e-cartoon moderation of celebrities in the cartoon space. The celebrity can be identified by comparing a cartoon face of the celebrity to facial images detected in the cartoon image(s) or video.
[0015] In a possible implementation form the character face is detected using a Generative Adversarial Network that is configured to compare cartoons of real faces to the detected character face. The aspects of the disclosed embodiments are directed to comparing cartoon faces to cartoon faces by converting real faces to cartoon images.
[0016] In a possible implementation form the predefined content guidelines are based on one or more of contextual information, user preferences, or platform-specific content policies.
[0017] In a possible implementation form the computing device is further configured to detect a character in the scene and identify whether the detected character corresponds to a targeted character. The aspects of the disclosed embodiments perform e- cartoon moderation of celebrities in the cartoon space. The search engine is configured to filter any inappropriate content concerning celebrities in a predetermined list and return moderated images to the user.
[0018] In a possible implementation form the computing device is further configured to detect a character face in the scene and identify whether the character face corresponds to a targeted character face. The aspects of the disclosed embodiments perform e- cartoon moderation of celebrities in the cartoon space. The search engine is configured to filter any inappropriate content concerning celebrities in a predetermined list and return moderated images to the user.
[0019] In a possible implementation form the predefined content guidelines include at least a first guideline and a second guideline, and the computing device is configured to: remove the content from the scene of the cartoon video when the content does not comply with the first guideline; or moderate the content in the scene of the cartoon when the content does not comply with the second guideline. The aspects of the disclosed embodiments perform visual content moderation from cartoon faces of celebrities by removing inappropriate content and returning moderated images.
[0020] In a possible implementation form, predefined content guidelines are based on one or more of contextual information, user preferences, or platform-specific content policies. The aspects of the disclosed embodiments perform visual content (pornography, violence, mockery) moderation from cartoon faces of celebrities by removing inappropriate content and returning moderated images.
[0021] In a possible implementation form the one or more images of the cartoon comprise cartoon characters and associated facial features. The aspects of the disclosed embodiments perform visual content moderation from cartoon faces of celebrities by removing inappropriate content and returning moderated images.
[0022] In a possible implementation form, the computing device includes a processor configured to execute non-transitory machine readable instructions, which when executed, cause the computing device to perform the processes of the possible implementation forms. The aspects of the disclosed embodiments perform visual content moderation from cartoon faces of celebrities by removing inappropriate content and returning moderated images.
[0023] According to a second aspect, the above and further advantages are obtained by a computer implemented method. In one embodiment, the method includes receiving one or more images as an input; detecting whether content in the one or more images complies with content guidelines; and moderating the content of the one or more images when an image of the one or more images does not comply with the content guidelines. The aspects of the disclosed embodiments are directed to e-cartoon content moderation and filtering inappropriate content from e-cartoons.
[0024] In a possible implementation form, the one or more images comprise a cartoon or a cartoon video. The aspects of the disclosed embodiments perform e-cartoon moderation of celebrities in the cartoon space.
[0025] In a possible implementation form, the method further includes moderate the content of the one or more images by one or more of removing the image from the one or more images or blocking access to the image. The aspects of the disclosed embodiments filter any inappropriate content concerning celebrities in a predetermined list and return moderated images to the user.
[0026] In a possible implementation form, the method further includes prior to moderating the content: determining if the image includes a celebrity; and moderating the content only if the image includes the celebrity.
[0027] In a possible implementation form, the method, when the method further includes determining if the image includes a celebrity by detecting a character face in the image; and identifying whether the character face corresponds to a celebrity from a list of celebrities. The aspects of the disclosed embodiments perform visual content moderation from cartoon faces of celebrities by removing inappropriate content and returning moderated images.
[0028] In a possible implementation form the character face is detected using a Generative Adversarial Network that is configured to compare cartoons of real faces to the detected character face. The aspects of the disclosed embodiments are directed to comparing cartoon faces to cartoon faces by converting real faces to cartoon images.
[0029] According to a third aspect, the above and further advantages are obtained by a computer program product. In one embodiment, the computer readable instructions are embodied on a non-transitory computer readable medium, which when executed by a computing device, are configured to carry out the methods of the possible implementation forms. The aspects of the disclosed embodiments are directed to using machine learning to moderate content in the cartoon space.
[0030] According to a fourth aspect, the above and further advantages are obtained by a computer-implemented method for moderating cartoon facial content using an artificial intelligence system. In one embodiment the method includes receiving input data comprising cartoon images and associated facial features; analyzing the input data using a trained machine learning model to detect facial characteristics within the cartoon images; generating a moderation decision based on the analysis, wherein the decision indicates whether the cartoon facial content complies with predefined content guidelines; automatically applying moderation actions to the cartoon facial content based on the moderation decision. The aspects of the disclosed embodiments are directed to using machine learning to moderate content in the cartoon space.
[0031] In a possible implementation form, the method further includes using a neural network to detect facial characteristics within the cartoon images, analyze the input images and generate a moderation decision based on the analysis. The aspects of the disclosed embodiments are directed to using machine learning to moderate content in the cartoon space.
[0032] In a possible implementation form the moderation actions include one or more of allowing, filtering or flagging the content. The aspects of the disclosed embodiments perform visual content moderation from cartoon faces of celebrities by removing inappropriate content and returning moderated images.
[0033] In a possible implementation form the moderation decision is further based on one or more of contextual information, user preferences, or platform-specific content policies. The aspects of the disclosed embodiments perform visual content (pornography, violence, mockery) moderation from cartoon faces of celebrities by removing inappropriate content and returning moderated images. In a possible implementation form a computer system with a processor, memory, and a communication interface is configured to perform the method of the possible implementation forms. The aspects of the disclosed embodiments are directed to content moderation of cartoon images across any suitable platform, such as social media platforms and search engines.
[0034] In a possible implementation form the content moderation apparatus comprises a mobile communication device. The aspects of the disclosed embodiments enable content moderation of cartoon images across any suitable platform, such as social media platforms and search engines.
[0035] In a possible implementation form the computing device comprises a convolutional neural network. The aspects of the disclosed embodiments are directed to using machine learning to detect and moderate cartoon content.
[0036] These and other aspects, implementation forms, and advantages of the exemplary embodiments will become apparent from the embodiments described herein considered in conjunction with the accompanying drawings. It is to be understood, however, that the description and drawings are designed solely for purposes of illustration and not as a definition of the limits of the disclosed invention, for which reference should be made to the appended claims. Additional aspects and advantages of the invention will be set forth in the description that follows, and in part will be obvious from the description, or may be learned by practice of the invention. Moreover, the aspects and advantages of the invention may be realized and obtained by means of the instrumentalities and combinations particularly pointed out in the appended claims.
[0037] BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In the following detailed portion of the present disclosure, the invention will be explained in more detail with reference to the example embodiments shown in the drawings, in which like references indicate like elements and:
[0039] Figure 1 illustrates exemplary architecture of a system incorporating aspects of the disclosed embodiments.
[0040] Figure 2 illustrates a schematic block diagram of an exemplary process flow incorporating aspects of the disclosed embodiments.
[0041] Figure 3 illustrates a schematic block diagram of an exemplary implementation incorporating aspects of the disclosed embodiments.
[0042] Figure 4 illustrates a schematic block diagram of an exemplary process flow incorporating aspects of the disclosed embodiments.
[0043] DETAILED DESCRIPTION OF THE DISCLOSED EMBODIMENTS
[0044] Figure 1 illustrates a diagram of an exemplary apparatus incorporating aspects of the disclosed embodiments. As shown in Figure 1 the apparatus 100 generally includes a content moderation device 104 that is configured to receive an input 102. The input 102 generally comprises one or more images. The content moderation device 104 is configured to determine whether the one or more images comply with content guidelines. If any one of the one or more images do not comply with the content guidelines, the content moderation device 104 is configured to provide moderated cartoon content 118 as an output. The aspects of the disclosed embodiments are generally directed to content moderation of cartoons, and in particular visual content moderation of cartoon characters that resemble well-known individuals, such as celebrities. A particular emphasis is made herein with respect to the faces of such individuals. Although, the aspects of the disclosed embodiments are configured to be applied to any aspect or feature of a cartoon, or a character that is portrayed in a cartoon. In the example of Figure 1, the content moderation device 104 generally comprises, and will be referred to herein as, a computing device. As shown in Figure 1, the computing device 104 is configured to receive one or more images as the input 102. In one embodiment, the input images 102 comprise cartoon images. While images are generally referred to herein, the aspects of the disclosed embodiments are not so limited. The aspects of the disclosed embodiments can be generalized to different kinds of virtual world and images or videos seen in cartoons, caricatures, sketches and drawings.
[0045] The computing device 104 is configured to detect whether content in the one or more images comply with certain content guidelines. This can also be described as whether the content of a scene in the one or more images complies with the content guidelines. For example, an image can include one or more scenes. The term “one or more images” as used herein can also include videos or video streams.
[0046] The content guidelines are generally rules or parameters used to ensure that the content published or shared by users complies with guidelines such as community guidelines, terms of service, and legal regulations set by the platform, website or other group. As described herein, the aspects of the disclosed embodiments are generally configured to categorize the content as acceptable, objectional or offensive. In alternate embodiments, the content can be categorized in any suitable manner.
[0047] If it is determined or detected that the content does not comply with the content guidelines, the computing device 104 is configured to moderate the content. This can include removing the image or scene with the objectional content from the one or more cartoon images. For example, if the content is detected in a single image or group of images, the image(s) with the objected to content can be prevented from being presented to or otherwise viewed by the user. If for example the content is in a video, the scene(s) from the video that include the objected to content can be removed or blocked from being viewed in any suitable manner.
[0048] The aspects of the disclosed embodiments are generally directed to the identification of characters and character faces of well- known individuals, generally referred to herein as celebrities. The term “celebrities” as is generally understood can include, but is not limited to, actors, musicians, athletes, politicians, authors, entrepreneurs, and other notable figures. However, the aspects of the disclosed embodiments are not so limited. In alternate embodiments, the aspects of the disclosed embodiments can be implemented with respect to any desired person, class, type or group of individuals.
[0049] Additionally, if users are of a certain age, or age group, the aspects of the disclosed embodiments can moderate content to ensure that the content is suitable for viewing by such users. In this example, the content guidelines would establish rules pertaining to the specific ages or age group. For example, it may be advantageous to prevent users in certain age groups from viewing content that includes violence or pornography. The aspects of the disclosed embodiments are configured to detect and moderate such content, as is generally described herein.
[0050] Figure 2 is a flowchart of an exemplary process incorporating aspects of the disclosed embodiments. As shown in Figure 2, the input 202 is in the form of one or more images, such as cartoon image(s) or a cartoon video. Although the aspects of the disclosed embodiment are described herein with respect to cartoons, the aspects of the disclosed embodiments are not so limited. For example, the aspects of the disclosed embodiments can be applied to any format, media or platform that enables an image or caricature of a person to be rendered, other than including a cartoon.
[0051] As illustrated in Figure 2, the content, context and semantics of the input cartoon image(s) or video 202 are detected, processed and analyzed using a video clip understanding block 204. The video clip understanding block 204 is configured to render a textual description of the content. In one embodiment, the video clip understanding block 204 is a video understanding artificial intelligence (Al) model configured to analyze and comprehend the content of videos. The output generated by the block 204 can include a detailed text description of the content within a scene(s) or image(s) of the input 202. In one embodiment, a Large Language Model (LLM) 206 is applied to the textual description of the content. The Large Language Model is generally configured to interpret the text description of the scene and classify the content of the scene according to one or more content guidelines. In one embodiment, an Al based Large Language Model such as PanGu-S™ or Chat-GPT™ can be used. In alternate embodiments, any suitable model can be used that understand and generate human-like text based on the input it receives.
[0052] The content guidelines, as the term is generally used herein, refers to the classification of content of a scene as acceptable, objectional or unacceptable. For purposes of the description herein, the terms “objectional”, “unacceptable” or “inappropriate” can be used to refer to content that includes one or more of pornography, violence or mockery. It will be understood however, that terms used herein are merely exemplary and any suitable classification can be implemented. The aspects of the disclosed embodiments are not intended to be limited to categories such as pornography, violence or mockery. Rather, the aspects of the disclosed embodiments can be configured to determine or detect any type or category or classification of content.
[0053] In one embodiment, a first guideline can be established for content that is deemed to be unacceptable, inappropriate or offensive content such as violence and pornography. According to the aspects of the disclosed embodiments, this type of content can or will be moderated or removed. In one embodiment, moderated can include preventing an image or images from being presented to, or otherwise viewed by the user. For example, if a user conducts a search via a search engine and content is returned that does not comply with the content guidelines, the content can be blocked from being viewed or accessed.
[0054] A second guideline can be for content that might be considered objectionable, subject to further review, such as mockery. Generally, this type of content may not rise to the level of offensive or inappropriate content that requires immediate modification and / or removal. The aspects of the disclosed embodiments enable further analysis of content that falls within this second guideline, prior to any modification or removal of scenes with such content.
[0055] The content guidelines described herein are merely examples. In alternate embodiments, any suitable criteria can be used to determine and set the content guidelines. The number of content guidelines is not limited to two. Rather, any suitable number of computational models can be trained for specific criteria, types or levels of differentiation and distinction of content. Once trained, such computational models can be applied within the aspects of the disclosed embodiments to analyze the inputs) and discern the different levels and types of content that can or should be moderated.
[0056] Referring back to Figure 2, the output of the Large Language Model 206 is used to determine 208 whether the scene includes any objectional content. If the output of the Large Language Model 206 indicates that there is no objectional content, the guidelines in this instance will indicate that the content of the scene can be retained 224. No further processing or analysis needs to be performed. The user will be able to view the content.
[0057] If it is determined 208 from the output of the Large Language Model 206 that the scene includes objectional content, in one embodiment the type of the objectional content is detected or determined. If the type of objectional content is determined 218 not to be mockery, the content of the scene is categorized as other objectional content. Other objectional content can include, but is not limited to offensive content, violence and pornography. In this example, the first predetermined guidelines can be applied to offensive content, which indicate that such content should be removed 222 from the scene, or otherwise rendered unviewable.
[0058] In one embodiment, removing 222 the offensive content can include, but is not limited to, removing the particular scene(s) or image frame(s), or modifying the scene so that the offensive content cannot be visualized. The aspects of the disclosed embodiments contemplate any suitable method for modifying, blocking or removing a scene to moderate offensive content. If it is determined 210 if the objectionable content is mockery, further analysis of the content is undertaken. In this example, mockery can correspond to the second content guideline.
[0059] In one embodiment, once the content is determined 210 to be mockery, face detection 214 is applied to the scene. As is illustrated in the example of Figure 2, a real-to-cartoon face model 230 is used to convert images of faces of individuals 232 into cartoon faces. The aspects of the disclosed embodiments rely on cartoon faces as inputs, with real faces in the face database 232. It is generally understood to be challenging to compare real faces to cartoon faces. A simpler approach is to compare real faces to real faces, or cartoon faces to cartoon faces. However, the aspects of the disclosed embodiments are configured to convert real faces into cartoon faces using style GAN (Generative Adversarial Network) transfer, also referred to as StyleGAN, and then perform face recognition in the cartoon space.
[0060] As is illustrated in Figure 2, the real-to-cartoon face model 230 is configured to convert the real face image in the real face database 232 to a cartoon face. During training, a database 232 of real faces is converted to cartoon faces 236 using the StyleGAN model 234. The cartoon faces 236 are transferred and face recognition 214 is performed in the cartoon space.
[0061] If a face is recognized 216, cartoon to cartoon face recognition 214 is then performed. The cartoon face recognition 214 generally includes determining 216 whether the detected cartoon face is among a list of targeted faces, such as a face of a celebrity. If the detected cartoon face is among the list of targeted faces, the scene is considered to be most likely a mockery about or that includes the celebrity. In that case, in order to prevent the scene or image from being viewed, the scene can be removed or filtered 222. If the detected cartoon face is not recognized, the scene is considered most likely not to be a mockery about or involving a celebrity. In that case, no action is taken, and the scene is maintained 224 for viewing.
[0062] The deep learning models of the disclosed embodiments can be trained on weakly labeled data, and adapted to specific capture conditions using unlabeled data. To reduce the cost and ambiguity of data annotation, weakly supervised learning (WSL) methods are proposed for training machine learning models using data with reduced supervision, either with incomplete (a subset of data is labelled), inexact (coarse grained labels), or inaccurate supervision (ambiguous labels). The aspects of the disclosed embodiments are not intended to be limited by the particular type of training or training model.
[0063] Figure 3 illustrates one example of a process flow of an exemplary implementation incorporating aspects of the disclosed embodiments. In this example implementation, also referred to as a e-cartoon safety system, the content moderation apparatus of the disclosed embodiments is used to analyze a video and potentially filter any inappropriate content, such as inappropriate content concerning a celebrity.
[0064] As illustrated in the example of Figure 3, a cartoon video is uploaded 302 to a social media video platform. For example, in one embodiment, the cartoon video can be uploaded by the user to the social media video platform. Although a cartoon video is referred to herein, the aspects of the disclosed embodiments are not so limited. In alternate embodiments, and suitable media and format can be uploaded. For example, a single image, or a batch of images could be uploaded.
[0065] In one embodiment, a video batch is formed 304. The video batch 304 is then applied to a cartoon content moderation system 306 incorporating aspects of the disclosed embodiments, such as that described with respect to Figures 1 and 2. Although a video batch is generally referred to herein, the aspects of the disclosed embodiments are not so limited. In alternate embodiments a single image or video can be applied.
[0066] The length of the uploaded video can be of any suitable length. In one embodiment, a maximum length of the uploaded video can be set by the system. For example, by default, a length of the video can be set to 10 seconds. In alternate embodiments, the length of the video can be any suitable length, other than including 10 seconds. When an uploaded video exceeds the predetermined length, the uploaded video can be segmented into sub-batches. The subbatches can be processed as described herein, sequentially within a loop, until sub-batches have been processed.
[0067] As shown in the example of Figure 3, in one embodiment, a pre-defined list of celebrities 308 is applied to the content moderation system 306. The pre-defined list 308 can be stored in a database, for example, and formed or updated in any suitable manner.
[0068] In one embodiment, a user can select or enter names of individuals to form the pre-defined list. The corresponding faces of the celebrities can then be identified and accessed in any suitable manner. As was discussed above, the aspects of the disclosed embodiments are configured to convert the real faces of the celebrities in the pre-defined list 308 into cartoon faces using style GAN transfer. Face recognition is then performed in the cartoon space.
[0069] The content moderation system 306 is configured to filter the video batch 304 of uploaded cartoon videos 302 against the predefined list of celebrities, as is generally described herein with respect to the embodiments illustrated in Figures 1 and 2. When a scene that includes a celebrity is detected, the scene can be analyzed to determine if the scene complies with the content guidelines. If scene does not comply, the scene can be moderated as is generally described herein.
[0070] For example, as shown in Figure 3, in one embodiment, it is determined 310, from the output of the content moderation 306, whether a celebrity from the list of pre-defined celebrities, has been recognized. If a celebrity is recognized, the image or video is deleted 312. If a celebrity is not recognized, the image or video is made available for viewing 314.
[0071] Figure 4 is another example of an implementation of the cartoon content moderation system or apparatus of the disclosed embodiments. In this example, the content moderation system of the disclosed embodiments is applied to, or implemented as part of, a search engine application. As a particular example, the aspects of the disclosed embodiments will be described herein with respect to the Petal Search Image Search Engine application. The Petal Search Image Search Engine Application is a search engine developed by Huawei. Although the Petal Search Engine is described herein, the aspects of the disclosed embodiments are not so limited. In alternate embodiments, the aspects of the disclosed embodiments can be applied to, or implemented within, any suitable search engine.
[0072] As is shown in the example of Figure 4, the user conducts a search 402 using a search engine incorporating the content moderation system of the disclosed embodiments. In this specific example, the search 402 is for a celebrity. A set of search results 404 is returned. It is determined or detected 406 whether there is a cartoon in the set of search results. If no cartoon(s) are detected in the set of search results, the process ends 416 and the set of search results is made available 418.
[0073] However, if there is a cartoon, the search engine in this example is configured to perform 408 cartoon moderation of celebrities in the cartoon space, using the content moderation apparatus of the disclosed embodiments. In this example, the scenes of the cartoon are analyzed to detect or otherwise identify scenes that include individuals or persons. The scenes can then be further analyzed, as is described for example with respect to Figures 2 and 3, to determine if the scenes comply with the content guidelines. In one embodiment, scenes that do not comply with the content guidelines can be flagged or otherwise identified.
[0074] In one embodiment, the identity of the person(s) in an image or scene that is flagged is determined 410. It is then determined 412 whether the identified person is a celebrity, or a person from the pre-defined list.
[0075] If a celebrity or person from the list is not identified, the process ends 416 and the final set of search results 418 is made available. However, if it is determined 412 that there is celebrity or listed person in the scene(s), the particular scene(s) or content is hidden or removed 414. The final set of search results 418 in this example will not include the hidden or removed content. The process ends 416 and the final set of search results 418 is made available. In one embodiment, the content moderation device 104 described in Figure 1 comprises a computing device. The computing device can comprise a hardware device, such as for example a computer processing device and / or memory and storage. The computer processing device can comprise a processor, Central Processing Unit (CPU), a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer and a microprocessor, for example. The computer processing device may be configured to carry out program code by performing arithmetical, logical, and input / output operations, according to the program code. Once the program code is loaded into a computer processing device, the computer processing device may be programmed to perform the program code, thereby transforming the computer processing device into a special purpose computer processing device. In a more specific example, when the program code is loaded into a processor, the processor becomes programmed to perform the program code and operations corresponding thereto, thereby transforming the processor into a special purpose processor.
[0076] The aspects of the disclosed embodiments provide solutions for performing content moderation from cartoon faces of celebrities. Large Language Models are fine-tuned to learn to detect inappropriate content from the input cartoon scene.
[0077] Accurate face detection is performed in the cartoon space by transferring the knowledge from face detection in real space. The real faces are converted into cartoon faces using style GAN (Generative Adversarial Network) transfer. Face recognition is then performed in the cartoon space.
[0078] The aspects of the disclosed embodiments utilize a specialized and adaptive architecture that can be trained on weakly labeled data, and adapted to specific list of known individuals or celebrities using unlabeled or weakly labeled data. In the model of the disclosed embodiments, the list of people in the database can be defined, changed and updated at any time. When adding or removing one person to the database, the model does not need to be retrained, due to this specialized and adaptive architecture. Moreover, when adding a person to the database, the name of the celebrity does not need to be specified, because the system can work without labeling or giving names to the faces.
[0079] The list of celebrities or individuals is flexible, and can be defined and provided by the user in an incremental basis. Moreover, the aspects of the disclosed embodiments can be generalized to different kinds of virtual world and images seen in cartoons, caricatures, sketches, drawings.
[0080] Thus, while there have been shown, described, and pointed out, fundamental novel features of the invention as applied to the exemplary embodiments thereof, it will be understood that various omissions, substitutions and changes in the form and details of devices and methods illustrated, and in their operation, may be made by those skilled in the art without departing from the spirit and scope of the presently disclosed invention. Further, it is expressly intended that all combinations of those elements, which perform substantially the same function in substantially the same way to achieve the same results, are within the scope of the invention. Moreover, it should be recognized that structures and / or elements shown and / or described in connection with any disclosed form or embodiment of the invention may be incorporated in any other disclosed or described or suggested form or embodiment as a general matter of design choice. It is the intention, therefore, to be limited only as indicated by the scope of the claims appended hereto.
Claims
CLAIMS1. An apparatus comprising: a computing device configured to: receive one or more images as an input; detect whether content in the one or more images complies with content guidelines; and moderate the content of the one or more images when an image of the one or more images does not comply with the content guidelines.
2. The apparatus according to claim 1 , wherein the one or more images comprise a cartoon or a cartoon video.
3. The apparatus according to any one of claim 1 or 2, wherein the computing device is configured to moderate the content of the one or more images by one or more of removing the image from the one or more images or blocking access to the image.
4. The apparatus according to any one of the preceding claims, wherein the computing device is further configured to, prior to moderating the content: determine if the image includes a celebrity; and moderate the content only if the image includes the celebrity.
5. The apparatus according to claim 4, wherein the computing device is further configured to determine if the image includes a celebrity by: detecting a character face in the image; and identifying whether the character face corresponds to a celebrity from a list of celebrities.
6. The apparatus according to claim 5, wherein the character face is detected using a Generative Adversarial Network that is configured to compare cartoons of real faces to the detected character face.
7. The apparatus according to any one of the preceding claims wherein the predefined content guidelines are based on one or more of contextual information, user preferences, or platform-specific content policies.
8. A computer implemented method, comprising: receiving one or more images as an input; detecting whether content in the one or more images complies with content guidelines; and moderating the content of the one or more images when an image of the one or more images does not comply with the content guidelines.
9. The computer implemented method according to claim 8, wherein the one or more images comprise a cartoon or a cartoon video.
10. The computer implemented method according to any one of claims 8 or 9, further comprising: moderate the content of the one or more images by one or more of removing the image from the one or more images or blocking access to the image.
11. The computer implemented method according to any one of claims 8 to 10, further comprising, prior to moderating the content: determining if the image includes a celebrity; and moderating the content only if the image includes the celebrity.
12. The computer implemented method according to any one of claims 8 to 11, further comprising determining if the image includes a celebrity by: detecting a character face in the image; and identifying whether the character face corresponds to a celebrity from a list of celebrities.
13. The computer implemented method according to any one of claims 8 to 12, wherein the character face is detected using a Generative Adversarial Network that is configured to compare cartoons of real faces to the detected character face.
14. The computer implemented method according to any one of claims 8 to 13, further comprising: providing a set of search results from a search engine as the one or more images for the input; and removing any search results from the set of search results that is determined to be moderated content.
15. A computer program product comprising computer readable instructions embodied on a non-transitory computer readable medium, which when executed by a computing device, are configured to carry out the method according to any one of claims 8 to 14.