Automated reviewer of platform content for harmful content using generative artificial intelligence

WO2025184818A8PCT designated stage Publication Date: 2025-10-02MICROSOFT TECHNOLOGY LICENSING LLC +11
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/080264
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-06
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing content moderation systems face inefficiencies due to the need for human review after automated flagging, which increases latency and costs, as they struggle to adapt to the diverse interpretations of harmful content across different platforms.

Method used

A system utilizing Generative Artificial Intelligence (GAI) to generate customized prompt templates based on platform-specific guidelines, allowing for automated and efficient detection of harmful content without extensive human intervention.

Benefits of technology

This approach enhances content moderation efficiency by aligning automated detection with community-specific guidelines, reducing latency and costs while maintaining high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024080264_02102025_PF_FP_ABST
    Figure CN2024080264_02102025_PF_FP_ABST
Patent Text Reader

Abstract

A device or method receives user input defining what content on a particular platform is to be blocked as harmful content; generates a customized prompt template for submission to a Generative Artificial Intelligence (GAI) annotation service, the customized prompt template instructing the GAI annotation service to label harmful content within an annotation request according to the user input defining what content on a particular platform is to be blocked as harmful content; receives an annotation request comprising platform content to be vetted for any harmful content based on the user input; generates a prompt for the GAI annotation service by combining the platform content with the customized prompt template; and receives an identification of harmful content that is to be blocked in the platform content from the GAI annotation service based specifically on the user input defining what content is harmful content.
Need to check novelty before this filing date? Find Prior Art

Description

AUTOMATED REVIEWER OF PLATFORM CONTENT FOR HARMFUL CONTENT USING GENERATIVE ARTIFICIAL INTELLIGENCEBACKGROUND

[0001] In the ever-evolving landscape of online platforms and digital communities, the challenge of policing harmful content has become a significant concern. Human review of all of the huge volume of electronic content being generated on most platforms is entirely impractical. Consequently, various industries have implemented content safety techniques, deploying automated light-weight classifiers to flag potentially harmful content. However, a significant hurdle lies in the diverse interpretations of what constitutes "harmful content" across different use cases. This variability still necessitates an additional layer of human review after content is flagged in an attempt at aligning the results with the unique community guidelines or terms of use for each platform. Unfortunately, this manual review introduces substantial costs and latency, limiting the efficiency of content moderation efforts.

[0002] Thus, a technical problem is presented of detecting harmful content in a manner that is tailored to a particular platform or digital community while minimizing subsequent human review and the associated increase in communication latency and other factors. In other words, there is a need for a more effective, streamlined and cost-effective approach to content moderation in a particular context or platform.SUMMARY

[0003] In one general aspect, the following description presents a device including a processor, and a memory storing executable instructions which, when executed by the processor, causes the processor, alone or in combination with other processors, to perform the following functions: receive user input defining what content on a particular platform is to be blocked as harmful content; based on the user input, generate a customized prompt template for submission to a Generative Artificial Intelligence (GAI) annotation service, the customized prompt template instructing the GAI annotation service to label harmful content within an annotation request according to the user input defining what content on a particular platform is to be blocked as harmful content; receive an annotation request comprising platform content to be vetted for any harmful content based on the user input; generate a prompt for the GAI annotation service by combining the platform content with the customized prompt template; and receive identification of harmful content that is to be blocked in the platform content from the GAI annotation service based specifically on the user input defining what content is harmful content.

[0004] In another general aspect, the following description presents a method of blocking harmful content from platform content of a platform for a particular electronic community. The method includes: receiving user input defining what content on a particular platform is to be blocked as harmful content; generating, based on the user input, a customized prompt template for submission to a Generative Artificial Intelligence (GAI) annotation service, the customized prompt template instructing the GAI annotation service to label harmful content within an annotation request according to the user input defining what content on a particular platform is to be blocked as harmful content; receiving an annotation request comprising platform content to be vetted for any harmful content based on the user input; generating a prompt for the GAI annotation service by combining the platform content with the customized prompt template; and receiving identification of harmful content that is to be blocked in the platform content from the GAI annotation service based specifically on the user input defining what content is harmful content.

[0005] In another general aspect, the following description presents a system for use in blocking harmful content from a platform serving an electronic community, the community having guidelines as to what constitutes harmful content, the system including: a processor in communication with the platform to receiving platform content to be vetted for harmful content; a memory storing executable instructions which, when executed by the processor, causes the processor, alone or in combination with other processors, to perform the following functions: generate a customized prompt template, based on the guidelines, for submission to a Generative Artificial Intelligence (GAI) annotation service, the customized prompt template instructing the GAI annotation service to label harmful content within an annotation request according to the guidelines; receive an annotation request comprising platform content to be vetted for any harmful content; generate a prompt for the GAI annotation service by combining the platform content with the customized prompt template; and generate, with the prompt, an identification of harmful content that is to be blocked in the platform content from the GAI annotation service based specifically on the guidelines defining what content is harmful content.

[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to implementations that solve any or all disadvantages noted in any part of this disclosure.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The drawing figures depict one or more implementations in accord with the present teachings, by way of example only, not by way of limitation. In the figures, like reference numerals refer to the same or similar elements. Furthermore, it should be understood that the drawings are not necessarily to scale.

[0008] FIG. 1 is a flowchart illustrating an example of automated content review under principles described herein.

[0009] FIG. 2A is a system diagram with additional detail of an example of automated content review under principles described herein.

[0010] FIG. 2B is another flowchart with additional detail of an example of automated content review under principles described herein.

[0011] FIG. 3A illustrates an example of user guidelines defining harmful content for a particular platform.

[0012] FIG. 3B illustrates an example of a customized prompt template for a Generative Artificial Intelligence using the guidelines of FIG. 3A.

[0013] FIG. 4 is a flowchart with additional detail of an example of automated content review under principles described herein.

[0014] FIG. 5A depicts a first level of prompt compression under principles described herein.

[0015] FIG. 5B depicts a second level of prompt compression under principles described herein.

[0016] FIG. 6 depicts a method of assigning a confidence level along with labeling for examples of harmful content under principles described herein.

[0017] FIG. 7 is a block diagram illustrating an example software architecture, various portions of which may be used in conjunction with various hardware architectures herein described.

[0018] FIG. 8 is a block diagram illustrating components of an example machine configured to read instructions from a machine-readable medium and perform any of the features described herein.DETAILED DESCRIPTION

[0019] A digital or electronic community may have written guidelines for appropriate communication on their platform. For example, these guidelines may specify one or more objectives such as:

[0020] Civility is at the core of our community. To promote a civil environment, our community prohibits certain types of content, including:

[0021] Real-World Sensitive Events: Our community prohibits content that recreates specific real-world events of a sensitive nature, mocks the victims of such an event, supports or promotes the perpetrators or outcome of such event or uses such an event for commercial purposes, including content about (1) mass acts of violence, (2) specific, real-world natural disasters, or (3) human or civil rights violations.

[0022] Violent Content and Gore: Our community prohibits content containing extreme violence or serious physical or psychological abuse including: (1) Animal abuse and torture, (2) realistic or real-world depictions of extreme violence, gore or death…

[0023] A more detailed version of this example is illustrated in FIG. 3A.

[0024] To implement these guidelines, the operators of the platform may direct user content through a light-weight, auto-classifier. This may be, for example, a keyword system that flags keywords associated with prohibited content. If the content is not flagged as harmful by the auto-classifier, it is approved. However, such an auto-classifier may inaccurately flag acceptable content as being potentially harmful. If the content is flagged by the auto-classifier, that content is then reviewed by a team of human content moderators who will judge whether the content is actually inappropriate and therefore should be blocked or whether the content is not in violation of the guidelines and is approved. While this approach may be effective, it is time-consuming and costly, adding significant latency to communication of content on the platform.

[0025] Accordingly, as noted above, a technical problem is presented by the lack of a system for the detection of harmful content in a manner that is tailored to a particular platform or digital community while limiting any need for human review and the associated detrimental increase in communication latency and other factors. In other words, there is a need for an effective, streamlined and cost-effective approach to content moderation based on the specific guidelines of a particular community. As will be described below, a technical solution to these technical problems is referred to as adaptive annotation.

[0026] Adaptive annotation is a service that augments human review of potentially harmful content. More specifically, this service is designed to empower users with the ability to enhance a human review process for potentially harmful content while adhering to community-specific guidelines. The primary goal is to achieve a balance between high accuracy and improved efficiency in content moderation. By leveraging the adaptive annotation service described herein, users can create customized categories and generate annotation prompts in plain text, aligning them with the specific nuances of their community guidelines or terms of use. This proactive approach enables the annotation of content, resulting in adapted detection outcomes that are finely tuned to the unique requirements of each digital community. Users can create customized categories and generate annotation prompts according to specific community guidelines / terms of use in plain text. Then, they can receive annotation of content flagged by the adaptive annotation service for further review. In this way, the adaptive annotation service provides adapted detection results of harmful content using the generated prompts specific to the user’s guidelines or terms of service.

[0027] FIG. 1 is a flowchart illustrating an example method of detecting harmful content based on user-specific guidelines according to the principles described herein. As shown in FIG. 1, a user 100 or user group has an application or service in which content is produced for a digital community. The application or service may be, for example, a social media platform, a discussion forum, a generative artificial intelligence, or other application / service where users or generative artificial intelligence are creating content for a digital community.

[0028] In such a scenario, the user group operating the platform may want to limit harmful content from appearing on the platform. In different contexts, harmful content could refer to any content that decreases the value of the platform or enjoyment of the platform to the intended audience. For example, harmful content is content considered offensive, threatening, a breach of privacy, or otherwise inappropriate to the digital community being served.

[0029] What content is considered “harmful” will vary depending on the user group, the digital community being served, the type of content being generated and other factors. Accordingly, content that is considered harmful for one platform may not be considered harmful on another platform. For this reason, a set of guidelines that define harmful content for one platform may not work well for a second platform. Thus, it is not feasible to create a single system that monitors, in the same way, harmful content on all platforms.

[0030] As used herein, the term “user” will refer to the administrator or administrative group of a platform with responsibility for moderating the content on the platform to minimize harmful content. The term “platform” is used to refer generally to any application or service where content is produced for a community and needs to be vetted.

[0031] As shown in FIG. 1, the approach being described herein allows a user to implement their own desired guidelines in a system that will assist with moderating harmful content on their platform. Referring to FIG. 1, the user 100 determines some user guidelines 101 that indicate or define what content is to be considered harmful on the user’s platform. The adaptive annotation service illustrated will utilize the user guidelines 101 to create a customized prompt template 102. This prompt template 102 is used to submit content from the user’s platform to a Generative Artificial Intelligence (GAI) model, such as a Large Language Model (LLM) .

[0032] The prompt template 102, using the user guidelines, instructs the GAI model how to review the content from the user’s platform for harmful content according to the user’s guidelines. Using the guidance from the prompt template, the GAI can then identify content that is considered harmful specifically by the user 100 without flagging content that is not considered harmful in the context of the user guidelines 101 that the user 100 has provided.

[0033] With the prompt template 102 prepared, the system is ready to moderate content from the platform for the user. Consequently, the user 100 submits content from the platform as an annotation request 103 to have the system identify anything in the content that should be further reviewed as potentially being harmful content. The annotation request 103 could be a series of requests streaming content continuously from the platform into the review system. Alternatively, the annotation request 103 may be submitted only periodically when the user 100 or some other light weight system has identified content from the platform that is to be reviewed.

[0034] The content of the annotation request 103 is combined with the customized prompt template 102 to generate a GAI prompt 104. This prompt is the submitted to the GAI annotation service 105. In the GAI annotation service 105 and GAI model, such as an LLM, will review the content of the annotation request 103 based on the user guidelines 101 incorporated into the prompt template 102. As a result, the GAI annotation service 105 returns an annotation of the content from the request 106 to the user 100. In this annotation 106, content that should be reviewed as potentially harmful is flagged or annotated to assist the user 100 with effective and timely review of the content.

[0035] In the more detailed description below, a number of different GAI models may be used for different purposes in the adaptive annotation service being described. The principal GAI is the GAI annotation service 105 which, as prompted, reviews and annotates platform content for anything harmful. Other GAI models, perhaps with less training and operating expense, will be mentioned below for performing other functions that support the adaptive annotation service ultimately provided by the GAI annotation service 105.

[0036] FIG. 2A is a system diagram further illustrating the details of the system. In the example given above, the user started with a set of written guidelines 101 defining what content is to be considered harmful. In such an example, the customized prompt template 102 may be generated as described below.

[0037] As shown in FIG. 2A to generate the customized prompt template, the system may submit the written user guidelines 101 to a custom prompt template generator 122 which makes a call to a GAI model 120. This GAI model may be a Large or Small Language Model. In some examples, GAI 120 is less developed than other GAIs described herein. The prompt to this GAI 120 will ask the GAI to extract labels and label definitions from the written user guidelines 101 and place that information in the structure of a customized prompt template. The GAI 120 then returns a resulting customized prompt template 102 that is specifically derived from the user’s written guidelines 101. Consequently, the user guidelines 101 do not need to be prepared in a form most readily ingestible by the system. Rather, any user, regardless of technical expertise, can prepare and provide natural language guidelines 101 that will then be enforced in content moderation by the system. The system solves for the user the technical problems of implementing guidelines that have only been expressed in natural language.

[0038] An example of a customized prompt template 102, based on the example of user guidelines 101 given above, might be as follows:

[0039] #Community Standards Civility Labeling Guidelines

[0040] ##How to Apply

[0041] ###Label Set

[0042] The label set includes “Non-harmful” (Label 0) , “Real-World Sensitive Events” (Label 1) , “Violent Content and Gore” (Label 2) , “Romantic and Sexual Content” (Label 3) , “Illegal and Regulated Goods or Activities” (Label 4) , “Profanity” (Label 5) and “Political Content (Label 6) .

[0043] ###Apply all the Labels Applicable

[0044] Given the attached piece of text, apply all the labels that match portions of the content.

[0045] For example, if a section of the content fits two or more of the labels, apply all the applicable labels in an annotation of that section of the content.

[0046] ###Notes

[0047] Please only focus on the scenarios listed in the Labeling Guidelines. Other scenarios not covered by the provided guidelines are out of scope for this task

[0048] ##Definitions of Labels

[0049] ###Label 0: non-harmful

[0050] ###Label 1: Real-World Sensitive Events …

[0051] A more detailed example of this customized prompt template is shown in FIG. 3B. The content to be vetted would then be appended to the customized prompt template 102 as will be described below.

[0052] Continuing in FIG. 2A, the user 100 then provides content from the platform for moderation in the form of an annotation request 103 as described above in FIG. 1. The content of the annotation request 103 is appended to the customized prompt template 102 by a prompt generator 124 to generate a prompt 104 for the GAI annotation service 105. The resulting prompt causes the GAI annotation service 105 to review the content and label potentially harmful content according to the types of content defined from the user guidelines 101 as potentially harmful.

[0053] The prompt 104 may also instruct the GAI annotation service 105 to discard content that is approved and considered safe. In this way, the volume of content from the platform or annotation request 103 is filtered to a more manageable volume, all of which is identified as being potentially harmful and labeled with a type of harmful content that may be present. The adaptive annotation service may also indicate the severity of each labeled section. In this way, no human review of the content may be needed or at least the volume of content presented for further user review is appropriately reduced to maximize the efficiency of any human review that is implemented.

[0054] Thus, in the present system, the prompt for the GAI annotation service 105 is automatically generated from the authoritative guidelines, e.g. text or multi-modal content, per category provided by the user in a simplified way. The user does not need to craft the prompt manually. Rather, as described above, the prompt is generated from the authoritative guidelines, e.g., automatically from plain text format to LLM-understandable format per category provided by the user using the prompts described for the GAI 120. The list of labels (classes) and detailed definitions per label (class) will be automatically extracted, and the output format will be automatically declared in the prompt. Besides the harmful annotation result (label) , reasoning can also be returned for interpretability. As further described below, a confidence level can be returned for filtering to a higher precision.

[0055] As noted above, FIGs. 3A and 3B illustrate, respectively, an example of user guidelines (FIG. 3A) and the resulting customized prompt template (FIG. 3B) . The prompt sent to the GAI 120 to produce the prompt template based on the guidelines will, in this example, instruct the GAI 120 to extract a list of labels (classes) and detailed definitions per label (class) for harmful content. The prompt will also instruct the GAI 120 to include in the prompt template 102, if possible, labeling examples (few-shots) extracted from the guidelines.

[0056] The prompt to the GAI 120 may also direct that the following instructions be included in the resulting customized prompt template, to (1) declare the input format and target output format, e.g., a label, reasoning or interpretation for placing the label and a confidence score that can be used for further filtering, and (2) include principles about how to handle out-of-scope cases and / or how to handle instances of harmful content to which multiple labels apply, meaning that the content is harmful in more than one category having a definition and label .

[0057] If no labeling examples (few-shots) are provided by the user guidelines, an additional prompt to the GAI 120 can instruct the GAI 120 to summarize the description for each label, generate diverse few-shots for the summarized description per label and transform the generated few-shots into examples of target output to be included in the customized prompt template 102. As used herein, the term “few-shot” is used synonymously with “example” to refer to an example of harmful content, likely along with a label that identifies a category or type of harmful content.

[0058] In some cases, however, the user may only have a vague idea of what content might be produced on the platform and when or whether any such content should be considered harmful. For example, the user 100 may only have in mind some keywords that would indicate harmful content in the context of their platform. This is referred to as a cold start. In such a case, the adaptive annotation service will accept those keywords as the user guidelines. As above, the system will submit the keyword set to the GAI model 120. The keywords are submitted with a prompt requesting the GAI model 120 to create a set of labels and corresponding definitions of potentially harmful content based on the keywords. Thus, GAI model 120 then returns a customized prompt template, as in the example above, structured for detecting harmful content based on those keywords as the indicators of the harmful content.

[0059] In a cold start, the system may also, based on the keywords, return to the user a natural language version of guidelines that they may adopt and publish for the community. That version of the guidelines can then be edited for accuracy or expanded as needed. In that case, the revised version can be used, as described above, as the basis for a new customized prompt template 102.

[0060] Referring to FIG. 2B, with a cold start based just on keywords, the system includes a cold start handler 125. The cold start handler 125 uses a taxonomy on a Responsible AI that is purpose built for the handler 125. In this taxonomy, each node is a high-level category such as violence, while each leaf node is a detailed sub-category of the parent node such as animal abuse, school bullying, etc. For each node (leaf or non-leaf) , there is a category name and a list of synonyms of the category name together with guidelines and prompts.

[0061] By operation of the cold start handler, the input keywords are mapped to any of the categories in the taxonomy by calculating the similarity between the input and the category name synonyms. Where a keyword is mapped to the taxonomy, the details from the taxonomy (e.g., guidelines and prompts) are used to describe the corresponding content to be flagged as objectionable in the subsequent prompt to the GAI 120. If the mapping fails, then the cold start handler creates a first version of the user guidelines using a LLM prompt to summarize detailed scenarios in the input category.

[0062] FIG. 4 illustrates another example of a system according to the principles described herein for filtering content from a platform for potentially harmful content. As shown in FIG. 4, the user inputs guidelines 101 specific to the user’s platform for identifying or defining objectionable or inappropriate content, referred to collectively as harmful content. The guidelines could be provided in JSON format, in purely plain text, as described above, or some other format. The guidelines 101 may include a definition of each label the user wishes to have applied to potentially harmful content. Optionally, the guidelines 101 may also include a list of detailed scenarios for each label with examples of harmful content to which that label would apply.

[0063] A text converter in the service will leverage an LLM prompt to the GAI 120, for example as shown in FIG. 2A, to extract key information from the text guidelines 101, including a definition of each label (class) , overall principles of current category annotation, emphasis on input format and the final output format, etc. Overall principles added to the resulting output of the GAI 120 can include how to handle cases where multiple labels (classes) hit or how to handle out-of-scope cases.

[0064] If the user 100 does not or is unable to provide examples (few-shots) of the harmful content for a given label, the system will generate such examples 131 using, for example, the GAI 120 of FIG. 2A. In a cold start situation, the GAI 120 can be prompted to propose examples of harmful content based on the keywords the user has provided. A few examples for each label (class) can be created to improve the ultimate annotation quality. Thus, the system can leverage an LLM prompt to generate detailed examples that match the definition of each label and severity rating instead of just copying a description and with diverse literary forms and writing styles.

[0065] As shown in FIG. 4, after a GAI prompt is generated, jailbreak detection 132 is performed. Specifically, jailbreak detection is applied to detect if there is a jailbreak intent in the provided guidelines 101. In this context, the term jailbreak is to describe a user’s attempts to use text input to make the model break its own rules which enables the users to bypass system guardrails and generate harmful content. The system implements jailbreak detection because the user may inject harmful content into the guidelines provided to the Adaptive Annotation Service to ask the LLM used by the Adaptive Annotation Service backend to do harmful things. With jailbreak detection 132, the Adaptive Annotation Service will reject those user requests with jailbreak intention. Specifically, the Adaptive Annotation Service will leverage a third additional prompt (besides the first prompt to extract the label definitions and the second prompt to generate few-shots) to instruct GAI 120 to detect if there is jailbreak intention in the user guidelines 101. If jailbreak detected, the user request will be returned with failing status code, which means the prompt template 102 is not generated successfully.

[0066] As in the previous examples, the result of processing the guidelines 101 is a customized prompt template 102. As shown in FIG. 4, the user can then input an annotation request 103 for content to be vetted. The content is then combined with the prompt template 102 to finalize a prompt to the main GAI annotation service 105.

[0067] In the example of FIG. 4, the prompt may also be compressed 133. Specifically, the generated prompt is compressed to reduce tokens so as to save annotation cost and reduce latency while avoiding significant quality decay. Prompt compression will be described in further detail below. In addition to identifying potentially harmful content in the annotation result, the prompt template can also request that the GAI annotation service 105 return reasoning in support of why content is flagged as potentially harmful to facilitate interpretability of the GAI output, i.e., the annotation of content from the annotation request 106.

[0068] The adaptive annotation service also supports dynamic few-shot selection 134 to improve annotation quality. Specifically, as noted above, there may be a pre-labeled data set containing a number of few-shot candidates provided by user 100 which may be collected and accumulated gradually in an account for that user. For each batch of sampled content in an annotation request 103, the system will dynamically select a few examples which are semantically relevant to the input content from the few-shot candidate set 134 to include in the prompt to the GAI annotation service 105 to guide the annotation of content from the request 106. Semantic similarity and relevance can be calculated based on any pairwise similarity algorithm including but not limited to cosine similarity between embedding vectors of two pieces of text, Pyramid Matching, etc. Diversity (of labels) will also be considered when selecting relevant examples. The adaptive annotation service, at function 134, will randomly pick some from those whose similarity with current input text is above a threshold. More few-shot candidates can be added to the few-shot corpora after each round of evaluation and iteration. Multi-modal input will be supported. Chat and non-chat inputs are supported.

[0069] For example, the user may be aware of phrases or ideas that are regularly expressed on the platform being moderated and which generally have a similar structure. The user may record these examples or few-shots as a corpus of prelabeled data. For example, the phrase “the game is amazing…” may be tagged with Label 0, indicating innocuous content. Similarly, a sentence such as: “The gameplay mechanics are smooth and intuitive, creating an enjoyable gaming experience for both beginners and experiences players” may also be tagged with Label 0. As described above, the system may analyze the content of the request to be vetted for these or semantically similar phrases.

[0070] When such a similar phrase is found, the prompt to the GAI annotation service 105 can be dynamically updated to include the matching examples with the Label 0. Consequently, the GAI annotation service 105 can readily and correctly tag those matching portions of the content to be vetted with Label 0. For example, assume that the content to be vetted includes the sentence: “I love this game. ” This phrase would be identified as semantically similar to the example: “The game is amazing. ” Accordingly, the example of “The game is amazing (Label 0) ” would be added to the prompt to the GAI annotation service 105. Consequently, the phrase “I love this game, ” would be labeled with Label 0 in the output to the user from GAI annotation service 105. Specifically, the Adaptive Annotation Service will be more likely to label “I love this game” with Label 0 if its similar examples have Label 0 from the few-shot candidate corpora. However, the Adaptive Annotation Service can provide the output based on both the guidelines and the few-shots.

[0071] The few-shot examples provided by the user can also be phrases that include harmful content with a corresponding label so that any such similar phrases in the content to be vetted cause those examples and their labels to be included in the prompt to the GAI annotation service 105. As indicated above, if there are no few-shots provided by the user, the initial GAI 120 can be prompted to produce some few-shots based on the user’s keywords for harmful content. In either case, the user can add to the few-shots corpus 141 over time.

[0072] As also shown in FIG. 4, an offline evaluation 135 may be conducted of the annotation of content from the request 106. If the system is not optimally identifying content that the user 100 wants to flag as harmful, the user guidelines 101 can be updated accordingly and the process of Fig. 4 can be iterated to improve the customized prompt template 102 and the resulting output of the GAI annotation service 105.

[0073] The compression of the prompt to the GAI annotation service 105 is explained with reference to FIGs. 5A and 5B. As will be appreciated, the larger the amount of data in the prompt to the GAI annotation service 105, the more bandwidth will be required to transmit the prompt and the more time and processing resources will be required at the GAI annotation service 105 to respond to the prompt. However, some parts of the prompt may be superfluous in a given instance and thus contribute unnecessarily to the demands on bandwidth, processing power and time. For example, depending on the content in the request 103 to be reviewed, some of the instructions in the customized prompt template 102 may simply not apply and can be omitted. Additionally, some of the content in the request 103 to be reviewed may be preemptively eliminated from the review process without a significant risk of missing harmful content. Thus, two steps can be performed in prompt compression 133. The compression 133 is performed based on information retrieval or probability theory.

[0074] Referring to FIG. 5A, as the first step, the prompt compressor 140 will segment the prompt into several pieces and tag which pieces of prompt segments are relevant to current input using another LLM prompt and feeding into another cheaper LLM. Only relevant parts will be retained and others will be removed from current input.

[0075] For example, the different label definitions along with the content of the request can be fed into a smaller, cheaper GAI 170, for example, an LLM, with the instruction to identify which of the label definitions might apply to the content to be vetted. As shown in FIG. 5A, labels 0-N may be defined in the prompt template 102. However, given the specific content in the current request that is to be reviewed for harmful content, only labels 1, 4 and 7 appear to apply. Accordingly, referring again to FIG. 4, the content to be reviewed 103 and the customized prompt template 102 are joined. Then, to compress the prompt, the call to the lighter GAI 170 is made. Based on the result, the instructions regarding all the labels, except 1, 4 and 7 are removed to compress the prompt before the prompt is submitted to the GAI annotation service 105.

[0076] As the second step, based on the compressed prompt from the first step, further compression is achieved on a smaller granularity. Specifically, the prompt compressor 140 will remove those words whose conditional probability is high given its context. More specifically, the system uses a pre-calculated table regarding n-gram probability (i.e. the probability of the n-th word given the first (n-1) words sequentially) .

[0077] LLMs operate, in general, by determining a word most likely to follow a previous word based on a massive training set and the input prompt as the starting point. In a similar approach, this phase of the prompt compressor 140 will calculate a conditional probability of each word or phrase in the content to be vetted given its context and based on the pre-calculated table regarding n-gram probability or using a small-scale GAI or LLM.

[0078] In the example of FIG. 5B, a first phrase is “I usually. ” In the given context, the prompt compressor 140 determines that the probability of the word “usually” following the word “I” is 90% (0.9) . Accordingly, the word “usually” can be omitted from the prompt to further compress the prompt without significantly impairing the ability of the service to detect harmful content. In the second part of the example of FIG. 5B, the input phrase is “create a bomb. ” As shown, the article “a” has a 0.87 probability of following “create. ” Accordingly, the word “a” can be eliminated to further compress the prompt. The probabilities, in the given context, for “create” and “bomb” are much lower (0.3 and 0.01, respectively) . Accordingly, neither word is omitted by the prompt compressor 140. A threshold can be set for the probability score at which a word is omitted from the prompt by the prompt compressor 140.

[0079] FIG. 6 depicts a process for providing labels for the few-shots according to principles described herein. As noted above, the user may provide labeled few-shots to assist the system in annotating the content to be moderated. Alternatively, if the user has not provided those few-shots, but has provided written community guidelines, the system can generate a number of few-shots based on the written guidelines. Or, as described above, if the user has only provide keywords and not written guidelines, the system may generate a number of few-shots using the keywords provided by the user. In these cases, the few-shots generated will need to be labeled. In other cases, the user may provide a few-shot (example) of harmful content, but not provide a label as to the type of harmful content, or the system may seek to verify a label that the user has provided for an example.

[0080] Consequently, as shown in FIG. 6, the system has a number of few-shots to be used to augment the prompt to the GAI annotation service 105. Each of these few-shots may be associated with a label identifying the type of harmful content that it typifies or such a label may be missing. If the few-shot is auto-generated, it will not necessarily have a label. Alternatively, the user may have omitted labeling the example.

[0081] Accordingly, for each few-shot, the system will leverage multiple models 161 to categorize or label it. These models 161 may be LLMs, likely smaller, cheaper models, that are prompted with a list of possible labels, each defined by a different type of harmful content, and an instruction to assign one of the labels to each of the examples from the collection 141.

[0082] A level of confidence in each such resulting level is determined by the voting of the different models. In this context, voting refers to how many models make the same prediction. In the example of FIG. 6, one of the few-shots is “You’re a dead man, Mr. Bond. ” When submitted to the voting models 161, one of three (1 / 3) predict that this content is objectionable and corresponds to Label 1. However, two of the three (2 / 3) models predict that this content is not harmful and is Labeled 0. The system could adopt the majority result, Label 0. Alternatively, the system may seek to avoid missing any potentially harmful content and may adopt Label 1 because one of the models so indicated.

[0083] In either case, the labeling may be associated with a confidence level. In the example of FIG. 6, the example is given Label 1, but with a note that this designation is with low confidence. The example, the label and the confidence level can then be incorporated into a prompt to the GAI annotation service 105 to finetune that prompt.

[0084] In another example of FIG. 6, the example in the collection 141 is “In GTA5, killing animals is a great way to make money quickly. ” This example does not have a label, likely because the user omitted applying a label when providing the example. When examples or few-shots are autogenerated, they are generated per label and thus have a label. In the illustrated example, when submitted to the voting models 161, the result is that all three (3 / 3) predict the statement to be harmful content associated with the Label 1 type of such content. Accordingly, this example is given Label 1 with full confidence. Again, the example, the label and the confidence level can then be incorporated into a prompt to the GAI annotation service 105 to finetune that prompt.

[0085] Different voting results can indicate the difficulty of inference. The system can map the confidence to a reasoning statement in natural language for LLM, as in the example of FIG. 6. This reasoning statement will be more readily processed by the GAI annotation service 105 as compared to a numeric confidence rating.

[0086] As noted above, when vetting a new input text for harmful content, the system will dynamically leverage the few-shots techniques described above to select a subset of relevant few-shots. The GAI annotation service 105 will then learn from the relevant few-shots about the label as well as the confidence (difficulty) of the prediction, and output the label and confidence when annotating the current input text. With the output confidence, then the user can filter the output by the confidence. For example, if the user wants to improve the precision (whether the predicted harmful input is really harmful traffic) , then he or she can focus on indications of harmful content made with low or no confidence and can provide additional examples or labeling to supplement where the system is lacking confidence in a labeling determination.

[0087] Again, there may be different approaches to the voting process. If the majority of models agree on the prediction for the few-shot text, then the system may assign the majority prediction and convey the confidence by adding a suffix in natural language such as “the prediction is with confidence” or “with full confidence” if the voting is unanimous. If the different models are evenly split on their predictions for the few-shot, this indicates the text is difficult for inference, the system can then assign one of the labels with moderate or low confidence. If a label of the few-shot is provided by user, then the system can compare the predictions by different models to the ground truth from the user. If the majority of models get the same prediction as the user’s ground truth, then the confidence statement may be “the prediction is with full confidence. ” If the majority of models get the wrong prediction compared to the user’s label, then the system may include a phrase such as “the prediction is with low confidence. ” If no models get the same prediction as the user’s label, then “the prediction is with no confidence. ”

[0088] Models participating in voting can be any light-weight models trained previously or any previous versions of prompts from each iteration. If there is a corpus of few-shot candidates, the LLM service  / application owner can also set them as the evaluation set and as the release gate of prompt iteration. They can try to update the guidelines for the same category, and decide to keep a best version from all the historical versions (not necessarily the final version) based on the agreement between the labels and prompt annotation results on the evaluation set.

[0089] FIG. 7 is a block diagram 700 illustrating an example software architecture 702, various portions of which may be used in conjunction with various hardware architectures herein described, which may implement any of the above-described features. FIG. 7 is a non-limiting example of a software architecture, and it will be appreciated that many other architectures may be implemented to facilitate the functionality described herein. The software architecture 702 may execute on hardware such as a machine 800 of FIG. 8 that includes, among other things, processors 810, memory 830, and input / output (I / O) components 850. A representative hardware layer 704 is illustrated and can represent, for example, the machine 800 of FIG. 8. The representative hardware layer 704 includes a processing unit 706 and associated executable instructions 708. The executable instructions 708 represent executable instructions of the software architecture 702, including implementation of the methods, modules and so forth described herein. The hardware layer 704 also includes a memory / storage 710, which also includes the executable instructions 708 and accompanying data. The hardware layer 704 may also include other hardware modules 712. Instructions 708 held by processing unit 706 may be portions of instructions 708 held by the memory / storage 710.

[0090] The example software architecture 702 may be conceptualized as layers, each providing various functionality. For example, the software architecture 702 may include layers and components such as an operating system (OS) 714, libraries 716, frameworks 718, applications 720, and a presentation layer 744. Operationally, the applications 720 and / or other components within the layers may invoke API calls 724 to other layers and receive corresponding results 726. The layers illustrated are representative in nature and other software architectures may include additional or different layers. For example, some mobile or special purpose operating systems may not provide the frameworks / middleware 718.

[0091] The OS 714 may manage hardware resources and provide common services. The OS 714 may include, for example, a kernel 728, services 730, and drivers 732. The kernel 728 may act as an abstraction layer between the hardware layer 704 and other software layers. For example, the kernel 728 may be responsible for memory management, processor management (for example, scheduling) , component management, networking, security settings, and so on. The services 730 may provide other common services for the other software layers. The drivers 732 may be responsible for controlling or interfacing with the underlying hardware layer 704. For instance, the drivers 732 may include display drivers, camera drivers, memory / storage drivers, peripheral device drivers (for example, via Universal Serial Bus (USB) ) , network and / or wireless communication drivers, audio drivers, and so forth depending on the hardware and / or software configuration.

[0092] The libraries 716 may provide a common infrastructure that may be used by the applications 720 and / or other components and / or layers. The libraries 716 typically provide functionality for use by other software modules to perform tasks, rather than rather than interacting directly with the OS 714. The libraries 716 may include system libraries 734 (for example, C standard library) that may provide functions such as memory allocation, string manipulation, file operations. In addition, the libraries 716 may include API libraries 736 such as media libraries (for example, supporting presentation and manipulation of image, sound, and / or video data formats) , graphics libraries (for example, an OpenGL library for rendering 2D and 3D graphics on a display) , database libraries (for example, SQLite or other relational database functions) , and web libraries (for example, WebKit that may provide web browsing functionality) . The libraries 716 may also include a wide variety of other libraries 738 to provide many functions for applications 720 and other software modules.

[0093] The frameworks 718 (also sometimes referred to as middleware) provide a higher-level common infrastructure that may be used by the applications 720 and / or other software modules. For example, the frameworks 718 may provide various graphic user interface (GUI) functions, high-level resource management, or high-level location services. The frameworks 718 may provide a broad spectrum of other APIs for applications 720 and / or other software modules.

[0094] The applications 720 include built-in applications 740 and / or third-party applications 742. Examples of built-in applications 740 may include, but are not limited to, a contacts application, a browser application, a location application, a media application, a messaging application, and / or a game application. Third-party applications 742 may include any applications developed by an entity other than the vendor of the particular platform. The applications 720 may use functions available via OS 714, libraries 716, frameworks 718, and presentation layer 744 to create user interfaces to interact with users.

[0095] Some software architectures use virtual machines, as illustrated by a virtual machine 748. The virtual machine 748 provides an execution environment where applications / modules can execute as if they were executing on a hardware machine (such as the machine 800 of FIG. 8, for example) . The virtual machine 748 may be hosted by a host OS (for example, OS 714) or hypervisor, and may have a virtual machine monitor 746 which manages operation of the virtual machine 748 and interoperation with the host operating system. A software architecture, which may be different from software architecture 702 outside of the virtual machine, executes within the virtual machine 748 such as an OS 750, libraries 752, frameworks 754, applications 756, and / or a presentation layer 758.

[0096] FIG. 8 is a block diagram illustrating components of an example machine 800 configured to read instructions from a machine-readable medium (for example, a machine-readable storage medium) and perform any of the features described herein. The example machine 800 is in the form of a computer system, within which instructions 816 (for example, in the form of software components) for causing the machine 800 to perform any of the features described herein may be executed.

[0097] As such, the instructions 816 may be used to implement modules or components described herein. The instructions 816 cause unprogrammed and / or unconfigured machine 800 to operate as a particular machine configured to carry out the described features. The machine 800 may be configured to operate as a standalone device or may be coupled (for example, networked) to other machines. In a networked deployment, the machine 800 may operate in the capacity of a server machine or a client machine in a server-client network environment, or as a node in a peer-to-peer or distributed network environment. Machine 800 may be embodied as, for example, a server computer, a client computer, a personal computer (PC) , a tablet computer, a laptop computer, a netbook, a set-top box (STB) , a gaming and / or entertainment system, a smart phone, a mobile device, a wearable device (for example, a smart watch) , and an Internet of Things (IoT) device. Further, although only a single machine 800 is illustrated, the term “machine” includes a collection of machines that individually or jointly execute the instructions 816.

[0098] The machine 800 may include processors 810, memory 830, and I / O components 850, which may be communicatively coupled via, for example, a bus 802. The bus 802 may include multiple buses coupling various elements of machine 800 via various bus technologies and protocols. In an example, the processors 810 (including, for example, a central processing unit (CPU) , a graphics processing unit (GPU) , a digital signal processor (DSP) , an ASIC, or a suitable combination thereof) may include one or more processors 812a to 812n that may execute the instructions 816 and process data. In some examples, one or more processors 810 may execute instructions provided or identified by one or more other processors 810. The term “processor” includes a multi-core processor including cores that may execute instructions contemporaneously. Although FIG. 8 shows multiple processors, the machine 800 may include a single processor with a single core, a single processor with multiple cores (for example, a multi-core processor) , multiple processors each with a single core, multiple processors each with multiple cores, or any combination thereof. In some examples, the machine 800 may include multiple processors distributed among multiple machines.

[0099] The memory / storage 830 may include a main memory 832, a static memory 834, or other memory, and a storage unit 836, both accessible to the processors 810 such as via the bus 802. The storage unit 836 and memory 832, 834 store instructions 816 embodying any one or more of the functions described herein. The memory / storage 830 may also store temporary, intermediate, and / or long-term data for processors 810. The instructions 816 may also reside, completely or partially, within the memory 832, 834, within the storage unit 836, within at least one of the processors 810 (for example, within a command buffer or cache memory) , within memory at least one of I / O components 850, or any suitable combination thereof, during execution thereof. Accordingly, the memory 832, 834, the storage unit 836, memory in processors 810, and memory in I / O components 850 are examples of machine-readable media.

[0100] As used herein, “machine-readable medium” refers to a device able to temporarily or permanently store instructions and data that cause machine 800 to operate in a specific fashion, and may include, but is not limited to, random-access memory (RAM) , read-only memory (ROM) , buffer memory, flash memory, optical storage media, magnetic storage media and devices, cache memory, network-accessible or cloud storage, other types of storage and / or any suitable combination thereof. The term “machine-readable medium” applies to a single medium, or combination of multiple media, used to store instructions (for example, instructions 816) for execution by a machine 800 such that the instructions, when executed by one or more processors 810 of the machine 800, cause the machine 800 to perform and one or more of the features described herein. Accordingly, a “machine-readable medium” may refer to a single storage device, as well as “cloud-based” storage systems or storage networks that include multiple storage apparatus or devices. The term “machine-readable medium” excludes signals per se.

[0101] The I / O components 850 may include a wide variety of hardware components adapted to receive input, provide output, produce output, transmit information, exchange information, capture measurements, and so on. The specific I / O components 850 included in a particular machine will depend on the type and / or function of the machine. For example, mobile devices such as mobile phones may include a touch input device, whereas a headless server or IoT device may not include such a touch input device. The particular examples of I / O components illustrated in FIG. 8 are in no way limiting, and other types of components may be included in machine 800. The grouping of I / O components 850 are merely for simplifying this discussion, and the grouping is in no way limiting. In various examples, the I / O components 850 may include user output components 852 and user input components 854. User output components 852 may include, for example, display components for displaying information (for example, a liquid crystal display (LCD) or a projector) , acoustic components (for example, speakers) , haptic components (for example, a vibratory motor or force-feedback device) , and / or other signal generators. User input components 854 may include, for example, alphanumeric input components (for example, a keyboard or a touch screen) , pointing components (for example, a mouse device, a touchpad, or another pointing instrument) , and / or tactile input components (for example, a physical button or a touch screen that provides location and / or force of touches or touch gestures) configured for receiving various user inputs, such as user commands and / or selections.

[0102] In some examples, the I / O components 850 may include biometric components 856, motion components 858, environmental components 860, and / or position components 862, among a wide array of other physical sensor components. The biometric components 856 may include, for example, components to detect body expressions (for example, facial expressions, vocal expressions, hand or body gestures, or eye tracking) , measure biosignals (for example, heart rate or brain waves) , and identify a person (for example, via voice-, retina-, fingerprint-, and / or facial-based identification) . The motion components 858 may include, for example, acceleration sensors (for example, an accelerometer) and rotation sensors (for example, a gyroscope) . The environmental components 860 may include, for example, illumination sensors, temperature sensors, humidity sensors, pressure sensors (for example, a barometer) , acoustic sensors (for example, a microphone used to detect ambient noise) , proximity sensors (for example, infrared sensing of nearby objects) , and / or other components that may provide indications, measurements, or signals corresponding to a surrounding physical environment. The position components 862 may include, for example, location sensors (for example, a Global Position System (GPS) receiver) , altitude sensors (for example, an air pressure sensor from which altitude may be derived) , and / or orientation sensors (for example, magnetometers) .

[0103] The I / O components 850 may include communication components 864, implementing a wide variety of technologies operable to couple the machine 800 to network (s) 870 and / or device (s) 880 via respective communicative couplings 872 and 882. The communication components 864 may include one or more network interface components or other suitable devices to interface with the network (s) 870. The communication components 864 may include, for example, components adapted to provide wired communication, wireless communication, cellular communication, Near Field Communication (NFC) , Bluetooth communication, Wi-Fi, and / or communication via other modalities. The device (s) 880 may include other machines or various peripheral devices (for example, coupled via USB) .

[0104] In some examples, the communication components 864 may detect identifiers or include components adapted to detect identifiers. For example, the communication components 864 may include Radio Frequency Identification (RFID) tag readers, NFC detectors, optical sensors (for example, one-or multi-dimensional bar codes, or other optical codes) , and / or acoustic detectors (for example, microphones to identify tagged audio signals) . In some examples, location information may be determined based on information from the communication components 864, such as, but not limited to, geo-location via Internet Protocol (IP) address, location via Wi-Fi, cellular, NFC, Bluetooth, or other wireless station identification and / or signal triangulation.

[0105] While various embodiments have been described, the description is intended to be exemplary, rather than limiting, and it is understood that many more embodiments and implementations are possible that are within the scope of the embodiments. Although many possible combinations of features are shown in the accompanying figures and discussed in this detailed description, many other combinations of the disclosed features are possible. Any feature of any embodiment may be used in combination with or substituted for any other feature or element in any other embodiment unless specifically restricted. Therefore, it will be understood that any of the features shown and / or discussed in the present disclosure may be implemented together in any suitable combination. Accordingly, the embodiments are not to be restricted except in light of the attached claims and their equivalents. Also, various modifications and changes may be made within the scope of the attached claims.

[0106] Generally, functions described herein (for example, the features illustrated in FIGS. 1-6) can be implemented using software, firmware, hardware (for example, fixed logic, finite state machines, and / or other circuits) , or a combination of these implementations. In the case of a software implementation, program code performs specified tasks when executed on a processor (for example, a CPU or CPUs) . The program code can be stored in one or more machine-readable memory devices. The features of the techniques described herein are system-independent, meaning that the techniques may be implemented on a variety of computing systems having a variety of processors. For example, implementations may include an entity (for example, software) that causes hardware to perform operations, e.g., processors functional blocks, and so on. For example, a hardware device may include a machine-readable medium that may be configured to maintain instructions that cause the hardware device, including an operating system executed thereon and associated hardware, to perform operations. Thus, the instructions may function to configure an operating system and associated hardware to perform the operations and thereby configure or otherwise adapt a hardware device to perform functions described above. The instructions may be provided by the machine-readable medium through a variety of different configurations to hardware elements that execute the instructions.

[0107] In the foregoing detailed description, numerous specific details were set forth by way of examples in order to provide a thorough understanding of the relevant teachings. It will be apparent to persons of ordinary skill, upon reading the description, that various aspects can be practiced without such details. In other instances, well known methods, procedures, components, and / or circuitry have been described at a relatively high-level, without detail, in order to avoid unnecessarily obscuring aspects of the present teachings.

[0108] While the foregoing has described what are considered to be the best mode and / or other examples, it is understood that various modifications may be made therein and that the subject matter disclosed herein may be implemented in various forms and examples, and that the teachings may be applied in numerous applications, only some of which have been described herein. It is intended by the following claims to claim any and all applications, modifications and variations that fall within the true scope of the present teachings.

[0109] Unless otherwise stated, all measurements, values, ratings, positions, magnitudes, sizes, and other specifications that are set forth in this specification, including in the claims that follow, are approximate, not exact. They are intended to have a reasonable range that is consistent with the functions to which they relate and with what is customary in the art to which they pertain.

[0110] The scope of protection is limited solely by the claims that now follow. That scope is intended and should be interpreted to be as broad as is consistent with the ordinary meaning of the language that is used in the claims when interpreted in light of this specification and the prosecution history that follows, and to encompass all structural and functional equivalents. Notwithstanding, none of the claims are intended to embrace subject matter that fails to satisfy the requirement of Sections 101, 102, or 103 of the Patent Act, nor should they be interpreted in such a way. Any unintended embracement of such subject matter is hereby disclaimed.

[0111] Except as stated immediately above, nothing that has been stated or illustrated is intended or should be interpreted to cause a dedication of any component, step, feature, object, benefit, advantage, or equivalent to the public, regardless of whether it is or is not recited in the claims.

[0112] It will be understood that the terms and expressions used herein have the ordinary meaning as is accorded to such terms and expressions with respect to their corresponding respective areas of inquiry and study except where specific meanings have otherwise been set forth herein.

[0113] Relational terms such as first and second and the like may be used solely to distinguish one entity or action from another without necessarily requiring or implying any actual such relationship or order between such entities or actions. The terms “comprises, ” “comprising, ” and any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element preceded by “a” or “an” does not, without further constraints, preclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0114] The Abstract of the Disclosure is provided to allow the reader to quickly identify the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various examples for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that any claim requires more features than the claim expressly recites. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed example. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.

Claims

1.A data processing system comprising:a processor (810) , anda memory (830) storing executable instructions which, when executed by the processor, causes the processor, alone or in combination with other processors (810) , to perform the following functions:receive user input (101) defining what content on a particular platform is to be blocked as harmful content;based on the user input (101) , generate a customized prompt template (102) for submission to a Generative Artificial Intelligence (GAI) annotation service (105) , the customized prompt template (102) instructing the GAI annotation service (105) to label harmful content within an annotation request (103) according to the user input defining what content on a particular platform is to be blocked as harmful content;receive an annotation request (103) comprising platform content to be vetted for harmful content based on the user input;generate a prompt (104) for the GAI annotation service (105) by combining the platform content with the customized prompt template (102) ;compress the prompt (104) prior to submission to the GAI annotation service (105) ; andgenerate, with the prompt (104) , an identification of harmful content that is to be blocked in the platform content from the GAI annotation service (105) based specifically on the user input defining what content is harmful content.2.The data processing system of claim 1, wherein the customized prompt template (102) includes a number of examples of harmful content.3.The data processing system of claim 2, wherein the examples of harmful content are generated by a GAI (120) based on the user input (101) .4.The data processing system of claim 2, wherein the examples of harmful content are each labeled based on a set of labels, each label applying to a different type of harmful content and having an associated definition of that type of harmful content based on the user input (101) .5.The data processing system of claim 2, wherein a GAI (120) is called to associate a label from a set of labels with each of the examples, each label applying to a different type of harmful content and having an associated definition of that type of harmful content based on the user input (101) .6.The data processing system of claim 5, wherein:the GAI (120) comprises multiple models (161) ; anda particular example in the customized prompt template (102) is associated with a label from the set of labels and a confidence statement based on output of the multiple models (161) .7.The data processing system of claim 6, wherein, when the multiple models (161) disagree about which label to associate with the particular example, the confidence statement associated with that particular example indicates a decreased confidence in the associated label.8.The data processing system of claim 6, wherein a label to be associated with the particular example is selected based on output that agrees from a majority of the multiple models (161) .9.The data processing system of any of claims 1-8, wherein the customized prompt template (102) is generated by a call to a GAI (120) based on the user input (101) with a request that the GAI (120) generate the customized prompt template (102) from the user input.10.The data processing system of claim 9, wherein the user input comprises descriptions of different types of harmful content to be blocked in the platform content, the call to the GAI (120) to generate the customized prompt template (102) instructing the GAI (120) to define a set of labels, each label applying to a different type of harmful content and having an associated definition of that type of harmful content based on the user input.11.The data processing system of claim 10, wherein the user input (101) only comprises keywords identifying the different types of harmful content.12.The data processing system of any of claims 1-11, wherein the customized prompt template (102) comprises a set of labels, each label applying to a different type of harmful content and having an associated definition of that type of harmful content based on the user input (101) .13.The data processing system of claim 12, wherein the identification of harmful content comprises an annotation of the platform content (106) that includes application of the set of labels to corresponding harmful content in the platform content.14.The data processing system of any of claims 1-13, further comprising executable instructions causing the processor (810) to block identified harmful content on the platform based on the identification of the harmful content.15.The data processing system of any of claims 1-14, wherein:the user input (101) provides or is used to generate a set of labels for different types of harmful content to be identified in the platform content; andcompressing the prompt (104) comprises making a call to a GAI (120) with a request (106) to identify a subset of the set of labels that apply to the platform content and removing information from the prompt (104) about labels not identified by the GAI (120) as applying to the platform content.16.The data processing system of any of claims 1-15, wherein compressing the prompt (104) comprises determining a probability for each successive word in the platform content and deleting words from the prompt (104) with a probability above a threshold.17.A method of blocking harmful content from platform content of a platform for a particular electronic community, the method comprising:receiving user input (101) defining what content on a particular platform is to be blocked as harmful content;generating, based on the user input (101) , a customized prompt template (102) for submission to a Generative Artificial Intelligence (GAI) (120) annotation service, the customized prompt template (102) instructing the GAI annotation service (105) to label harmful content within an annotation request (103) according to the user input defining what content on a particular platform is to be blocked as harmful content;receiving an annotation request (103) comprising platform content to be vetted for any harmful content based on the user input;generating a prompt (104) for the GAI annotation service (105) by combining the platform content with the customized prompt template (102) ; andreceiving identification of harmful content that is to be blocked in the platform content from the GAI annotation service (105) based specifically on the user input defining what content is harmful content.18.The method of claim 17, further comprising generating the customized prompt template (102) with a call to a GAI (120) based on the user input and with a request (106) that the GAI (120) generate the customized prompt template (102) from the user input.19.A system for use in blocking harmful content from a platform serving an electronic community, the community having guidelines as to what constitutes harmful content, the system comprising:a processor (810) in communication with the platform to receiving platform content to be vetted for harmful content;a memory (830) storing executable instructions which, when executed by the processor, causes the processor, alone or in combination with other processors (810) , to perform the following functions:generate a customized prompt template (102) , based on the guidelines, for submission to a Generative Artificial Intelligence (GAI) annotation service (105) , the customized prompt template (102) instructing the GAI annotation service (105) to label harmful content within an annotation request (103) according to the guidelines;receive an annotation request (103) comprising platform content to be vetted for any harmful content;generate a prompt (104) for the GAI annotation service (105) by combining the platform content with the customized prompt template (102) ; andgenerate, with the prompt (104) , an identification of harmful content that is to be blocked in the platform content from the GAI annotation service (105) based specifically on the guidelines defining what content is harmful content.20.The system of claim 19, further comprising executable instructions causing the processor to compress the prompt (104) prior to submission to the GAI annotation service (105).