Dynamic batch processing of cue query
By using meta-model topology and dynamic batch processing technology, the problems of computational resource overload and content moderation in machine learning model systems are solved, enabling efficient and flexible multi-model collaboration and content moderation, thereby improving user experience and computational efficiency.
Patent Information
- Application Number
- CN202480030592.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-12
- Filing Date
- 2024-06-11
- Publication Date
- 2025-12-02
AI Technical Summary
Existing machine learning model systems suffer from computational resource overload when processing multiple input prompts, resulting in complex user experiences, difficulties in model coordination, and challenges in content review, especially in the context of interconnected frameworks for multiple machine learning models, which can easily generate inappropriate content.
By employing a meta-model topology and dynamic batch processing technology, and configuring batch processing criteria for different processing nodes, the model graph generation and review process is simplified, enabling distributed deployment and efficient computing.
It improves computational efficiency, simplifies user experience, reduces the risk of inappropriate content generation, and enables flexible collaboration among multi-model frameworks and efficient content moderation.
Smart Images

Figure CN121058014A_ABST
Abstract
Description
Background Technology
[0001] A typical machine learning model is configured to receive an input cue word and generate output based on that cue word. The output generated by the machine learning model is produced after applying the layers of the machine learning model to that input cue word.
[0002] In some cases, large language models (LLMs) (such as ChatGPT) are also configured to generate output based on an initial seed cue and intermediate cue words derived from the output associated with that initial cue, and these intermediate cue words can be part of an interactive dialogue between the user and the LLM interface. For example, a typical system might apply an LLM to an initial input cue word and generate preliminary output, which is further used to construct derived intermediate cue words. Subsequently, new intermediate outputs can be generated by applying the LLM to the newly derived intermediate cue words. This loop can continue as part of an interactive dialogue until the desired final output is generated, based on the complete set of initial and intermediate cue words processed by the LLM.
[0003] In some cases, an LLM will include multiple independent but interconnected machine learning models, or at least enable users to interact with these models, which are applied to different prompts being processed. However, as the number and complexity of the machine learning models used to process input prompts increase, the user experience (including the writing and interrelation of different prompts) becomes more complex and cumbersome, as users must switch between different models used to process various input prompts and between model interfaces.
[0004] Conventional LLMs, and the network of models (including interconnected sets of machine learning models) that further process and analyze the outputs from LLMs, typically serve a large and diverse user base across different enterprises. Therefore, computer hardware (such as central processing units (CPUs), graphics processing units (GPUs), and neural processing units (NPUs)) that handle the workloads of different models can easily become overloaded when processing multiple input prompts from different users.
[0005] It should also be noted that existing models are constantly evolving. As new models are developed and their interrelationships with existing models become more complex, coordinating and integrating the functionality of different models within a conventional system's framework becomes quite difficult.
[0006] In some cases, machine learning models may also produce outputs containing inappropriate content, such as hate speech, content that may offend some users, or other types of inappropriate content that may violate specific corporate policies. This is especially true for systems that utilize frameworks that interconnect multiple machine learning models that consider different sets of parameters and link intermediate outputs and prompts from different models together.
[0007] In light of the above, there is a need to improve the systems and methods for generating, deploying, and auditing artificial intelligence frameworks, including large language models and / or other types of machine learning models for generative tasks; and in particular, there is a need to improve the systems and methods for facilitating the auditing of content processed and generated by multiple machine learning model frameworks.
[0008] The subject matter claimed herein is not limited to embodiments that address any shortcomings or operate solely in environments such as those described above. Rather, this background information is merely intended to illustrate an example technical field in which certain embodiments described herein can be practiced. Summary of the Invention
[0009] The disclosed embodiments include, or can be used, to deploy and modify metamodel topologies and to plan the generation of model graphs based on metamodel topologies. The disclosed embodiments also include methods and systems for processing prompts and generating corresponding outputs using model graphs derived from metamodel topologies.
[0010] Some disclosed embodiments relate to systems and methods for performing dynamic batching of data processing requests for a machine learning model. The machine learning model includes different processing nodes configured to perform different functions on input data, such as data prompts. Each node includes batching criteria for controlling the number of processing requests to be batched together and / or the timing of submitting the batched processing requests. Each node's batching criteria include a threshold that includes the number of processing requests and / or the duration of the batch, when the threshold is exceeded, triggering the scheduling and transmission of the batched processing requests.
[0011] During runtime, the system identifies different processing requests it receives and routes them to different batch processing caches or queues, each corresponding to a node assigned to process a different request. The system prohibits the transmission of batches of processing requests from the cache or queue until the batch processing criteria for a particular node are met. For example, in response to determining that the batch processing criteria for a specific node are not met, the system prohibits the transfer of cached processing requests from the batch request cache to that specific node. Alternatively, or subsequently, in response to determining that the batch processing criteria for a specific node or its corresponding batch processing queue have been met, the system triggers the transmission and routing of processing requests from the batch processing queue to the node.
[0012] This summary is provided to introduce, in a simplified form, some concepts that will be further described in the following detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to help determine the scope of the claimed subject matter. Attached Figure Description
[0013] To illustrate how the advantages and features of the systems and methods described herein are obtained, the embodiments briefly described above will be described in more detail with reference to the specific embodiments shown in the accompanying drawings. It should be understood that these drawings depict only typical embodiments of the systems and methods described herein and should not be considered as limiting their scope; certain systems and methods will be described and explained with additional features and details using the drawings, in which:
[0014] Figure 1 An example computer environment is shown, in which a computing system is integrated and / or used to perform one or more aspects of the disclosed embodiments.
[0015] Figure 2 An example computer architecture is shown that configures a meta-model topology as a content moderation system.
[0016] Figures 3-13 The diagram illustrates the data flow between components in an example computer architecture designed to facilitate the use of metamodel topologies in client-oriented systems.
[0017] Figures 14-17 An example process is shown for generating and pruning model graph instances derived from the metamodel topology.
[0018] Figure 18 Example implementations are shown for generating and executing model graphs for each segment of the complete prompt word.
[0019] Figures 19-20An example embodiment of batching data processing requests at one or more nodes in a model graph is shown.
[0020] Figure 21 An example embodiment is shown with a flowchart of multiple actions associated with a method for deploying a metamodel topology.
[0021] Figure 22 An example embodiment is shown with a flowchart of multiple actions associated with a method for processing data using a metamodel topology.
[0022] Figure 23 An example embodiment is shown with a flowchart of multiple actions associated with a method for segmenting data cues and processing fragments using multiple instances of a model graph.
[0023] Figure 24 An example embodiment is shown with a flowchart of multiple actions associated with a method for dynamic batch processing of prompt words. Detailed Implementation
[0024] The disclosed embodiments include system methods for generating and utilizing metamodel topologies. A metamodel topology refers to multiple interdependent models or functions that have been abstracted into multi-layered or multi-level networks, where only some inputs and outputs of the metamodel topology are visible through a user-facing system. Different models or functions are configured to perform different discrete tasks to process different input data. For example, some models or functions are configured to facilitate content moderation, annotation or tag generation, content modification, content generation, etc. As described in more detail below, different models and functions are configured as neural networks, pattern matching algorithms, or used to perform other data operations. The framework of the metamodel topology also allows for the efficient integration of new models into the metamodel topology, making the inputs and outputs of the models included in the metamodel topology compatible with each other.
[0025] In some cases, the metamodel topology is configured to perform content moderation on data prompts generated by a large language model or multiple different model networks. Some embodiments include using multiple model graph instances to process fragments of data prompts, where each model graph instance is tailored to a specific fragment of the data prompt. Some embodiments include automatic expansion of model graph instances. Furthermore, the disclosed embodiments also include dynamic batching of prompt queries. Dynamic batching is equally applicable to any data processing request, including both user-generated prompts and complete prompt content generated by the model.
[0026] The disclosed embodiments offer several technical advantages over conventional systems. For example, the disclosed embodiments facilitate the distributed deployment of service components and enforce a common pattern for orchestrating and coordinating applications using a dynamic model framework that can scalably adapt to an ever-growing number and variety of models. By implementing the system in this way, the disclosed embodiments significantly improve computational efficiency in handling requests received across multiple enterprises and users. Furthermore, computational efficiency is improved across different tasks and applications where the metamodel topology is trained and configured. Such frameworks also define interfaces, protocols, and libraries, enabling teams across organizations to deploy their models into production using procedural and uniform specifications.
[0027] The disclosed embodiments also improve the user experience by abstracting away the large number of models, tags, and taxonomies created across different modalities, such as speech, text, and images. These embodiments allow these models to be combined into a consistent set of content moderation configuration objects and policies in a simple, "no-code" manner. Therefore, the underlying network is abstracted away from the user when any changes are made to the models, where the user is presented only with a simplified user interface corresponding to a single, holistic metamodel capable of performing content moderation based on different policies associated with the different prompts being processed. With such embodiments, users can use a high-quality system without needing to understand the details of the individual models included in the system.
[0028] In some embodiments, the system includes a single model, such as a large language model or a generative pre-trained (GPT) model. Alternatively, the system includes multiple interconnected smaller models. Regardless of the approach, all models included in the system conform to a global labeling pattern, which includes a predefined set of labels and classifications. Advantageously, any changes made to the backend of the system's models will not affect the simplicity of user interaction with the front-end interface.
[0029] The disclosed batch processing technology for processing requests at different nodes can also help improve the efficiency of processing at different nodes, while enabling the meta-model to serve multiple different prompts from different users in parallel.
[0030] Overall, the disclosed embodiments facilitate close collaboration among numerous different model owners and achieve cross-model standardization in terms of labeling and taxonomy. The disclosed embodiments also promote the scalability of the meta-model framework and provide flexibility in the combination of interconnected models, while offering technical solutions for reducing model execution latency. Implementation of computing systems
[0031] First, shift your attention Figure 1The illustration shows a computing system 110 integrated within a computing environment 100, which also includes (via network 140) multiple client systems 120 and multiple remote systems 130 communicating with the computing system 110. These systems include and / or can be used to implement the disclosed embodiments.
[0032] As an example, computing system 110 includes one or more processors (such as one or more hardware processors 111), and a storage device (i.e., (a plurality of) hardware storage devices 113) storing computer-readable instructions 114 that can be executed by (a plurality of) processors 111 to cause computing system 110 to implement one or more aspects of the disclosed embodiments. Computing system 110 also includes (a plurality of) user interfaces 112 for receiving user input (such as prompts processed by computing system 110) and for presenting output, for example, based on the processed prompts. Computing system 110 may also include corresponding input / output (I / O) hardware devices for receiving user input and presenting corresponding output.
[0033] As shown in the figure, the client system 120 also includes one or more processors 121, one or more user interfaces 122, one or more input / output (I / O) devices, and one or more hardware storage devices 123 storing computer-readable instructions 124 that can be executed by the processors 121 to perform the functions described herein.
[0034] Although not shown in the figures, the remote systems 130 may also include a processor, a user interface, (I / O) devices, and a storage device containing computer-readable instructions executable by the processor to perform the functions described herein.
[0035] like Figure 1 As shown, hardware storage devices 113 and 123 are presented as independent local storage units. However, it should be understood that each of these storage devices can be configured to include distributed storage devices with local and / or remote distribution. Additionally, in this scenario, computing system 110 and / or client systems 120 can both be considered distributed systems, whose combinations of any components can be maintained and / or executed locally and / or remotely. In some cases, multiple distributed systems perform similar and / or shared tasks to achieve the disclosed functionality, such as in a distributed cloud environment.
[0036] In some cases, computing system 110, client systems (multiple) 120, and / or remote systems (multiple) 130 are configured to generate, train, and / or use various machine learning models during the generation and deployment of meta-model topology 115 and corresponding model graphs (e.g., model graph 116).
[0037] The metamodel topology 115 includes multiple models, or more generally, multiple functions used to process the input data during inference. The metamodel topology also has a global labeling pattern 119, which includes multiple tags of interest (ROIs) for labeling different parts of the input data. In some cases, each model or function included in the metamodel topology 115 follows the principle of generating a label value corresponding to at least one ROI selected from the global labeling pattern. Models or functions associated with similar ROIs are also grouped or associated together within the metamodel topology.
[0038] The models mentioned in this document are sometimes configured as machine learning models or models derived from machine learning, such as deep learning models, algorithms, and / or neural networks. Such models can include LLMs. Some LLMs are specific types of large generative models (LGMs) capable of receiving multimodal inputs and generating multimodal outputs, including processing real-time streaming content. In some cases, these models can also be configured as engines or processing systems (e.g., computing systems integrated within computing system 110), where each engine includes one or more processors (e.g., hardware processors 111) and computer-readable instructions 114 corresponding to computing system 110.
[0039] In some cases, the term "model" also refers to an abstraction, which represents the process of using input to classify or regress one or more labels of a global label pattern 119.
[0040] Some models are programmatically defined as different model subtypes, which can be plugged into the framework and invoked by services. Models can also invoke other models in a composite manner and can define functions on inputs and internal model outputs to serve as both outputs and internal model inputs simultaneously (see [link to documentation]). Figure 14 In some cases, these models or functions are only responsible for scheduling tasks between remote endpoints hosting the actual models, or they can directly execute logical operations.
[0041] In some configurations, the model is a set of numerical weights embedded in a data structure, while the engine is a separate piece of code that is configured to load the model when executed and compute the model's output in the context of the input.
[0042] The functions included in the metamodel topology are configured to perform various operations on the input data to generate function outputs that include label values corresponding to labels of interest from global label patterns. For example, these functions are configured to perform logical or applied operations on the input data. Examples of executable logic include pattern matching, such as n-gram matching, block lists, or regular expressions. Some pattern matching is also performed using a trie data structure, where each trie node is a set of similar patterns, such as bigrams, regex, or wildcards. Some functions are local neural networks or lightweight alternative classifiers, such as random forests and decision trees, Bayesian networks, or regression models. Such models are configured to perform tasks such as determining whether a portion of a data stream is a clear positive case, thus avoiding calls to downstream models that provide additional analysis or improve accuracy when predicting positive cases.
[0043] In some cases, each function generates labels of interest for specific parts of the input data, or generates intermediate outputs that are used as input to downstream functions, or combined with other intermediate outputs to generate the final labels of interest. In other cases, the model outputs only labels from the global label pattern 119. This is a predefined constraint of the framework, introduced to ensure the simplification and maintainability of the registered model subtypes. Specifically, this constraint facilitates future refactoring of the model ensemble approach.
[0044] In some cases, any model in the system may output any label value registered in global label pattern 119. However, each model specification will specify the output it enables in the candidate set of global label pattern 119.
[0045] Additionally or alternatively, the metamodel topology 115 includes multiple models configured to receive an initial seed prompt or intermediate prompt and generate new outputs based on that initial seed prompt or intermediate prompt; the metamodel topology 115 also includes multiple models configured to process these new outputs. For example, in some cases, the metamodel topology is configured to include models that can invoke one or more different nodes or layers of a large language model (such as the GPT model).
[0046] By implementing the system in this way (essentially, where the orchestrator 318 includes a model graph 320 integrated with the LLM 314), the system is able to simultaneously generate and process new outputs in alternating sequences. For example, in some cases, the first model / node of the integrated model graph can receive a first new initial seed cue and generate a new output based on that new initial seed cue. Subsequently, the second model analyzes whether there is anything unexpected in the new output and generates values for specific tags in the global tag pattern 119.
[0047] In certain situations, if the value for a specific label exceeds a predetermined threshold, the system will modify the initial seed word and apply the modified initial seed word to the first model to generate a second new output. The second model then analyzes / processes this second new output and generates a new value for the specific label in the global label pattern. The purpose of modifying the initial seed word is to guide the first model to generate a second new output with a lower value for that specific label in the global label pattern.
[0048] Model graph 116 includes multiple nodes that represent unique functions configured to perform operations on input data and are configured to generate label values corresponding to one or more labels of interest included in the global label pattern 119 associated with the model graph.
[0049] In some cases, global tagging pattern 119 is configured as a taxonomy for the structured organization and encapsulation of the different models included in the meta-model topology 115. Some global tagging patterns are configured as content moderation patterns, responsible artificial intelligence (AI) patterns, or "safe AI" patterns, which include tags of interest used to facilitate the monitoring and labeling of unwanted model-generated outputs. Global tagging patterns can be updated and modified as new models and new tags of interest are developed and defined.
[0050] By providing a meta-model topology defined by a global labeling pattern in this way, the model framework improves upon conventional systems that lack access to large datasets of multi-labeled examples and a general pattern for providing attribute supersets for managing internal datasets and model outputs. Additionally, the framework provides a method for easily integrating independently created and trained models into the meta-model topology without further training or fine-tuning. Instead, these models / functions and their outputs conform to the global labeling pattern. This improves computational efficiency because the system does not require further training of the models before or during integration into the meta-model topology. The computational efficiency also extends to the fact that the meta-model topology itself does not require additional training when a new model is integrated into its framework.
[0051] This allows distributed service components to be integrated into any cross-company model contribution, such as neural network models, n-gram guard lists, embedding clusters, composite models, and policies incorporated into the metamodel topology. The exposed system can host all models within the entire metamodel framework. Alternatively, the system can use the framework to interface with one or more remote model processing systems and manage workload offloading.
[0052] In addition to storing computer-readable instructions 114, meta-model topology 115, global label patterns 119, and model graphs 116, the (multiple) hardware storage devices 113 also store various electronic data, including data cues 117. These data cues 117 can be generated by human users or artificial users. Data cues are of various types, including initial input cues used as input to generative pre-trained models to generate complete cues (e.g., the output of processed input cues).
[0053] The meta-model topology 115 and its corresponding model graph (e.g., model graph 116) can also be used to process complete cue words generated by generative pre-trained models or other types of models. In some cases, these complete cue words, or the complete content of the initial input cue words, are generated as streaming content. Additionally, any data cue words include unimodal data content (e.g., text-only) or multimodal data content (e.g., text and image content). In some cases, generative pre-trained models generate complete cue words token by token at extremely high speed. These tokens can then be buffered and / or segmented before being processed through the model graph. Therefore, streaming content or stream can refer to user-generated cue words, model-generated cue words, or a combination of both.
[0054] In some cases, data cues 117 may include electronic content obtained from one or more external sources, such as emails, text messages, research papers, websites, news articles, academic papers, movies and television programs, audio files such as music or podcasts, or other media content. Therefore, data cues 117 may also include electronic content of various modalities, including text, audio, images, and / or video. It should be noted that, therefore, in some cases, different models and functions processing data of similar modalities may also be grouped or associated within the metamodel topology 115.
[0055] Additionally, in some cases, data cues 117 are segmented into multiple data cue segments (e.g., segment 118). Data cues 117 can be segmented in a variety of different ways. For example, they can be segmented by sentence, by paragraph, or by a predetermined length of text segment (e.g., by word count). In some cases, data cues 117 are segmented based on topic or theme.
[0056] In some cases, data cues 117 (including the initial seed cues) are first segmented and then routed to a Large Language Model (LLM) for processing to generate output, which is then processed by the metamodel topology. Alternatively, data cues may include output from the LLM, which is first segmented and then routed to the metamodel topology for policy compliance and content moderation processing.
[0057] Each segment can be processed separately (e.g., in parallel) by multiple instances of model graph 116. Each instance of model graph 116 can be pruned based on metadata or policy information associated with each segment of the data prompt. In some cases, the data processing request (e.g., the data prompt) includes a header or other specification to specify the policy to be used. For example, the entire policy can be given inline and include an identifier in the data processing request that enables the system to use the identifier to retrieve the policy from remote storage, or the policy can be built into the binary source code executed during deployment. The policy contains information on (i) which segmenter components to use and / or (ii) which tags in the global label pattern are of interest. It should be understood that the system can select from multiple segmenters, each configured to perform different segmentation policies for different tags included in the global label pattern. Each segmenter produces segments, and each resulting segment is processed through the entire model graph instantiated for that segment.
[0058] Now let's turn our attention to... Figure 2 This diagram illustrates an example computer architecture used to configure a metamodel topology as a content moderation system. Metamodel topology 115 and its corresponding model graph (e.g., model graph 116) are configured to perform content moderation as part of a responsible artificial intelligence (RAI) system. Specifically, the RAI system monitors outputs generated by one or more different models (e.g., large language models and generative pre-trained models such as ChatGPT).
[0059] Figure 2The diagram illustrates a RAI regional endpoint 202 communicating with either a content moderation proxy 201 or a reverse proxy. The RAI regional endpoint 202 facilitates connections between different components of the RAI framework and third-party systems. The reverse proxy receives data processing requests, including initial prompts provided to a remote generative model. The output of the remote generative model is then routed from the reverse proxy to the RAI endpoint 206. In some cases, communication between the reverse proxy and the RAI regional endpoint is a bidirectional remote procedure call (BIDI-GRPC) to allow for the transfer of electronic content input and output in streaming applications, where initial input prompts are streamed, and fully moderated / annotated prompts are streamed output.
[0060] In one embodiment, in a streaming application where data is segmented, when the system receives a data stream (e.g., text and / or images), a strategy-dependent segmenter performs segmentation on the data stream and executes a graph instance for each segment. When the system receives the results of each graph run, it streams the results back to one or more systems as streaming output in response to RPC requests. In some cases, the results are streamed back to user-facing systems. Each result message indicates the start / end position of the application stream. The result message is configured as a summary report similar to its corresponding segment and uses text offsets to indicate the position. Images are assumed to be zero-length text located at the character position in the data stream.
[0061] RAI regional endpoint 202 is shown as RAI endpoint 206 communicating with RAI publish-subscribe topic 208, RAI result store 212, and RAI async worker 210. RAI publish-subscribe topic 208 coordinates and routes processing requests between RAI endpoint 206 and RAI async worker 210.
[0062] RAI Results Store 212 stores the following: RAI Reports 214, RAI Metrics 216, and RAI Alerts 218. RAI Reports 214 include performance reports, detailed information on various tasks, and updates to the RAI system. RAI Metrics 216 include metrics such as processing efficiency, the number of currently running model graph instances, and pruning time. RAI Alerts 218 include alerts such as system failure alerts, system completion alerts, and tagged alerts when content is identified and / or marked as violating specific audit or compliance policies.
[0063] RAI asynchronous worker 210 coordinates between RAI endpoint 206 and RAI publish-subscribe topic 208 to orchestrate different processing requests and model instantiations on the asynchronous processing path of the exposed framework.
[0064] RAI region endpoint 202 also includes RAI meta-model 220 (e.g., Figure 1 The RAI metamodel topology 220 communicates with the event response server 228 and various RAI endpoints. These RAI endpoints focus on different tags of interest, which users may wish to monitor or use as a basis for tagging and filtering inappropriate content. Examples of RAI endpoints include personally identifiable information (e.g., endpoint 222), hate speech endpoints (e.g., endpoint 224), and pornography endpoints (e.g., endpoint 226). Each of these endpoints corresponds to a specific tag of interest (e.g., PPI, hate speech, or pornography) and can be used to tag, label, and / or modify input data.
[0065] Different tags of interest corresponding to different RAI endpoints are stored in a global tag pattern, such as Figure 1 The global label pattern 119. The RAI metamodel 220 includes one or more RAI model ensembles (e.g., RAI model ensemble 230), which includes multiple models and functions for performing various operations on the input data at different times during the inference process of processing the input data.
[0066] Overall, systems such as RAI System 200 advantageously provide a framework for monitoring content generated and / or processed by different machine learning models to prevent harmful or potentially harmful content from reaching client-facing systems and / or otherwise enforce policy compliance. This improves user experience by protecting users from receiving unwanted generated content. Additionally, this framework also allows model administrators to monitor the parameters of different machine learning models. If a model irresponsibly generates content, the model administrator can receive alerts in real time and can pause content generation or reconfigure / retrain the model to avoid generating harmful content. Meta-model topology and deployment
[0067] Now let's turn our attention to... Figures 3-13 This illustrates example data flows between components of an example computer architecture used to facilitate the integration of the metamodel with client-oriented systems. It should be understood that while the metamodel topology is configured for RAI systems in some applications, the disclosed embodiments are applicable to a wide variety of different applications. System 300 advantageously supports the flexible integration of any number of models, and the models can be changed at any time.
[0068] System 300 also supports the evolution of data patterns for tags, which can be added or replaced over time. Tags include binary classification, categorical regression, or other combinations. The exposed framework is also configured to support different subjects (e.g., users, tenants, resource objects) so that their contributions meet the acceptability criteria for each tag.
[0069] System 300 supports filtering of synchronous requests / responses and asynchronous monitoring. For example, contributors can deploy computationally intensive models only to asynchronous paths, as the latency risks incurred on synchronous paths may not yield valuable benefits. In another example, contributors can deploy undertrained models only to asynchronous paths, providing opportunities for manual annotation and subsequent training of low-confidence prediction samples. Similarly, some models may be used solely for monitoring rather than providing client-facing output.
[0070] Additionally, the same model can be deployed at two different runpoints on the filtering and monitoring paths. The filtering path conservatively prioritizes reducing the false positive rate, sacrificing recall for higher precision; while the monitoring path may be adjusted to be more tolerant of low-probability predictions, in order to optimize for higher recall. In summary, the primary goal of the synchronous or filtering path is to block content or label warnings, while the primary goal of the asynchronous or monitoring path is to label the model, function, and content for subsequent processing.
[0071] Now let's turn our attention to... Figure 3 The diagram illustrates the computer architecture, including client 302, a cognitive service gateway layer (e.g., API gateway 304), a front-end ingress 306, and a data plane 308. Front-end ingress 306 allows data traffic to enter through API gateway 304. Data plane 308 is configured to proxy API requests, apply policies, and collect any data required for the requests. Client 302 provides the system with an initial prompt (such as "We should"), which is then routed through the routing logic of data plane 308 to engine stack deployment 310 (see [link to deployment details]). Figure 4 ).
[0072] For example, such as Figure 4 As shown, engine stack deployment 310 includes a reverse proxy 312 (e.g., representing content moderation proxy 201) and an engine API 313. In some cases, the engine API is configured to communicate / connect to an LLM 314 (such as ChatGPT). Although LLM 314 is shown as part of engine stack deployment 310, LLM 314 can also reside outside of engine stack deployment 310 and may even include models and nodes referenced in model diagram 320.
[0073] Once reverse proxy 312 receives the initial prompt word, it is routed to LLM 314 via the engine API. In some cases, reverse proxy 312 is used as a reverse proxy for data moderation configuration. Reverse proxy 312 is also referred to as a content moderation proxy (CMP). Subsequently, LLM 314 is applied to process the initial prompt word, thereby generating output corresponding to the initial prompt word.
[0074] Now let's turn our attention to... Figure 5 As shown in the figure, after the initial prompt is processed by LLM 314, LLM generates output based on the initial prompt. For example, as... Figure 5 As shown, LLM 314 generates output (i.e., an intermediate or complete result of the initial prompt), including "We should break your lamp," which is a prediction of what should follow "We should" in the initial prompt. The output, including the intermediate or complete prompt, is then routed back to the reverse proxy 312 via engine API 313.
[0075] Now, as Figure 6 As shown, the prompt word output is routed to orchestrator 318 via any middleware (e.g., middleware 316) for processing based on content moderation and policy compliance. Middleware 316 serves as a front-end service layer for the processes performed by orchestrator 318, such as authentication, load balancing, DNS service discovery, and DDoS protection. As shown, an instance of orchestrator 318 is generated for the output received from reverse proxy 312. In some cases, orchestrator 318 is configured as a RAI orchestrator. In this regard, different orchestrator instances and corresponding model graph instances can be created for each prompt word or prompt word fragment processed, as described in more detail below.
[0076] As shown in the figure, the orchestrator 318 includes a model diagram 320, which is configured as a content moderation diagram or a policy compliance diagram.
[0077] Model diagram 320 shows Figure 1 The model is shown in Figure 116. Additionally, as... Figures 3-7 The model shown in Figure 320 is in Figure 15 This is explained in more detail. For example... Figures 3-13 The model shown in Figure 320 is both a function library and its practical application. During certain data processing requests, multiple orchestrator instances are generated, and they are executed in parallel / simultaneously with each other.
[0078] Now let's turn our attention to... Figure 7Once an instance of the orchestrator is generated, the system identifies a strategy (i.e., specification information for pruning model graph 320), such as by using the strategy runner 322. The strategy runner 322 facilitates data transfer between the strategy registration service 334 and the orchestrator 318. For example, a strategy is loaded from the strategy registration service 334. This strategy (i.e., the loaded strategy) specifies how the system prunes the model graph 320 included in the orchestrator 318. The model graph 320 is compiled at runtime.
[0079] In some cases, policies are pre-registered by registering the code in a shared repository (e.g., policy registration service 334). The policy runner can then apply one or more of the registered policies along the synchronization path via model diagram 320. The policy runner selects the type of policy to be evaluated in model diagram 320 by identifying the context of the user, tenant, or application object to which the request belongs. A subject-to-policy mapping function can be used to map the authenticated identity of a subject (e.g., a user) to the type of policy to be applied, or to facilitate user-defined policy instances for further configuration. Alternatively, policies can be identified based on policy information included in the output sent from reverse proxy 312 to orchestrator 318.
[0080] Now let's turn our attention to... Figure 8 This illustrates how model graph 320 is pruned according to a loaded strategy. During pruning, one or more nodes of model graph 320 are omitted from the model graph instance based on the strategy information identified in the loaded strategy. In some cases, the strategy is associated with a specific subset of tags identified from the global tag pattern associated with model graph 320.
[0081] like Figure 8 As shown, with Figure 7 Compared to the model graph 320, some nodes and edges of the model graph 320 have been pruned. Figures 8-13 The model shown in Figure 320 is in Figure 17 This was further explained in the text. It should be understood that... Figure 8 The pruning process in the diagram (e.g., “(2) Pruned diagram”) is for reference. Figures 14-17 The pruning process is described further.
[0082] Now let's turn our attention to... Figure 9 The output routed from reverse proxy 312 (“We should have smashed your light”) (which is generated based on the input prompt “we should” from client 302; see also...) Figures 3-5 Each node of model diagram 320 is processed. For example, as shown in Figure 320. Figure 9As shown, a node is configured to invoke an instance of a specific model (e.g., model A instance 324) via middleware 316. In some cases, one or more nodes are configured to perform logic or apply functions to the input data. Examples of executable logic include pattern matching, such as n-gram matching, block lists, or regular expressions. Some pattern matching is also performed via a trie data structure, where each trie node is a set of similar patterns, such as tuples, regular expressions, or wildcards. Some models are local neural networks or lightweight alternative classifiers, such as random forests and decision trees, Bayesian networks, or regression models. Such models are configured to perform tasks such as determining whether a portion of a data stream is a clear positive case, thereby avoiding invoking downstream models that provide additional analysis or improve accuracy when predicting positive cases. Additionally or alternatively, one or more nodes are configured to invoke remotely managed models and / or be configured to invoke locally managed models. Such models may include neural networks or other machine learning models configured for language identification. In this case, the result can be a label, which is fed as an intermediate label to the downstream model node to determine which remote endpoint the request will be dispatched to (e.g., the remote endpoint for English hate speech may be different from the remote endpoints for other languages).
[0083] In some cases, the system utilizes real-time event-driven topology sorting to process data cues. Event-driven topology sorting schedules the execution of each generated graph instance. This allows the system to maximize the parallelism of graph execution and ensures that graph execution completes as quickly as possible. This allows for reduced latency, especially in real-time streaming applications. This improves the user experience by returning reviewed content to the user with increased time efficiency. Event-driven topology sorting is an algorithm that allows the system to determine which nodes are ready to be executed. Those nodes ready to be executed are executed as quickly as possible when processing a specific data processing request. When a node is successfully executed and returns its output, the system marks that node as complete, and the topology sorting algorithm continues execution for that specific data processing request.
[0084] The output of the invoked model instance is returned to model graph 320 (as indicated by the bidirectional arrow between model A instance 324 and the associated node in model graph 320). The output from model A instance 324 continues to be processed in different nodes of the content moderation graph (e.g., a node invokes model B instance 326).
[0085] Now let's turn our attention to... Figure 10 After the LLM output is processed through the entire model graph 320, the system outputs one or more values of one or more different labels to the policy runner 322 for classifying, filtering, and tagging the LLM output. For example, ... Figure 10 As shown, after model graph 320 is routed through policy runner 322, the LLM output "We should smash your light" returns a value of 1.00 under the Violence label. In some cases, this value is configured as a label for the LLM output (e.g., the full or intermediate prompt word generated by LLM 314 in response to the previous input prompt word processed by LLM 314).
[0086] Now let's turn our attention to... Figure 11 As shown in the figure, a label (e.g., brute force 1.00) is appended to the LLM output ("We should smash your light"). After the orchestrator 318 routes the modified or labeled tooltip to the reverse proxy 312 through middleware 316, the modified or labeled tooltip is returned to the client 302 by the reverse proxy 312.
[0087] like Figure 12 As shown, the LLM output and tag values (i.e., RAI annotations) are also routed from orchestrator 318 to responsible AI monitoring agent 330 via work queue 328. Work queue 328 is used to cache and store samples, such as review inference results and / or, optionally, association identifiers re-associated with prompt words or full text, for additional monitoring and potential follow-up processing. Inference results can be logged as log events and / or returned to the system that originally generated the request or the user-facing system on the real-time request path. Log results can be further processed downstream in the pipeline using online analytical processing (OLAP) methods, such as pattern identification, or used to detect malicious users in content moderation scenarios.
[0088] refer to Figure 13 The label value (e.g., Violence: 1.00) is returned to the alert engine 332, which is configured to generate an alert notification to indicate that the LLM output includes labeled content that may violate a preset policy. This alert can be passed to the client or system administrator to address undertrained models, improper model use, or other label value content.
[0089] It should be understood that the computer architecture (more specifically, the model complex) may include any configuration of the metamodel topology and its corresponding model graph to perform various functions, including but not limited to policy compliance and content moderation.
[0090] Implemented as described above, the results (e.g., intermediate or complete prompts) processed by machine learning models (such as LLM or ChatGPT types) are analyzed, labeled, flagged, and / or modified before being returned to the client-facing system (i.e., the user who issued the request / prompt). A reverse proxy is configured to facilitate the return of results between the LLM and the moderation service. When a prompt is flagged as containing content that violates the policy, the system can block any output from the LLM or other models from being routed to the client. Audit Model Diagram
[0091] Now let's turn our attention to... Figures 14-16 These diagrams illustrate an example process of generating a content moderation graph from a metamodel topology and pruning it.
[0092] Figure 14 A meta-model topology 500 is shown, which includes multiple functions and models. Once the system identifies multiple functions and models, it organizes them according to different levels and groupings. For example, the first level includes an alpha model 402, which has input data points 416 and one or more output data points (e.g., output a, output b, output c, output d, output e, and output f). The alpha model 402 is shown as having multiple models and functions organized at different sub-levels. For example, the alpha model 402 (located at the second level of the meta-model topology 500) includes a beta model 404, a gamma model 410, a zeta model 412, and an eta model 414, as well as several additional functions in the alpha level: function f2(), function f3(), function f4(), and intermediate variables.
[0093] Beta model 404 includes multiple models and functions located at the third level, including beta1 model 406, beta2 model 408, and function f9(). Beta1 model 406 includes function f7(), while beta2 model 408 includes function f8(). Beta model 404 is shown as having input data points 405 and multiple outputs (e.g., output a, intermediate output d, and output e).
[0094] The gamma model 410 includes a function f5() at the third level of the meta-model topology, which generates output c based on the input received at input data point 411. The zeta model 412 is shown receiving input data at input data point 413, which is routed to functions f6() and f10(). The output of function f6() is also routed through zeta intermediate variables to generate output d. Additionally, the output of function f6() is also routed through zeta intermediate variables as input to function f10(). Function f10() then generates output f. The eta model 414 is configured to receive input data at input data point 413, which is then routed through function f1() to generate intermediate output d.
[0095] In the alpha or top-level of the meta-model topology 500, output 'a' comes directly from output 'a' generated by beta model 404. Output 'b' is based on the output generated by function f3(). Output 'd' from beta model 404 is routed to function f3() as input via function f4() and intermediate variables. Output 'c' from gamma model 410 is also routed to function f3() as input. Output 'd' from zeta model 412 is also routed to function f3() as input.
[0096] Output c is directly routed from the output c generated by gamma model 410. Output d at the alpha level is based on output d from eta model 414 and output d from zeta model 412, both of which are routed as inputs through function f2(). Output e is directly routed from output e of beta model 404. Finally, output f at the alpha level is directly routed from output f generated by zeta model 412.
[0097] In some cases, different models and functions are organized based on what type of modality they are configured to receive as input data. For example, beta model 404, gamma model 410, and zeta model 412 receive text-based input, while eta model 414 receives custom input. In other examples (not included in...) Figure 14 As shown in the diagram, beta model 404 can receive image-based input, gamma model 410 can receive audio-based input, and zeta model 412 can receive text-based input. Alternatively, beta model 404 can receive multimodal input, wherein beta1 model 406 receives the video-based portion of the multimodal input, and beta2 model 408 receives the audio-based portion of the multimodal input.
[0098] In some cases, each model at the second level (e.g., beta model 404, gamma model 410, zeta model 412, and eta model 414) corresponds to a specific tag of interest. For example, beta model 404 could correspond to hate speech, gamma model 410 to violence, zeta model 412 to pornography, and eta model 414 to personally identifiable information (PPI). In this configuration, beta model 404 generates at least two distinct tag values for hate speech: a binary hate speech tag from beta1 model 406 and a hate speech severity level tag value from beta2 model 408.
[0099] Gamma model 410 generates at least one violence label value (e.g., a binary label value). Similar to beta model 404, zeta model 412 generates at least two distinct pornographic content label values, including a binary label (e.g., output d) and a severity level output (e.g., output f). Similar to gamma model 410, eta model 414 generates at least one label value for PPI, including a binary label value (e.g., intermediate output d). It should be understood that functions and models can also be organized according to a combination of interest label grouping and modality grouping.
[0100] Contributors can add their models to the system by writing subtypes of model classes and registering their code in a globally shared repository. Contributors define their models by specifying: a typed declaration of any internal models; the outputs expected to be enabled by the global tagging pattern (which can be invalidated in case of timeouts or other exceptions); and any additional inputs besides the input text. Input adaptation can be done within functions included in the model. Contributors also need to specify how the outputs are derived from the inputs, internal model outputs, and / or intermediate variables, ranging from simple direct passing to complex logic in instantiated scientific libraries. Specifications are also defined to define how intermediate variables are derived from the inputs, internal model outputs, and / or intermediate values, and how internal model inputs are derived from the inputs, internal model outputs, and / or intermediate variables.
[0101] Examples of such models include a neural network classifier for binary classification of hate speech, which generates a one-dimensional output. This classifier is task-specific. Therefore, global label mode 119 will include the binary label "Is_Hate_Speech", whose label value represents the probability or likelihood that the content is hate speech. The classifier's output is directly mapped to this label, while all other labels are disabled. In other words, in global label mode 119, only this label is enabled.
[0102] Another example of the model includes a profanity protection list, which uses a hard-coded function to output regression scores based on the weights and frequencies of terms. However, assuming the global label pattern only includes the labels "Profanity_Light" and "Profanity_Extreme", the contributed model subtype will define the corresponding code to adapt the regression scores to the labels.
[0103] In some cases, a model within the framework can be viewed as a stub, placeholder, handle, or marshal of one or more underlying predictive models. Most complex models (including neural network classifiers and regressors) involve service-to-service calls. These calls are executed by functions defined in the model hierarchy, and the framework processes them asynchronously. In some cases, these complex models are hosted by an AML hosting platform that can be deployed collaboratively with the model graph. It's important to note that remote scheduling can be performed within model functions. In other cases, the model graph is instantiated within the orchestrator, allowing complex models to be hosted by a dedicated platform whose continuous deployment workflow can be shared with the orchestrator.
[0104] A model in the metamodel topology 500 can call multiple underlying model endpoints. The contributed function is configured to adapt the output of the underlying model to labels in the global labeling pattern. Unlike complex models, lightweight tasks and remote scheduling can be executed directly within the contributed model. These tasks are executed on an orchestration thread pool. Low-level lightweight models (such as shallow neural networks, keyword guard lists, and non-neural methods) can be executed directly within framework-level functions.
[0105] Each model specification also defines timeouts and provides a cancellation mechanism for remote calls. When such exceptions occur, enabled outputs will be invalidated. Null propagation occurs when one of the node's inputs is null. In this way, even if a timeout or other null exception occurs, non-null paths in the graph can still propagate, thus still generating some non-null labels.
[0106] Now let's turn our attention to... Figure 15 This figure illustrates an example embodiment of model diagram 510. After all models and functions are organized according to their different levels, groupings, and dependencies, the system converts the metamodel topology 500 into model diagram 510.
[0107] Model graph 510 includes multiple nodes and multiple edges, where each of the multiple functions in the meta-model topology 500 is represented as a node in the model graph, and each of the multiple edges connecting at least two nodes in the model graph represents the data dependency relationship between different functions in the meta-model topology 500.
[0108] When a data processing request is received that includes a data prompt word, the system generates an instance of a model diagram, such as... Figure 15 As shown. Subsequently, the model graph instance can be pruned to omit nodes irrelevant to the data cue. For example, the system identifies one or more tags of interest to generate tag values based on the data cue. Based on the identified tags of interest, the system can identify a subset of nodes associated with one or more tags of interest from multiple nodes in the model graph. Before applying the model graph to the data cue, the system modifies the model graph instance by omitting one or more nodes in the model graph that are not included in the identified subset of nodes associated with one or more tags of interest.
[0109] Depth-first search can be used to identify irrelevant nodes. For example, ... Figure 16 As shown, alpha outputs a and b are identified as the desired outputs (as indicated by the left arrow). The system determines that alpha outputs d, e, and f are irrelevant (as indicated by the "X" above the corresponding nodes).
[0110] Now let's turn our attention to... Figure 17 This figure illustrates an example embodiment of the pruned model diagram 510. (As shown...) Figure 17 As shown, after irrelevant output nodes are identified and omitted, the model graph is further pruned to omit any nodes that precede these irrelevant nodes in terms of dependencies. Therefore, once the model graph instance is further pruned in this manner, the pruned model graph instance (e.g., model graph 530) can be applied to input data prompts. segmentation
[0111] Now let's turn our attention to... Figure 18 It illustrates an example implementation of generating multiple instances of a content moderation graph (or diagram) for each segment of an input prompt word, where each instance includes a different pruned version of the content moderation graph.
[0112] During processing, the system is configured to identify data prompt words (e.g., data prompt word 602) and segment the prompt word into multiple segments (e.g., input segment 604 and input segment 606). Subsequently, for each segment of the prompt word, the system generates a separate graph instance (e.g., instance 608 and instance 610), where the graph instance corresponds to the specific segment.
[0113] The system then prunes the graph instance based on the different strategy information for each segment. For example, instance 608 is shown as omitting at least four different output nodes, while instance 610 is shown as omitting at least three different output nodes. In this way, the system can generate customized graph instances for each segment, where each instance omits different nodes from the underlying basic review graph based on the different strategies applied to different prompt word segments.
[0114] For example, in the first segment, policy information might instruct a scanning task for segments containing violent content; while in subsequent segments, policy information might instruct a scanning task for segments containing hate speech. By implementing the method in this way, the system can process input prompts in a more accurate and granular manner because each graph instance is specifically tailored for different segments. Additionally, this improves computational efficiency because multiple graphs can run concurrently, while also saving processing memory by loading only those nodes most relevant to a specific segment of a particular graph instance.
[0115] It should be understood that, in some cases, each segment is generated by one or more segmenters after the data stream is sent through the system. These segments are generated as early as possible, depending on when the data stream is processed. The data stream is continuous, but these segments are generated through buffering. The system generates a model graph for each segment and performs a complete model graph operation on each segment via an algorithm such as topological sorting.
[0116] By implementing the system in this way, it gains the technical advantages of facilitating low-latency streaming processing. For example, these segmenters are configured to handle both streaming and non-streaming data prompts. Therefore, some strategies are specifically configured for streaming scenarios and optimized for generating incremental fragments. The optimal fragment length should neither be too short (i.e., including sufficient tokens to provide context during inference) nor excessively buffered. In contrast, non-streaming strategies (or segmenters) are configured to select the optimal fragmentation method. One approach is to use a punctuation-sensitive scrolling window or a scrolling window sensitive to the regular expressions of the detection clauses. Batch processing
[0117] Now let's turn our attention to... Figures 19-20It illustrates an example embodiment of batch processing data processing requests at one or more nodes of a model graph, wherein the model graph may be the data audit graph described above.
[0118] Figure 19 Multiple model graph instances generated for multiple tenants are shown. For example, instances 702, 704, and 706 are shown as instantiated for tenant A, while instances 708, 710, and 712 are shown as instantiated for tenant B.
[0119] In some systems, each model graph instance is generated in response to different data processing requests from one or more different users or in response to receiving different data cues. (In some systems, a single data processing request may include one data cue, or alternatively, may include multiple data cues). It should also be understood that these different model graph instances may be generated based on fragments of data cues, with each instance being pruned according to policy information associated with each fragment of the data cues. Therefore, in some systems, one or more instances of the model graph may comprise the same subset of nodes, while in others, one or more instances of the model graph may comprise different subsets of nodes.
[0120] In order to improve Figure 19 The computational and processing efficiency of the meta-model topology in the multiple examples shown is achieved by batching different sets of data processing requests for a specific node. Efficiency can be achieved by reducing the number of times the registry state changes when a node receives different requests. Specifically, by batching requests, a node can utilize the same common set of registry entries when processing multiple requests, thereby reducing the inefficiency caused by resetting the registry for each different request if they were received as separate, unbatched requests.
[0121] like Figure 19 As shown, the system identifies multiple model graph instances from tenant A and tenant B for node f5 (e.g., representing...). Figure 17 Data processing requests from node f5 in instance 702. For example, for tenant A: data processing request 714 from instance 702, data processing request 716 from instance 704, and data processing request 718 from instance 706. For tenant B: data processing request 720 from instance 708, and data processing request 722 from instance 710. In this case, no data processing request is received from instance 712 (i.e., instance 712 may have been pruned to omit node f5).
[0122] Subsequently, multiple data processing requests (e.g., batch 724) are transmitted and temporarily stored in the batch processing cache corresponding to node f5. The system identifies one or more batch processing criteria associated with node f5. For example, such as... Figure 20 As shown, batch processing criteria 726 includes node maximum value 728, time elapsed since the last request 730, and other batch processing criteria 732.
[0123] The maximum node size of 728 indicates the maximum number of data processing requests a node can handle as a single input. This is a threshold for the maximum queue size of batch requests. Specifically, the maximum node size of 728 indicates the maximum number of data processing requests that can be included in the batch queue or cache allocated to a particular node before being scheduled for processing. Larger batches are generally more hardware efficient but can lead to increased latency. Therefore, setting a maximum batch size ensures that batches do not become so large that latency increases to unacceptable levels, impacting the user experience, especially in streaming applications.
[0124] The timeout criterion of 730 seconds since the last request is configured. This means that if a certain amount of time has elapsed since the last data processing request was received, the system will send a batch of data processing requests, including fewer than the maximum number of data processing requests, even if the maximum number of data processing requests has not been reached. This prevents system delays or unnecessary processing time extensions if / when the batch queue does not reach the threshold within the specified time.
[0125] Back Figure 19 The system identifies that node f5 has a maximum batch processing request count of five. Therefore, once five different data processing requests are received, batch 724 is transferred to node f5 for processing.
[0126] It should be understood that dynamic batching can be executed either on the orchestrator process side within the framework (e.g., as a framework feature) or at remote models (i.e., models called by one or more functions in the metamodel topology). Typically, there are multiple orchestrator processes, each running the graph framework, and for any given remote model, there are also multiple remote model processes. Therefore, dynamic batching can be applied on the orchestrator side, in which case batching requests are transmitted and routed to the remote models once batching criteria are met. Alternatively, each node in the model graph submits a request to a remote model process, and the dynamic batching algorithm is then applied in the remote process. The batch is then sent via forward inference once it is ready in the remote process. Depending on how dynamic batching is performed, the system imports different libraries (e.g., libraries configured for orchestrator-based batching or libraries configured for remote model dynamic batching). Example Method
[0127] The following discussion will involve several methods and method actions. Although these method actions may be discussed in a specific order, or shown in a flowchart as occurring in a specific order, a specific execution order is not required unless explicitly stated, or because the execution of an action depends on the completion of another action and must follow that order.
[0128] Now let's turn our attention to... Figure 21 It illustrates an example embodiment with a flowchart having multiple actions (e.g., actions 810, 820, 830, 840, 850, 860, and 870) that are associated with a method for deploying a metamodel topology and generating an audit graph from the metamodel topology, implemented by a computing system (e.g., computing system 110).
[0129] The first action illustrated includes an action (action 810) for accessing a metamodel topology containing multiple functions (e.g., metamodel topology 115). Each of the multiple functions is configured to perform a unique operation on the input data to generate a function output that includes label values corresponding to the labels of interest included in the global label pattern (e.g., global label pattern 119). This configuration of the metamodel topology advantageously provides a concise framework for processing data cues. Additionally, because multiple functions conform to the global label pattern, the output of the metamodel topology (or the corresponding model graph) is uniform. This uniform output makes subsequent downstream analysis or operations more accurate and efficient.
[0130] In some cases, the system may access new functions that are not included in the metamodel topology (Action 820). By identifying new functions integrated into the metamodel topology, the metamodel topology can be continuously improved, updated, and extended to provide improved data processing services and adapt to new applications or domains.
[0131] To integrate a new function into the metamodel topology, the system identifies a specific tag of interest associated with the new function from the global label pattern (action 830) and configures the new function to output a tag value corresponding to the tag of interest from the global label pattern (action 840). By allowing new functions to be integrated into the metamodel topology, the system ensures that the metamodel topology has the latest improved functions and provides additional operations not previously included in the metamodel topology.
[0132] After configuring the new function, the system integrates the new function into the metamodel topology (Action 850). Once the new function is integrated into the metamodel topology, the system converts the metamodel topology into a model graph (e.g., model graph 116) (Action 860).
[0133] The model graph includes multiple nodes (e.g., Figure 15 The model graph contains nodes f1, f2, f3, etc., and multiple edges. Each of the multiple functions in the metamodel topology is represented as a discrete node in the model graph. Similarly, each of the multiple edges connects two or more nodes within the model graph and represents a data dependency between different functions existing in the metamodel topology. For example, in some cases, the data dependency between two different functions is defined by an edge, such that the output of one model is used as the input of another. In other cases, the data dependency between different functions is defined based on the level of abstraction of a particular node within the model graph. For example, some nodes represent independent functions or models, while others represent a set of models or collections of models that collectively generate specific outputs for use downstream in the model graph. Some functions are organized according to which tag of interest from the global label pattern they are configured to output. Finally, the system deploys the model graph to remote applications (e.g., multiple user interfaces 122 of (multiple) client systems 120) configured to receive data cues (e.g., data cue 117) (action 870).
[0134] In some cases, the method also includes actions for processing input data using a model graph. For example, in this case, the system receives a data processing request (e.g., at a remote application (e.g., client 302)). Figure 3 The prompt word "we should" is used in response to a received data processing request. The system generates an instance of a model diagram (e.g., orchestrator 318) and applies the model diagram to the data prompt word corresponding to the data processing request.
[0135] Data cues can be generated in several different ways. For example, some data cues are generated by human users. Others are automatically generated by AI users or other computer users. Furthermore, model graphs can be applied to different types of data cues. Some data cues are initial cues, used as input to machine learning models (such as generative pre-trained language models), which then use these initial cues to generate complete cues. Therefore, model graphs can also be applied to complete cues. Additionally, it should be understood that data cues can also include any audiovisual content obtained from one or more external data sources.
[0136] In some cases, the method also includes actions for configuring different functions within the metamodel topology. For example, some systems determine which tags of interest from a global labeling pattern are associated with different functions included in the metamodel topology. For instance, a system might identify sets of functions corresponding to hate speech tags included in a global labeling pattern, while different sets of functions correspond to violence tags included in the same pattern. Within a specific set of functions, even if these functions correspond to the same tags of interest, there may be some data dependencies between them. For example, within the set of functions corresponding to hate speech tags, some functions might be configured for binary prediction, while others might be configured for regression prediction or hierarchical prediction (e.g., severity level).
[0137] These systems then generate one or more subsets of association functions, each subset comprising functions that follow the principle of generating output forms corresponding to similar or equivalent labels of interest (see [link to documentation]). Figure 14 (The dependency links between the model and the function are shown). For example, in some configurations, it is advantageous to execute the binary function first, and then increase the granularity of the analysis so that the severity level function is only executed if the output of the binary function indicates a positive probability of generating a hate speech label, where there is a dependency between the binary function and the severity level function that depends on the output of the binary function.
[0138] Some metamodel topologies include a single input data entry point (e.g., input data point 413), allowing users to interact with the metamodel topology through this single input point without interacting with the intermediate outputs generated by the individual functions within the metamodel topology. By implementing the system according to these embodiments, the user experience is improved, where users can submit data prompts in a streamlined manner and receive approved outputs because all underlying models and data dependencies between models are abstracted and hidden from the user interface.
[0139] Some users may want to update the metamodel topology using a new or improved function that outputs tag values of interest that are not yet included in the existing global tag pattern associated with the metamodel topology. In this case, the system accesses the new function, which was not previously included in the latest version of the metamodel topology. After identifying the new tags of interest associated with the new function, the system integrates the new function into the metamodel topology and updates the global tag pattern by adding the new tags of interest associated with the new function.
[0140] Some new functions integrated into the metamodel topology may output label values for labels of interest associated with another label of interest already existing in the global label pattern. When the system determines that a new function is similar to a specific function already included in the metamodel topology (e.g., beta1 model 406), the system links the new function (e.g., beta2 model 408) with existing functions in the metamodel topology. The system may determine similarity between a new function and existing functions based on the similarity between the first label of interest generated by the new function and the second label of interest associated with a specific existing function among multiple functions included in the metamodel topology.
[0141] The system can link two functions in different ways. In some cases, the input of one function may depend on the output of another. Additionally or alternatively, similar functions are linked to an abstract model associated with a specific label of interest or a similar set of labels of interest. For example, in some cases, the label of interest associated with an existing function is a binary label of interest (e.g., the function outputs a binary value for that label of interest), while the label of interest associated with a new function is a severity level label of interest (e.g., the function outputs a label value between a predetermined maximum label value and a predetermined minimum label value).
[0142] A global tagging pattern can include many different tags of interest. For example, tags of interest may include hate speech warning tags (e.g., endpoint 224), pornography warning tags (e.g., endpoint 226), violence warning tags, personally identifiable information tags (e.g., endpoint 222), or combinations thereof. It should be understood that tags of interest may also include topic tags that users may be interested in or component tags used to identify certain components of data cues.
[0143] Now let's turn our attention to... Figure 22 It illustrates an example embodiment of a flowchart with multiple actions (e.g., actions 910, 920, 930, 940, and 950), which are associated with a method implemented by a computing system (e.g., computing system 110) for performing content moderation on data prompt words (e.g., data prompt word 117) using a model diagram (e.g., model diagram 116).
[0144] The first action shown includes an action (action 910) to access a metamodel topology containing multiple functions (e.g., metamodel topology 115). Each of the multiple functions is configured to operate on the input data to generate a function output that includes label values corresponding to labels of interest included in the global labeling pattern associated with the metamodel topology. This configuration of the metamodel topology advantageously provides a streamlined framework for handling data cues. Additionally, because the multiple functions conform to the global labeling pattern, the output from the metamodel topology (or the corresponding model graph) is uniform. This uniform output makes further downstream analysis or manipulation more accurate and efficient.
[0145] After, simultaneously with, or before accessing the metamodel topology, the system receives a data cue (e.g., data cue 117) (action 920). In response to receiving the data cue, the system generates an instance of the model graph (orchestrator 318) (action 930) and applies that model graph instance to the data cue (action 940). By instantiating individual instances of the model graph, the system is able to improve processing capabilities, including the ability to tailor each instance to a specific data cue. This not only provides a customized experience for processing the data cue but also improves the use of computer memory, as the model graph includes only nodes relevant to processing that data cue and / or the desired application.
[0146] Based on applying the model graph instance to the data prompt, the system generates a tag value corresponding to at least one tag of interest (action 950) included in the global tagging pattern for that data prompt. This tag value is based on the output of the model graph. By generating values for one or more different tags of interest, users and model administrators can monitor the content generated by the AI system, as well as other content submitted to the system. This value can then be used to determine whether the content should be modified before returning the complete prompt to the user-facing system, or whether the machine learning model should be adjusted or improved to avoid such content.
[0147] In some cases, the method also includes actions to modify the data hint based on the model graph output. For example, some systems determine that the label value of a data hint is equal to or exceeds a predetermined threshold. When a label value is determined to be equal to or exceed the predetermined threshold, the system identifies one or more tokens associated with that label value in the data hint. Before returning the processed data hint to the client-facing system, the system modifies the data hint by removing or replacing one or more tokens. After modifying the data hint, the system displays the data hint on the user interface of the client-facing system.
[0148] Additionally or alternatively, certain instances of the model graph are configured to audit the content included in the data hints based on generating tag values corresponding to multiple tags of interest included in the global tagging pattern.
[0149] In some cases, the method also includes actions to modify the model graph based on information associated with the data cues. For example, some systems identify one or more tags of interest (ROIs) that require the generation of label values based on the data cues. The system also identifies a subset of functions included in the model graph that are associated with the one or more ROIs. Before applying the model graph to the data cues, the system modifies the model graph instance by omitting one or more functions that are not included in the identified subset of functions associated with the one or more ROIs.
[0150] Some systems generate multiple model graph instances based on data cue word segmentation. For example, some methods involve segmenting the data cue word into multiple segments and generating multiple model graph instances. Each model graph instance corresponds to a different segment among the multiple segments of the data cue word. The system then applies each model graph instance to its corresponding segment. Subsequently, the system generates multiple intermediate outputs, where each intermediate output includes one or more label values of one or more tags of interest associated with its corresponding data cue word segment. Because these outputs correspond to specific segments of the data cue word rather than the entire cue word, they are considered intermediate outputs. Therefore, the system generates a final output for the data cue word based on the combination of these intermediate outputs.
[0151] It should be understood that the system can segment different types of data prompts, including complete prompts generated by a large language model based on the initial input prompts.
[0152] As the number of model graph instances increases, the system is also configured to process data processing requests associated with different data cues and fragments of those cues by batching data processing requests for specific nodes in the model graph. For example, some methods also include actions that identify batching criteria for specific nodes included in the model graph. The system also identifies and routes one or more data processing requests for a specific node to the batching cache or queue corresponding to that specific node.
[0153] The system determines whether batch processing criteria have been met. In response to determining that batch processing criteria have not been met, the system prohibits the transmission of one or more processing requests in the batch processing cache; alternatively, in response to determining that batch processing criteria have been met, the system schedules and routes batches of one or more processing requests from the batch processing cache / queue to the specific node corresponding to that batch processing cache for processing.
[0154] Now let's turn our attention to... Figure 23 It illustrates an example embodiment with flowcharts having multiple actions (e.g., actions 1010, 1020, 1030, 1040, 1050, 1060, 1070, 1080, and 1090) that are associated with a method implemented by a computing system (e.g., computing system 110) for performing content moderation on segmented data prompts (e.g., fragments 118) using a model diagram (e.g., model diagram 116).
[0155] The first action shown includes accessing a model graph (action 1010) derived from a metamodel topology (e.g., metamodel topology 115) that comprises multiple functions. This model graph includes multiple nodes representing unique functions configured to perform operations on input data and to generate label values corresponding to one or more labels of interest in a global label pattern associated with the model graph. The system also receives a data cue word (e.g., data cue word 117) (action 1020) and segments the data cue word into multiple fragments (e.g., fragment 118) (action 1030). By segmenting the data cue word into multiple fragments, the system can fine-tune and customize the processing of the data cue word based on each fragment.
[0156] For each segment of the data prompt, the system executes a series of actions (action 1040, action 1050, action 1060, action 1070, and action 1080). For example, for a specific segment of the data prompt, the system identifies the policy information included in that data prompt (e.g., Figure 7 The loaded strategy specifies the subset of nodes in the model graph that should be used when processing the fragment, and generates an instance of the model graph (action 1040). A model graph instance (e.g., orchestrator 318) is generated for each fragment (action 1050).
[0157] Therefore, based on the policy information, the system then prunes the model graph instance to include the subset of nodes specified in the policy information, so that the model graph instance now omits at least one node from the previously visited model graph (action 1060) (see...). Figure 14-17 In some cases, within the model graph instance generated for a specific fragment, a depth-first search is performed to identify the subset of nodes specified for that specific fragment in the policy information.
[0158] After pruning the model graph instance, the system applies the model graph instance to its corresponding fragment (action 1070) and generates intermediate output, which includes the tag values of one or more tags of interest associated with a specific fragment among multiple fragments of data prompt words (action 1080).
[0159] Finally, after generating all intermediate outputs, the system generates the final output (action 1090) for the data cue words based on the combination of intermediate outputs generated for multiple segments.
[0160] In some cases, pruning the model graph of a specific segment among multiple segments also includes: identifying one or more interest labels from policy information associated with the specific segment; identifying one or more nodes of the model graph that are not associated with the one or more interest labels from the policy information; and omitting one or more nodes of the model graph that are not associated with the one or more interest labels identified from the policy information in the model graph instance.
[0161] and Figure 23 Some related methods also include actions such as: identifying one or more additional nodes that can be skipped during runtime processing of data cues using the model graph, which are included in the pruned data audit graph instance; and skipping the identified one or more additional nodes when processing a specific segment among multiple segments using the pruned data audit graph instance.
[0162] Data cues can be generated from various sources. For example, data cues can include initial cues generated by users (human or otherwise); large language models that have received initial cues generated by users; or even complete cues generated by large language models based on electronic content obtained from various sources.
[0163] Policy information can also be configured or identified according to different embodiments. For example, in some cases, policy information is user-defined and appended to data prompts. Alternatively, policy information can be automatically generated based on analyzing data prompts to determine which tags of interest apply to those prompts. In other cases, policy information is selected from multiple stored policies, each associated with a specific user or enterprise. It should be understood that policy information can originate from a combination of the embodiments described above.
[0164] and Figure 23 Some related methods also include actions for identifying the final output of the entire data cues based on a combination of intermediate outputs from each model graph instance generated for multiple fragments of the data cues. Based on the final output (i.e., one or more label values of one or more tags of interest for the data cues), the system modifies multiple fragments of the data cues by removing portions of one or more fragments of the multiple fragments before displaying them to the user's display.
[0165] After generating the final output from multiple model graph instances, data prompts are displayed to the user on the user's display based on the hardware configuration and model interface utilized by the user system. For example, in some embodiments, data prompts are displayed to the user as text output in an LLM format or a format provided by other models applied to the user prompts. In other embodiments, data prompts are displayed on the user interface as modified or annotated prompts, which include annotations generated by the editor 318 and / or modified based on the annotations generated by the editor 318.
[0166] As an example, the system can determine that the tag value of a data hint fragment is equal to or exceeds a predetermined threshold, and identify one or more words in the data hint fragment that are associated with the violation tag value. Then, before displaying the hint (or hint fragment) to the user, the system can modify the data hint by removing or replacing one or more of the identified violation words in the data hint fragment, and display the modified version of the hint or hint fragment on the user interface / display.
[0167] In some cases, data cues comprise multiple individual segments that can be processed one by one by the orchestrator instances. In these cases, the system identifies the final output from multiple model graph instances 320 based on a combination of intermediate outputs generated for each segment of the data cues. When combined, the final output includes label values for one or more tags of interest and annotations for labeling one or more segments of the data cues. Subsequently, before displaying the data cues to the user, the system modifies the data cues by annotating one or more segments of the data cues using the label values of one or more tags of interest identified in the final output. The modified final output can then be displayed on the user's display as a composite modified data cues.
[0168] As described above, data cues analyzed by metamodel topology are sometimes modified before being transmitted / displayed on the client-facing system. For example, the system may determine that the label value of a specific segment of a data cue (i.e., the output from an instance of the model graph) is equal to or exceeds a predetermined threshold. The system may also identify one or more words in a specific segment of the data cue that are associated with label values exceeding the threshold. Then, before displaying the data cue segment to the client-facing system, the system can modify the data cue segment by removing or replacing one or more identified words, and display the modified data cue segment at the user interface of the client-facing system.
[0169] It should be understood that the metamodel topology advantageously includes a single input data entry point (e.g., via a reverse proxy), allowing users to interface with this single entry point without connecting to intermediate or low-level function interfaces included in the metamodel topology. By implementing the system in this way, the user experience is improved and simplified because intermediate outputs are abstracted and hidden from the user interface. Even more generally, the system can be implemented such that it can be directly invoked by any client using a bidirectional streaming interface or as a simple non-streaming request-response interface (e.g., an HTTP interface). For example, in non-LLM scenarios, some third-party developers could directly invoke the service for content moderation purposes.
[0170] In some cases where multiple instances of a model graph are generated for different fragments, the system identifies batching criteria for specific nodes included in the different instances of the model graph. The system also identifies one or more processing requests for specific nodes across different instances of the model graph and routes one or more processing requests to a batching cache corresponding to the specific node appearing in one or more fragment-based model graph instances. The system then determines whether the batching criteria have been met. In response to determining that the batching criteria have not been met, the system prohibits the transmission of one or more processing requests from the batching cache to the specific node; alternatively, in response to determining that the batching criteria have been met, the system routes one or more processing requests as a batch to the specific node for processing.
[0171] In certain situations, where multiple data processing requests for different nodes are received across different model instances generated for each fragment, the system can automatically expand the model graph (and / or individual nodes) to improve the processing efficiency of different data requests. Automatic expansion refers to the process of proactively instantiating instances of the model graph (and / or individual nodes) based on the predicted number of instances required to efficiently execute the received data processing requests for a specific node. For example, the system predicts the number of instances of a specific function / model associated with a specific node needed to process the data input based on how many times a specific node is retained across multiple instances of the model graph generated for multiple fragments of data prompts. Subsequently, the system automatically expands the model graph based on the predicted number of instances of the specific function to provide the predicted number of instances of the specific function for the specific node. This can occur during runtime or before runtime.
[0172] Intermediate outputs for different instances of the model graph include multiple distinct label values. For example, some intermediate outputs for a specific instance of the model graph may include binary label values, while others may include severity level label values. Furthermore, different label values are associated with different tags of interest. For instance, one or more tags of interest may include hate speech warning labels, pornography warning labels, violence warning labels, or PPI warning labels, among others.
[0173] The final output of the data cue is generated by combining intermediate outputs from different instances of the model graph, such that the final output of the entire data cue includes multiple label values corresponding to multiple tags of interest. Some label values correspond to the same tag of interest (i.e., the same binary label value and severity level label value).
[0174] Now let's turn our attention to... Figure 24 It illustrates an example embodiment with a flowchart of multiple actions (e.g., actions 1110, 1120, 1130, 1140, 1150, and 1160) associated with a method for dynamic batch processing implemented by a computing system (e.g., computing system 110) to perform prompt word queries (e.g., data prompt word 117) using a model graph (e.g., model graph 116).
[0175] The first action shown includes actions for identifying one or more instances of the model (e.g., instance 702, instance 704, etc.), which includes different processing nodes (action 1110) configured to perform different functions on the input data. It should be understood that the model can be configured according to different applications. For example, according to the embodiments described herein, in one application, the model is represented by a model graph configured as a data audit graph.
[0176] The system also identifies at least one batching criterion (e.g., batching criterion 726) for a specific node (e.g., node f5) across different processing nodes (Action 1120). Examples of batching criteria include: waiting to receive a minimum number of processing requests before transmitting a batch of processing requests based on the maximum number of processing requests a specific node can handle; waiting a predetermined amount of time between receiving batch processing requests; or other batching criteria. Batching criteria include thresholds that trigger the transmission of a batch of processing requests to a specific node. During runtime, the system identifies one or more processing requests for a specific node (e.g., data processing request 714, data processing request 716, etc.) and routes them to the batching cache corresponding to that specific node (e.g., batch 724) (Action 1130). The system then periodically (e.g., after receiving each data processing request or after a predetermined time interval) determines whether the batching criteria have been met (Action 1140).
[0177] In response to determining that batch processing criteria have not been met, the system prohibits the transmission of one or more processing requests (e.g., batch 724) in the batch cache to the specific node (action 1150). Alternatively, in response to determining that batch processing criteria have been met, the system routes one or more processing requests (e.g., batch 724) as a batch to the specific node for processing (action 1160). After transmitting the batch request, the system subsequently processes the batch containing the one or more processing requests at the specific node.
[0178] By batching different processing requests, the system can adjust the processing time for each iteration of the input data and improve the processing efficiency of the hardware that accommodates different instances of the model. These technical benefits can be realized, especially when the computing system includes GPUs capable of efficiently batching data processing requests.
[0179] In some cases, a batch corresponds to one or more requests received from multiple instances of a model. Additionally or alternatively, each of the multiple instances of the model is generated for a specific fragment of the input dataset. In some systems, a batch corresponds to one or more requests received from multiple users. Additionally or alternatively, a batch corresponds to one or more requests received from multiple enterprises.
[0180] The system can identify various batch processing criteria. For example, in some cases, batch processing criteria are based on the minimum or maximum number of data processing requests, and / or on the maximum wait time between received data processing requests. In other cases, depending on the model's processing specifications, the maximum wait time is less than one millisecond, or even less than a few milliseconds.
[0181] and Figure 24 Some related methods include additional actions for automatically scaling the model / its layers to improve processing efficiency for different batches of data processing requests. In these actions, the system predicts the number of instances of a specific function / model associated with a specific node required to process the data input, based on how many times a specific node is retained across multiple instances of the model. The system then automatically scales the model based on the predicted number of instances of the specific function to provide the predicted number of instances of the specific function for the specific node at runtime.
[0182] In some systems, automatic scaling is performed by instantiating multiple instances of the model for different fragments of data input. Additionally or alternatively, multiple instances of the model are instantiated for multiple data inputs from different users within the enterprise. In some systems, multiple instances of the model are instantiated for multiple data inputs from different enterprises.
[0183] When instantiating an instance of the model for different segments of the data input, the method includes actions for segmenting the input data prompts into multiple segments.
[0184] Then, for each of the multiple segments, the system is configured to: (i) identify strategy information included in the input prompt, which specifies a subset of nodes in the data review graph to be utilized when processing a particular segment of the multiple segments; (ii) generate an instance of the model; (iii) prune the instance model to include the subset of nodes specified in the strategy information, such that the instance of the model omits at least one node of the accessed data review graph; and (iv) generate intermediate outputs, which include one or more tags of interest associated with a particular segment of the multiple segments. Finally, the system generates a final output based on the combination of each intermediate output generated for the multiple segments.
[0185] The segmented data cues (i.e., the input data for the meta-model topology) include different formats. For example, some input data includes different electronic content, including data cues. Some data cues include initial cues generated by the user, while others include complete cues generated by a large language model based on the received initial cues generated by the user.
[0186] Similar to automatically scaling up calls to different nodes of local or remote models, the system can also predict the number of model instances required to process batches of data processing requests based on the amount of input data and automatically scale up the number of model instances based on the predicted number of model instances to provide the predicted number of model instances at runtime.
[0187] In light of the foregoing, it should be understood that the disclosed embodiments offer numerous technical advantages over conventional systems, methods, and frameworks. For example, the systems and methods described herein advantageously facilitate the distribution of service components necessary for orchestrating the evaluation of a growing number of models. This evaluation also extends to tags and policies for content moderation schemes. This framework also creates definitions for interfaces, protocols, and libraries, enabling teams across organizations to deploy their models into production using procedural and uniform specifications.
[0188] The disclosed embodiments also improve the user experience by abstracting away the large number of models, tags, and taxonomies created for different modalities, such as voice, text, and images. Such embodiments allow these models to be combined into a consistent set of content moderation configuration objects and policies in a simple, "no-code" manner. Therefore, for any changes to the models, the underlying network is abstracted away from the user, who is presented with a simplified user interface corresponding to a single, holistic content moderation metamodel. Using such embodiments, users can leverage a higher-quality system without needing to understand the details behind each model included in the system.
[0189] In some embodiments, the system includes a single model, such as an LLM model or a generative pre-trained (GPT) model. Alternatively, the system includes multiple combined models. In either case, the models(s) included in the system conform to a global labeling pattern that includes predefined labels and a set of taxonomies. Advantageously, any backend changes to the models(s) of the system will not affect the simplicity of user interaction with the frontend interface.
[0190] In summary, the disclosed embodiments facilitate close collaboration among multiple different model owners, provide standardization of labels and taxonomy across models, offer flexibility in the combination of models used when deploying metamodels for specific users, and facilitate low-latency execution. Example computing system
[0191] It should be understood that the disclosed embodiments may include: being practiced or implemented by a computer system configured with a computer storage device storing computer-executable instructions that, when executed by one or more processing systems (e.g., one or more hardware processors) of the computer system, cause various functions to be performed, such as the actions described above.
[0192] Embodiments within the scope of this disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available medium accessible by a general-purpose or special-purpose computer system. Computer-readable media storing computer-executable instructions are physical storage media. Computer-readable media carrying computer-executable instructions are transmission media. Therefore, by way of example, embodiments of this disclosure may include at least two distinct types of computer-readable media: physical computer-readable storage media and transmission computer-readable media.
[0193] Physical computer-readable storage media include random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), optical disc ROM (CD-ROM), or other optical disc storage devices (such as optical discs (CD), digital video discs (DVD), etc.), magnetic disk storage devices or other magnetic storage devices, or any other hardware storage device that can be used to store required program code in the form of computer-executable instructions or data structures and is accessible by a general-purpose or special-purpose computer.
[0194] When information is transmitted or provided to a computer via a network or other communication connection (hardwired, wireless, or a combination of hardwired and wireless), the computer treats that connection as a transmission medium. The transmission medium may include a network and / or a data link, which may be used to carry the required program code in the form of computer-executable instructions or data structures, and may be accessible by a general-purpose or special-purpose computer. Combinations of the above are also included within the scope of computer-readable media.
[0195] Furthermore, upon arrival at various computer system components, program code in the form of computer-executable instructions or data structures can be automatically transferred from the transmission computer-readable medium to the physical computer-readable storage medium (and vice versa). For example, computer-executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (e.g., a network interface card (NIC)) and then ultimately transferred to the computer system RAM and / or a less volatile computer-readable physical storage medium at the computer system level. Therefore, computer-readable physical storage media can be included in computer system components that also (or even primarily) utilize the transmission medium.
[0196] Computer-executable instructions include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a specific function or group of functions. For example, computer-executable instructions can be binary code, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or actions described above. Rather, the described features and actions are disclosed as exemplary forms of implementing the claims.
[0197] Those skilled in the art will understand that this invention can be practiced in network computing environments with a variety of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics devices, network PCs, minicomputers, mainframes, mobile phones, PDAs, pagers, routers, switches, etc. This invention can also be practiced in distributed system environments, where local and remote computer systems connected via network links (via hardwired data links, wireless data links, or a combination of hardwired and wireless data links) perform tasks. In a distributed system environment, program modules can reside in local and remote memory storage devices.
[0198] Alternatively or additionally, the functions described herein may be performed at least in part by one or more hardware logic components. For example, illustrative types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), etc.
[0199] For clarification, unless otherwise specified, the terms "set" and "subset" are intended to exclude the empty set; therefore, a "set" is defined as a non-empty set, a "superset" as a non-empty superset, and a "subset" as a non-empty subset. Unless otherwise specified, the term "subset" does not include the entirety of its supersets (i.e., a superset contains at least one item not included in the subset). Unless otherwise specified, a "superset" may include at least one additional element, and a "subset" may exclude at least one element.
[0200] The invention may be embodied in other specific forms without departing from the spirit or characteristics thereof. The described embodiments are to be considered illustrative rather than restrictive in all respects. Therefore, the scope of the invention is indicated by the appended claims, and not by the foregoing description. All modifications within the equivalent meaning and scope of the claims are included within their scope.
Claims
1. A method implemented by a computing system for managing batch processing of data processing requests, the method comprising: Identify one or more instances of a model, the model including different processing nodes configured to perform different functions on input data; Identify batch processing criteria for a specific node among the different processing nodes, the batch processing criteria including a threshold for triggering the transmission of a batch processing request to the specific node; Identify one or more processing requests for the specific node, and route the one or more processing requests for the specific node to the batch cache corresponding to the specific node; Determine whether the batch processing criteria have been met; In response to determining that the batch processing criteria are not met, transmission of the one or more processing requests in the batch processing cache to the specific node is prohibited; Alternatively, in response to determining that the batch processing criteria have been met, the one or more processing requests are routed as a batch to the specific node for processing.
2. The method according to claim 1, further comprising: The batch of one or more processing requests is processed by the specific node.
3. The method of claim 1, wherein the batch corresponds to one or more requests received from a plurality of instances of the model.
4. The method of claim 1, wherein each of the plurality of instances of the model is generated for a specific segment of the input dataset.
5. The method of claim 1, wherein the batch corresponds to one or more requests received from a plurality of users.
6. The method of claim 1, wherein the batch corresponds to one or more requests received from a plurality of enterprises.
7. The method according to claim 1, wherein the model is a data audit graph.
8. The method of claim 1, wherein the batch processing criterion is based on the maximum number of data processing requests.
9. The method of claim 1, wherein the batch processing criterion is based on the maximum waiting time between received data processing requests.
10. The method of claim 9, wherein the maximum waiting time is less than one millisecond.
11. The method of claim 1, further comprising: The number of instances of a specific function associated with the specific node that will be needed to process the data input is predicted based on the number of times the specific node is retained in the plurality of instances of the model. as well as The model is automatically expanded based on the predicted number of instances of the specific function to provide the predicted number of instances of the specific function for the specific node at runtime.
12. The method of claim 11, wherein the plurality of instances of the model are instantiated for different segments of the data input.
13. The method of claim 11, wherein the plurality of instances of the model are instantiated for multiple data inputs from different users across an enterprise.
14. The method of claim 11, wherein the plurality of instances of the model are instantiated for multiple data inputs across different enterprises.
15. The method according to claim 14, wherein the input prompt word is divided into multiple segments; For each of the plurality of segments: (i) Identifying strategy information included in the input prompt, the strategy information specifying a subset of nodes of the data audit graph to be utilized when processing a particular segment of the plurality of segments; (ii) Generate an instance of the model; (iii) Prune the instance model to include the subset of nodes specified in the policy information, such that the instance of the model now omits at least one node of the accessed data audit graph; as well as (iv) Generate intermediate output, the intermediate output including one or more tags of interest associated with a specific segment among the plurality of segments; as well as The final output is generated based on the combination of each intermediate output generated for the multiple fragments.