Systems, methods, and computer-accessible medium for providing and / or generating characteristic analysis of genai systems
The system analyzes and refines GenAI outputs by identifying prompt units causing specific characteristics, allowing for effective control and improvement of GenAI systems.
Patent Information
- Application Number
- PCT/US2025/013225
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-26
- Filing Date
- 2025-01-27
- Publication Date
- 2025-07-31
AI Technical Summary
Current systems lack the ability to analyze and explain the characteristics of generative AI (GenAI) systems, such as bias, hate speech, and content moderation, which are crucial for various applications including document assessment and chat systems, and there is no known method to identify and control undesirable outputs.
A system and method that provides characteristic analysis of GenAI systems by determining the units within a prompt that cause specific characteristics in the output, assigning scores, and iteratively refining prompts to achieve target scores, while explaining the characteristics and potentially reprogramming the system based on these analyses.
Enables effective identification and control of undesirable outputs in GenAI systems by providing detailed explanations and iterative prompt refinement, enhancing the system's performance and adherence to design goals.
Smart Images

Figure US2025013225_31072025_PF_FP_ABST
Abstract
Description
SYSTEMS, METHODS, AND COMPUTER-ACCESSIBLE MEDIUM FOR PROVIDING AND / OR GENERATING CHARACTERISTIC ANALYSIS OF GENAI SYSTEMSCROSS REFERENCE TO RELATED APPLICATION(S)
[0001] This application relates to U.S. Provisional Patent Application No. 63 / 625,444, filed on January 26, 2024, the entire disclosure of which is incorporated herein by reference.FIELD OF THE DISCLOSURE
[0002] The present disclosure relates generally to analysis of artificial intelligence (Al) systems, and more particularly to methods, systems, and computer-accessible medium for providing and / or generating characteristic analysis of GenAI systems. In particular, the present disclosure provides (i) explanations for why GenAI systems generate outputs with certain characteristics, and (ii) exemplary7mechanisms to explore (‘Ted team”) the generation behavior of GenAI system(s) to discover inputs that lead to the generation(s) of output(s) with certain exemplary characteristics.BACKGROUND INFORMATION
[0003] Increasingly, it may be interesting and / or important to analyze GenAI systems and their outputs to decide whether they exhibit certain characteristics, to determine what inputs produce outputs with certain characteristics, and to explain how the input that can lead the genAI system to produce output exhibiting those characteristics. Most commonly, the characteristics of interest can include whether the outputs are biased, whether they are fabricated, whether they contain hate speech, whether they violate policies or otherwise require content moderation actions, etc. Various applications may require an identification of task-specific characteristics. For example, if the GenAI system is requested to provide assessments of documents, there may be interest in whether the assessments satisfy some properties, such as fidelity or correctness. For a deployed chat system, it may be interesting to assess whether the system and / or outputs thereof exhibits certain personality characteristics, such as empathy, and / or to measure it along the Big Five personality traits. It is also useful to characterize generated output as being “Al generated”, for example to help identify' and control student cheating. However, no such system, method or computer-accessible medium is known to exist.
[0004] Thus, it may be beneficial to provide exemplary systems, methods and computer- accessible medium can overcome at least some of the deficiencies described herein above.SUMMARY OF EXEMPLARY EMBODIMENTS
[0005] For example, in view of the above, once it is being decided whether the system exhibits one or more such characteristics, systems, methods and computer-accessible medium that incorporate the GenAI system can be provided to facilitate the above preferences and / or needs.
[0006] To that end, it is possible to provide exemplary systems, methods and non- transitory computer-accessible medium according to exemplary embodiments of the present disclosure, which can provide and / or generate characteristic analysis of GenAI systems.
[0007] In some exemplary' embodiments of the present disclosure, exemplary systems, methods, and computer accessible medium can be provided which can supply a prompt to a genAI system, determine at least one characteristic in an output generated by the genAI, and determine at least one unit within the prompt that caused the characteristic to be present in the output. The unit(s) can comprise one or more of a token, a word, a sentence, an n-gram, and / or a natural language parsing object. Further, the characteristic(s) can be assigned a score, and the score can represent a degree to which the output exhibits the characteristic(s). The score can be used as an input to a search procedure, and the search procedure can determine at least one replacement unit to create and / or generate a subsequent prompt.
[0008] In further exemplary embodiments of the present disclosure, exemplary systems, methods, and computer accessible medium, the subsequent prompt can be supplied to the genAI system and a subsequent score can be assigned to the at least one characteristic in a subsequent output generated by the genAI system on the subsequent prompt. The subsequent score can be compared against the score to determine whether the at least one replacement unit effected the at least one characteristic. Additionally, the exemplary' process can be utilized to determine at least one replacement unit, create / generate a subsequent prompt with the unit(s), supply the subsequent prompt to the gen Al system, assign a subsequent score, and compare scores is iterated until a target score is reached. Furthermore, the search procedure can output at least one explanation for the characteristic(s), and the explanation(s) can comprise the replacement unit(s). The genAI system can be reprogrammed based on the explanation(s).
[0009] In some exemplary embodiments of the present disclosure, exemplary systems, methods, and computer accessible medium can be provided which for explaining generativeAl (“genAI”) system behavior, including supplying a prompt to a genAT system comprising an LLM, a generated output, and a characteristic classifier, and determining, by the characteristic classifier, at least one characteristic of the generated output and a score that is above a predetermined value. The score could also be below a predetermined value. Further, the value can comprise a threshold based on the characteristic.
[0010] These and other objects, features and advantages of the exemplary embodiments of the present disclosure will become apparent upon reading the following detailed description of the exemplary embodiments of the present disclosure, when taken in conjunction with the accompanying claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Further objects, features and advantages of the present disclosure will become apparent from the following detailed description taken in conjunction with the accompanying Figures showing illustrative embodiments of the present disclosure, in which:
[0012] Figure 1 is a block / functional diagram of an exemplary GenAI system applying a machine-learned model iteratively to generate its output according to an exemplary embodiment of the present disclosure;
[0013] Figure 2 is a block / functional diagram of an exemplary embodiment of an architecture according to the present disclosure;
[0014] Figure 3 is a block / functional diagram system for generating GenAI characteristic explanations according to an exemplary embodiment of the present disclosure;
[0015] Figure 4 is a block / functional diagram showing the details of a Search Algorithm for the system of Figure 3 according to an exemplary embodiment of the present disclosure;
[0016] Figure 5 is block / functional diagram of another exemplary' embodiment of the architecture providing an exemplary characteristic classification according to the present disclosure; and
[0017] Figure 6 is a block / functional diagram of an exemplary system which is configured, designed and / or programmed to implement the exemplary architectures and / or methods according to various exemplary' embodiments of the present disclosure.
[0018] Throughout the drawings, the same reference numerals and characters, unless otherwise stated, are used to denote like features, elements, components or portions of the illustrated embodiments. Moreover, while the present disclosure will now be described in detail with reference to the figures, it is done so in connection with the illustrativeembodiments and is not limited by the particular embodiments illustrated in the figures and the appended claims.DETAILED DESCRIPTION OF EXEMPLARY EMBODIMENTS
[0019] The following description of exemplary embodiments provides non-limiting representative examples referencing numerals to particularly describe features and teachings of different aspects of the present disclosure. The exemplary embodiments described should be recognized as capable of implementation separately, or in combination, with other exemplary embodiments from the description of the exemplary embodiments. A person of ordinary skill in the art reviewing the description of the exemplary embodiments should be able to learn and understand the different described aspects of the present disclosure. The description of the exemplary embodiments should facilitate understanding of the exemplary embodiments of the present disclosure to such an extent that other implementations, not specifically covered but within the knowledge of a person of skill in the art having read the description of embodiments, would be understood to be consistent with an application of the exemplary embodiments of the present disclosure.
[0020] Generative Al systems are different from traditional machine-leaming-based Al systems in that, e.g., they do not simply produce an estimate of a number (for example, how much will this customer be worth to us) or an estimate of a likelihood (for example, what is the probability7that this web page contains objectionable content). Instead, they “generate"’ complex outputs, such as sentences, paragraphs, or documents of text, pictures, videos, audio (e.g.. music), and so on. This can be done by applying a traditional machine-learned model repeatedly. Figure 1 shows a block functional diagram of an illustrative example of this process. In the exemplary process in Figure 1, the GenAI system 105 is the left (outer) box. The generated output 160 (e.g., a text document) is represented by the right-most box. The exemplary GenAI system can incorporate a machine-learned model, e.g.. Model 130, and can use it iteratively, as represented by the arrow pointing back to the left. More specifically, the GenAI system 105 can accept as input a prompt 110 from the user, as well as (e.g., usually) some additional input (e.g., in the same format as the prompt) that the system designers add, but the user may not see. e.g., the context 120. The explanations can be desired, e.g., only for the original prompt 110. although they could apply to the context 120 or a combination of prompt 1 10 and context 120. The input to the Model will be referred to as the prompt henceforth.
[0021] The exemplary Model 130 can accept the prompt as input and can produce as output a set of scores over all the possible “tokens” in the output. These can be considered as a probability distribution over tokens, although the scores may not be, e.g., valid probabilities. That may not really matter. This is represented by Pred 140. for prediction. The next step in the exemplary text generative process is represented by Next Word Logic 150. This exemplary logic can take the output probability distribution (Pred 140), and use it to select the next “word” to add to the generated output 160. That “next word” is appended to the prompt and the process iterates with this augmented prompt, producing the next next word and so on. At some particular point, the next-word logic 150 can select the “End of Transmission” token (or some other stopping logic will be hit) and the system will produce the sequence of generated words as output 160.
[0022] Prior work, such as, e.g., by Martens & Provost (2014) and further work thereafter, have provided “counterfactual” explanations to explain the input / output behavior of standard (non-generative) Al models, including those produced by machine learning systems. These methods do not apply to generative Al systems, because those systems do not provide a classification or a classification score — instead they provide generated text or images or sound, etc. The exemplary systems, methods and computer-accessible medium according to an exemplary embodiment of the present disclosure can treat as one single component, the multi-component system, including the (itself multi-component) generative Al system, its generated output, and a Characteristic Classification Model (CCM) component, that assesses whether and / or to what degree the generated output exhibits the characteristic of interest, as illustrated in Figure 2 and described in detail herein. Technically, a generative Al system for text operates on “tokens” instead of words. Tokens can include words, and also other aspects of text, like punctuation, coding characters, syllables, and even letters. “Words” can be used in the present disclosure to provide a further exemplary explanation, and in no way limiting.
[0023] The CCM of the exemplary systems, methods and computer-accessible medium according to an exemplar}7embodiment of the present disclosure can comprise a machine learned model itself (e.g., from a different ML process). According to other exemplar}' embodiments, the CCM can be or include a set of rules, or take some other form (for example, all big social media platforms have content moderation systems that combine Al models and large teams of humans). The CCM component of the exemplary systems, methods and computer-accessible medium according to an exemplary' embodiment of thepresent disclosure can produce a numeric “score” as an output. This exemplary score can usually be interpreted as the degree to which the input to the CCM component (e.g., the generated output of the genAI sy stem) exhibits the characteristics. Thus, e g., if the characteristic is: “is the output toxic?”, a toxicity CCM component can produce a score that can be interpreted as either an estimated level of toxicity or an estimated likelihood score (e.g., a probability estimate) that the particular Al-generated output is toxic. This exemplary' CCM output score can be compared to a threshold for decision making. For example, if the estimated probability of toxicity is greater than 0.3, then some action can be taken (such as, e.g., investigating why more deeply).
[0024] The term “counterfactual” can be used in many different ways in Al / machine learning / data science. It is a general term that can mean, e.g., “what would have happened otherwise.” Thus, when performing the causal estimation one can employ a method (there may be a lot of them) for estimating the counterfactual. For example, in order to estimate the efficacy of a medical treatment, it can be beneficial to examine outcomes when someone was given the treatment, with what would have happened had they not been given the treatment. Those can be called, e.g., the two counterfactuals, and more specifically, e.g., the outcome that is not observed is often called the counterfactual outcome. The term “counterfactual explanations” can cause confusion because the term may be used not to discuss causality “in the real world,” and instead may refer to a very specific causal question about the behavior of a system. For example, the counterfactual in “counterfactual” explanations (and there can be a large number of different variants now) can be: what would have been the output of the system if it had been given this other (usually modified) input, rather than the original input. That is still a causal analysis, although the subject of the analysis is the Al model behavior. The exemplary' systems, methods and computer-accessible medium according to an exemplary embodiment of the present disclosure can extend this notion to the behavior of an entire generative Al system plus the CCM.
[0025] Original work on counterfactual explanations can be found in the explanations of the patent publications by Martens and Provost, and a related published paper (e.g., MISQ 2014), U.S. Pat. 9,836,455, the entirety' of which is incorporated by reference herein. Those exemplary systems, methods and configurations can be applied here, e.g., the large system (GenAI system characteristic classifier 200 shown in Figure 2) can be considered to be the unit of analysis — the thing being explained.
[0026] The exemplary systems and methods described in U. S. Patent No. 9,311 ,599, which is incorporated herein by reference in its entirety, can be applicable to the exemplary GenAI system characteristic classifier 200 shown in Figure 2. In the exemplary' systems, methods and computer-accessible medium according to an exemplary’ embodiment of the present disclosure, humans may not be challenged just to find inputs where the system makes a simple error, but more generally humans can be challenged to find inputs that give outputs that exhibit a particular characteristic (as indicated by the CCM output) or particularly high CCM output scores. In many situations, this can still retain the essence of "‘Beat the Machine” concepts described in the patent identified above, in those cases (e.g., like toxicity) where a GenAI system was designed not to exhibit the characteristic (e.g., not to be toxic). In these exemplary' cases, producing toxic output can be considered an error that the humans are trying to find — just not such a simple predictive model error as in prior work. In particular, the additional "‘characteristic classification” machinery can be used to get an assessment automatically of whether the generated output indeed exhibits this erroneous behavior. Similarly, consider a model that is designed so that its output can be identified as being Al generated (for example, to help detect student cheating). Finding inputs that lead to generated outputs for which the CCM (which in this specific example is an “is it Al generated detector”) produces a low score (i.e., probably not Al generated) would again be an error that developers would prefer to locate and deal with. On the other hand, there are applications where the goal could be to find desirable characteristics. For example, the CCM may be a “lyricism” classifier, the goal is to figure out how to prompt the system to generate lyrical poetry.
[0027] The current system is related to the Beat the Machine (BtM) system in the abovereferenced patent by making the following replacements to the BTM system, all referring to Fig. 1 in that prior patent. Anywhere the prior patent refers to error, classifier error, predictive model(ing) error, incorrectly classified, etc., it can be replaced with “target characteristic score”; the intention is that the system / method is not just looking for errors but more generally for generated outputs that when input to the characteristic classifier produce a score that satisfies one or more specified properties. Examples of these properties include: being above a threshold, being below a threshold, being in a certain range, equaling a particular value, etc.
[0028] In box 120 in Fig. 1 , the instance received would be a prompt for the LLM-based system, that corresponds to a generated output by the genAI system and subsequently to a classification or score output by the characteristic classifier.
[0029] In box 130 the classification retrieved from the predictive model, along with its confidence, would now be whether or not the characteristic classification score satisfies the one or more specified properties, along with the score itself replacing the confidence.
[0030] In box 140, determining that the instance has been incorrectly classified would be replaced with determining that the instance produces a characteristic output with the one or more specified properties (the generated output being incorrect could be one characteristic output; the score would represent its confidence). The offensive content example from box 140 is illustrative. Now instead of looking for webpages that a simple classifier classified incorrectly as not offensive (for example), the system / methods would be looking for prompts that lead the GenAI system to generate outputs that then get characterized as being offensive.
[0031] In box 150, the payment would be based on the characteristic score. So for example, the annotator user might get paid more if he found prompts that led to generated outputs that were scored as being more offensive.
[0032] And in box 170, the training would be to train and improve the GenAI system, for example via adding the example to the training data of an internal large language model, using the example to feed a reinforcement learning with human feedback (RLHF) process, or other methods for training and improvement from data.
[0033] The example given below shows an example prompt that a user suggested, from which the specified GenAI system generated the output shown, which then received a high characteristic score for toxicity — the property of interest.
[0034] Turning to Fig 17 of the prior patent, once again instead of annotating URLs or other documents, the annotator user is supplying prompts, and would have a corresponding interface (as with box 1710 and 1720). The Classification Model (box 1740) would be replaced with the entire characteristic analysis system (box 200 in Fig 2). The Novelty of Error model (1750) could then be replaced with aNovely of Prompt model, as it may not make sense to reward users for submitting ver7similar prompts once they found one that satisfied a specified characteristic property. Similarly, the system could incorporate a Prevalence model for prompts, analogous to the BtM Prevalence model (box 1760).
[0035] In addition, as shown in Fig 5, with the recent advances of large-language models and GenAI systems, the human annotator that comes up with prompts to “beat the machine”can be replaced by another GenAI system. This GenAI system itself could be rewarded along the same lines as the humans were, albeit the actual rewards would be different. (Indeed, for the GenAI system one option would be simply to reward with a numeric reward or point system, since the GenAI system can be programmed to maximize rewards, and in that case monetary rewards would not be necessary.)
[0036] Other prior work for explanation focuses on token level explanation, i.e., what led to the production of the next word. Among the most popular of these methods are SHAP and LIME (see Mosca et al., 2022 and Mardaoui and Garreau 2022). These do not involve a CCM or could only be viewed as a trivial CCM which is the next token. Another avenue of prior work is on adversarial networks which often have a generator / discriminator architecture (see de Rosa and Papa, 2023). The purpose of these methods is not explanation, but instead model parameter tuning.
[0037] The exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure can be used with an exemplar}’ architecture to provide and / or generate characteristic analysis of GenAI systems.
[0038] In one exemplary embodiment of the exemplary systems, methods and computer- accessible medium according to exemplary embodiments of the present disclosure, it is possible to assume that it is being decided whether, and / or to what degree, individual generated outputs of the Al system (e.g., GenAI system) exhibit the characteristics. When it is preferred to draw conclusions about the system behavior more generally, it is possible to aggregate over the decisions about some collection of specific outputs - for example, the outputs generated by a sample of user queries, or those generated over multiple runs with the same specific query. This assumption can facilitate the discussion, e g., only about the individual output characterization, and can avoid certain complications of the many different ways it is possible to aggregate across different outputs. Nonetheless, the description of the exemplary embodiments described herein are applicable in numerous other ways, and without any limitations.
[0039] Figure 2 shows a block / functional diagram of an architecture according to an exemplary embodiment of the present disclosure. For example. Figure 2 shows the architecture of the genAI system 205 with the characteristic classification / scoring from characteristic classification model 270. It is possible to refer to the GenAI system 205 (e.g.. comprising prompt 210, context 220, model 230, pred 240, and next word logic 250) as a component, e.g., GAIsys 205, whose input can be the original query 210, and any otherinputs given (such as, e g., documents produced for a Retrieval Augmented Generation (RAG) process). For example, this exemplary input can be whatever can be the initial input to the underlying machine-learned model, e.g., the "Large Language Model’" (LLM) for a chat system. The output of the GAIsys component can be the generated output.
[0040] The above description is for the sake of clarity and is not in any way limiting, as there is also the input and the output of the learned Al model (e.g., the LLM), which is technically not the same as the I / O of the GAIsys. Thus, technically, the output of a system like ChatGPT can be referred to as a multi-word chat response. Further, the output of the LLM is a probability distribution over individual tokens. Such single token can be selected per inference run of the LLM. For example, the Al chat system can use this data as a component in the system’s processing, executing it multiple times, each time with a different input to the LLM, to generate the multi-word output. In such exemplary' embodiments of the present disclosure one can view, the entire chat system as a component, and it is not necessary' to consider its inner workings thereof.
[0041] Returning to Figure 2, the output of the GAIsys component 260, can then serve as the input to another inference component, characteristic classification model 270, which can decide whether and / or to what extent the output exhibits the characteristic - which can be referred to herein as, e.g., the ‘’scoring module.” For example, the scoring module can be or include an Al model itself and / or can be any component that takes the genAIsys output as its input, and produces a score or classification 280.
[0042] According to the exemplary embodiments of the present disclosure, based on the exemplary non-limiting definitions and functions discussed herein, it is possible to treat the combined Al systems as an Al inference system itself, as reflected in Figure 2 as the GenAI system characteristic classifier 200. The resulting exemplary' Al system can be or behave as an exemplary' Al inference system (as opposed to the GenAI system). Thus, it is possible to apply and / or modify the methods - for example, for producing explanations, for red-teaming, etc. Brief elaborations of each of those examples are provided as follows.Exemplary Explaining
[0043] A challenge for modem generative Al systems can be an understanding why they generated the particular output. For example, this can involve understanding why such generative Al systems exhibit a current characteristic, such as, e.g., why do they produce biased or hateful output? As with other types of Al systems, there are several different thingsthis can mean. One of them can be, e.g., what is it about the input that led it to produce biased / hateful output? This can be different from, e.g., what is it about the internals of the model that led it to produce biased or hateful output. Further differences can be, e.g., what was it about the training data used to learn the model that led it to, ultimately, produce biased or hateful outputs.
[0044] Provided herein is a discussion of, e.g., how this now can be addressed by prior solutions, and also how it is possible to adapt those solutions, according to the exemplary embodiments of the present disclosure.
[0045] Figure 3 illustrates exemplary details of a system for finding GenAl characteristic explanations in the exemplary systems, methods and computer-accessible medium according to exemplary' embodiments of the present disclosure. The set of generated explanations 390 of Figure 3 can be the output of the explanation system. The search procedure 380 is described in further detail with respect to Figure 4 herein. Figure 3 further shows that an explanation system operates on the prompt input to the GenAl system 370 — which is shown as the leftmost blue box inside the GenAl system characteristic classifier 305. As described above, alternatively this process can apply to the combination of the original user prompt and the context, or even just on the context (which might itself include documents produced by a retrieval-augmented generation (RAG) system). The input prompt is parsed into whatever units on which the explanation system has decided to focus: tokens, words, sentences, n- grams, natural language parsing objects (e.g., noun phrases, verb phrases), documents. The search procedure 380 can replace zero or more of these units, normally in an iterative manner, with the selected replacement units 360 (e.g.. underscores, mask tokens, attention masking, spaces, synonyms, other tokens of similar or greater probability, paraphrases, even nothing at all — an empty' unit). These exemplary modified prompts can run through the GenAl system characteristic classifier 305, which itself can run them through the genAI system 350, get a generated output 340, and then run that through the CCM 330, to get a score 320 for this modified prompt.
[0046] Figure 4 shows a block / functional diagram providing exemplary' details of the search procedure 380. For example, the exemplary search procedure 380 takes as input (on the right) the characteristic score 320 for the most recent (e.g., modified) prompt, as well as a prompt (not shown explicitly in the figure), e.g., it can take as input the original prompt. The exemplary procedure 380 has two or more outputs. At the end of the search process (or possibly as it goes along), it can output the set of explanations 390 it has found. During theexemplary search process, it outputs modified (‘'candidate”) prompts 360 to run through the characteristic system. Figure 4 shows an exploded view of a block / functional diagram of the exemplary search procedure 380. At the far right of Figure 4. taking the characteristic score 320 as input, the exemplary procedure 380 can generate an evaluation score 410 for the focal (modified) prompt. This modified prompt can be called “candidate i”, for the ithcandidate that it is exploring. The evaluation score 410 is not necessarily just the characteristic score. For example, it can be the difference between the current characteristic score and the (saved) score from the original prompt. It can also take into account other aspects of the modified prompt, for example, a cost function over the modification, where the cost relates to how much the modification is “messing up” the prompt. Next logic can be applied to judge whether explanation candidate i is a valid explanation 420 — using whatever function or logic is appropriate for the domain. For example, it may be desired not to have generated text with toxicity characteristic scores, e.g., above 0.4; then the evaluation score can just be the characteristic score, and the validity logic 420 can compare that score to see if it exceeds, e.g., 0.4. Take the example scenario where the score for candidate does not exceed, e.g., 0.4. Assuming that the score for the original prompt exceeded 0.4 (which can be the impetus for trying to explain the toxic behavior of the GenAI system for the original prompt), then in this example candidate explanation i is a valid explanation. (Generally, there may be other criteria a valid candidate would need to satisfy, which can also be applied in block 420.) This can be saved in the working set of explanations at, e.g., block 430.
[0047] Regardless of whether candidate i is a valid explanation, the search can check its stopping criteria at block 440 — which can be based on the number of explanations explored, the number found, the clock time passed, the compute used, whether all explanations of at least a certain size have been tested, etc. If the stopping criteria are met, then the search algorithm can move on to post-process the working set of explanations at block 450.
[0048] Post-processing the working set of explanations can be implemented by, e.g.. outputting them. It can also be more complex as well, possibly picking and choosing among them, conducting further testing of them for ranking or filtering, or even modifying them. So for example, a post-process may examine the whole set of explanations to eliminate any explanations that are simple supersets of other explanations. An explanation can be a set of input units (e.g., let the replacement unit be implicit for simplicity). Thus, for example, postprocessing 450 can eliminate explanation {drivers, can’t, but} because it sees that {can’t, but} alone is an explanation. So the former explanation is redundant. Post-processing block450 also could experiment with the explanations, possibly even running alternatives through the process above. So for example, if the working set contains only {drivers, can’t, but}, the exemplary system could attempt the cardinality-2 subsets to see if they also are explanations. If post-processing 450 finds that {can’t, but} is also an explanation, the system may decide only to show {can’t, but} as the former explanation again is redundant. That’s just an example for a situation where redundant explanations are unwanted.
[0049] At block 440, if the search algorithm decides not to stop then it can choose the next candidate explanation to consider — candidate i+1 at 460. The search procedure can generate the corresponding prompt at block 470, normally by replacing the selected input units with the replacement units, and the process iterates. Alternatively, multiple new candidate explanations (i+1, i+2, etc.) can be produced at block 460. Additional logic can also be applied to the generation of the next candidate explanation(s), for example, the system might prefer explanations that replace consecutive input units, e.g., consecutive sequences of words.Exemplary Red Teaming
[0050] Organizations implementing Al systems can be concerned that they may produce output that is objectionable or otherwise poses risk to the organization. Of a particular concern can be “unknown unknowns”, where the system does not provide indications of its own risk or provides incorrect indications of its own risk, and standard evaluations do not reveal potentially risky behavior. For generative Al systems, this can, for example, be a corporate chat system that produces unexpected biased / hateful output in a sensitive business context. One exemplary method for uncovering / exploring / assessing this risk can be to have humans explicitly tasked to “Beat the Machine” - meaning to come up with inputs that produce output with the specified (and in this case undesirable) characteristics. (See, e.g., Attenberg, J., Ipeirotis, P., & Provost, F. “Beat the machine: Challenging humans to find a predictive model's ’unknown unknowns’,” Journal of Data and Information Quality JDIQ), 6(1), 1-17, 2015) and U.S. Patent No. 9,311,599. More generally, this sort of “red teaming” can be used to find prompt inputs that produce any characteristic output of interest.
[0051] Such understanding can facilitate the utilization of the application of many prior methods and methodologies to generative Al systems. For example, RLHF can be an instance or an example thereof. The exemplary characteristic can be “thumbs up or thumbs down”,e.g., initially provided by a human (or a crowd in aggregate), and later provided by a learned Al model.
[0052] The exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure can be compared to the patent described herein. That prior patent describes the use of, e.g., a hate speech classifier as an example, which can be similar to the current example of toxic (possibly hateful) output of a generative Al model. Thus, the two red-teaming cases can be compared to see — in this specific example — what are the similarities and differences. The prior patent discusses how a document, specifically a web page, may be given to the hate speech classifier. The prior system engages humans to try to find web pages that in fact are toxic in this way (hateful), yet the hate-speech classifier is confident that they are not toxic. Thus, finding places where the hate-speech classifier can be wrong. The exemplary' systems, methods and computer- accessible medium according to exemplary embodiments of the present disclosure can be similar in that it can also engage humans to find particular behaviors of the Al system, and more specifically, in this example can place where its behavior is “wrong” with respect to the design goals (not producing toxic outputs). The difference can be that instead of challenging humans to find the documents that the system gets wrong in this way (misclassifies them as non-toxic), exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure can challenge humans to find prompts that lead the GenAI system to generate documents that are “wrong” in this way (mistakenly are toxic). The exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure can incentivize users to submit prompts to the system, prompts that can produce GenAI output (documents) that are undesirable or exhibit other characteristics of interest.
[0053] For example, some exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure may decide not to reward prompts that they already “know” produce characteristic (e.g., toxic) outputs. So the reward may not be based simply on the characteristic score, but (similarly to the explanations search evaluation) may include additional information. An example of this could be: receipt of submissions showing a certain sort of prompt produces toxic output, but the GenAI system has not yet been “fixed.” Repeated prompts of this nature may not be rewarded. That could be implemented in many ways, for example, using a “near neighbor” similarity function to penalize prompts that are too similar to some database of identified prompts.Exemplary Counterfactual explanations: Why does the system produce output with this characteristic ?
[0054] When reviewing or working with an Al system that generates output with some particular characteristic -especially a negative one- it can be interesting to determine or understand why the system just produced hateful or biased content. There may be different interpretations of this issue, and as an example, an exemplary discussion herein can focus on one exemplary (but non-limiting) issue. Specifically, it is possible to, e.g., focus on the question of “what was it about the input that led it to produce output with this characteristic?”
[0055] Once it is reformulated how the system is viewed, as shown in Figure 2, it is possible to see that the technology and discussion included in Martens and Provost 2014 and the elaboration of that technology described by Fernandez-Loria et al. 2023, can be adapted to provide a mechanism for answering such questions. For example, viewing the joint GenAI+output+CCM system (represented by the outermost box in Figure 2) that comprises the generative Al system and the characteristic classification system according to exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure, it is possible to see that this exemplary system takes as input a text document (e.g., the prompt plus any other context taken as input) and produces an output. It is possible to apply the techniques of Martens and Provost 2014 and Fernandez-Loria et al. 2023 to such exemplary systems, to provide explanations for why the prompt+context input led to the characteristic classification output.
[0056] The above explanation may not necessarily solve all problems with understanding the behavior of the system. Nonetheless, the exemplary systems, methods and computer- accessible medium according to exemplary embodiments of the present disclosure can provide an understanding how the inputs drive characteristics of the output is one important task in explaining the behavior of Al systems. For example, it is possible to aggregate explanations for individual prompts to create explanations for families (or populations) of prompts. As an example, it is possible to take a collection of prompts from real usage of the system-that exhibit a characteristic-and by aggregating their explanations explain what is it about current usage that is leading the system to exhibit the focal characteristic (e.g., bias).
[0057] It is possible to generate prompts with variations to explore the behavior of the system more thoroughly. Thus, for example, once explanations are obtained for why some user prompts led to biased output, it is possible generate variations of the prompts and examine the sensitivity of the characteristic behavior to the exact wording of the prompts.Further, similarly to the discussion herein above, it is possible to aggregate across the explanations to provide a more general understanding of the system behavior.Exemplary Adapting the prior explanation methods to prompt-based GenAI systems
[0058] The prior methods can be applied directly, as words or tokens can be removed from the prompts, and possibly replaced with alternatives (see, e.g., Fernandez-Loria et al., 2023). However, according to various exemplary embodiments of the present disclosure, it is possible to provide better ways to modify the input to produce counterfactual explanations of characteristic behavior. For example, instead of focusing on individual words for removal / replacement, it is possible to focus on phrases, n-grams (sequences of contiguous words or tokens), part-of-speech groupings (e.g., as produced by a text parser such as the Stanford parser https: / / nlp.stanford.edu / software / lex-parser.shtml), entity-based replacement, heuristic function span prediction, attention guided focal tokens using agglomerative clustering, etc
[0059] There are different ways to replace the text, e g., by deletion (resulting in a shorter prompt). Further, the text can be masked (e.g., so that the LLM does not use it in inference) and / or replaced with alternative text. One way to replace the text with an alternative text can be to use either a generative Al system or another such Al system to suggest alternative words, phrases, text sequences, or prompts. The exemplary system can generate many and various alternative prompts with similar meaning, and look at the characteristic classification output and / or differences thereof.
[0060] More generally, the process of finding explanations can be a search process. The space being searched can be the space of all combinations of input ‘"units.” Input units are the result of dividing up the input into pieces. Consider a GenAT system that takes a text prompt as input — essentially it takes a text document as input. The document can be divided up in different ways, resulting in different input units. So for example, the input units could be tokens or words or n-grams or sentences. The input units can be result of applying a NLP parser to the document, yielding (for example) aggregations of document words corresponding to noun phrases, verb phrases, and so on.
[0061] The exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure can employ search “operator(s)” that can take particular input units and replace them with replacement input units. The replacement units can be the sorts of things discussed above: e.g., nothing (deleting the input unit),underscores or some other placeholder character, masking tokens, masked attention, synonyms, Al-generated paraphrases, and so on.
[0062] In the exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure, the search space can then be the space of all sequences of applications of the search operator, in other words, most or all combinations of replacements of (some) input units by replacement units. To be clear, this can include every combination from just replacing one input unit, to all possible pairs of replacements, to replacing every unit in the input document.
[0063] The exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure can give feedback to the search algorithm by running the GenAI system with the modified input prompt, then taking the generated output (another, different document) and inputting that to the CCM module. The CCM model can then score the new document (e.g.. the generated output). This score can then be compared to other generated scores. For example, it may be useful to compare to the score for the original document, which gives information on how much the replacements affected the overall characteristic score. Or alternatively, exemplary systems, methods and computer- accessible medium according to exemplary embodiments of the present disclosure can compare the new score to the best score from a prior search round, which for example can allow the implementation of a greedy search strategy (or more sophisticated search strategies like beam search). Specifically, a greedy search in this space can start by looking at all singleton replacements for the original input. After getting feedback for each from the CCM, it can choose the one replacement that reduced the score the most (this can be called the “best’’ single replacement). Next exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure can initiate a second round of search, but this time starting with this best single replacement. So now it can try- adding all (remaining) token replacements one-by-one, trying to find the “best” second replacement. The node evaluation / feedback of exemplary- systems, methods and computer- accessible medium according to exemplary embodiments of the present disclosure can also take into account other factors. For example, perhaps different replacements are more or less desirable or more or less costly for the ultimate application. Then the search node evaluation can combine the reduction in CCM score with the benefit or cost of the replacement. For example, if a user wants to generate a document about driver quality, it might be very costly to remove the word “drivers” from the prompt. Exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure might then create a score that is the (combined) benefit per unit cost of the (joint) replacement. In fact, some input units may be deemed to be infinitely costly to remove. If a benefit per unit cost is used, search evaluation could lead those input units to never be chosen.
[0064] The search can also take into account a “goal threshold” on the CCM score or other score function — a score level that would indicate a satisfactory explanation. For example, when the greedy procedure of the exemplary systems, methods and computer- accessible medium according to exemplary embodiments of the present disclosure finds a combination of replacements that moves the score below the threshold, this combination can be saved or highlighted as a viable explanation.
[0065] Described above is only one example search strategy7of exemplar}7systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure, and the exemplary embodiments can include many search procedures. For example, a beam search algorithm might be similar, but instead of saving the single “best” combination of replacements so far, it saves some number w (the “beam width”) of “best-so- far” explanations. Alternatively, fast combinatorial solvers can be used to search for high- scoring explanations. A few (of the very many) additional alternatives for searching for exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure include: A* search, best-first search, iterative deepening search, cached iterative weakening search, OPUS, spans suggested by another generative Al system, etc.
[0066] Generally, there can be multiple valid explanations for a given input prompt — e.g., multiple replacement combinations that bring the score below the threshold. There also could be zero or one. In exemplar}7systems, methods and computer-accessible medium according to exemplar}7embodiments of the present disclosure, different particulars explanation applications can demand different numbers of explanations: e.g.. just one, all of them, as many as can be found within a certain limit on time or compute.Table 1. An overview of a run of a genAI system that produces a toxic output.
[0067] Table 1 provides a simple illustrative example and presents a typical run of a generative Al (LLM-based) system according to the exemplary embodiments of the present disclosure. The particular GenAI model is OpenAT’s GPT-2 XL (1.61 billion parameters).The bottom right of table 1 shows the input and output of the system. Presented in this way, the input can be considered the prompt and the output as the “continuation.” Thus, the input is, “You know ... I really can't say this ... but ... it is pretty’ obvious that the worst drivers in this country are” The text that follows the prompt is the continuation — this is the text thatGPT-2 XL generated: “ black. The problem is that the media is not reporting this.” Not surprisingly, a toxicity CCM classifies this output as being toxic with a score of 0.983 (out of a [0,1] range). Specifically for the classifier for the toxicity characteristic, an off-the-shelf opensource toxicity model using a RoBERTa model fine-turned on three datasets from Jigsaw can be used.
[0068] Table 2 shows an explanation result according to the exemplary' embodiments of the present disclosure.Table 2. The exemplary search uncovers the explanation '‘remove the words ‘can’t’ and ‘but’ (replacing them with underscores) and the output no longer is toxic.
[0069] In this example, the input units can be simply the words. The replacement unit can be the underscore. The explanation is then the set of words {‘can’t’, ‘but’}, which can be interpreted as: if the words ‘can't’ and ‘but’ are removed from the prompt, the GenAI system will no longer produce toxic output. This exemplary result is reflected in Table 2, where (again, under Text Generation) the black text now shows the prompt after the replacement of the explanation words by underscores, and the generated output is quite different, “the ones who are driving the most. And I think that's a pretty good indication of what' s”. The toxicity score from the classifier is 0.
[0070] The particular notion of counterfactual at work can be seen here in this explanation: the original input produced toxic output. The exemplary search using the exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure, is looking for inputs that are very similar, although are not the original inputs (counterfactual inputs) that lead to significantly different outputs with respect to the characteristic of interest (here, toxicity).
[0071] Despite its simplicity, this is a revealing example with respect to the value of characteristic counterfactual explanations. It appears that the system has been trained to not be toxic (for example, via reinforcement learning with human feedback or other means). Indeed, it looks like it has been trained to say that it does not want to answer when the answer from the raw training data might lead it to be toxic. However, many “gotcha” experiences have taught that people can engineer prompts to get around such guard-rail training. For example, guard-rail training can be different from post-generation guard rails. The former tries to train the model to avoid certain generation behavior. The latter notices that the model has generated something objectionable, and then takes an action. For example, refusing to answer or giving a stock answer. The explanation here seems to be highlighting the “why”: by prompting the machine with the speaker being reluctant “can’t” to “say” the toxic content, “but” feeling like it is necessary to do it, it bypasses the training that kept the machine from generating the toxic output (that it does not generate with the exact same prompt, except for these two words).
[0072] If there are multiple instances of a word (or other input unit), an exemplary procedure for the exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure may have to select whether to remove individual occurrences of the word (input unit) or all the occurrences of identical input units.Exemplary reasons why systems that find such explanations are needed.
[0073] The prior work by Martens & Provost 2014, as well as the large amount of subsequent work on counterfactual explanations for non-generative Al systems / machine- leamed models, discusses many reasons it can be important to be able to explain Al system I / O behavior — from justifying behavior to customers and managers, helping to comply with regulations, and debugging the machine learning process, all the way to understanding that there are fundamental problems with the training data.
[0074] There are additional important reasons for generative Al systems, including the extreme difficulty of examining the workings of the genAI model. This small model in the example in Tables 1 and 2 has 1.6 billion parameters — meaning the behavior of the model is defined by 1.6 billion numbers that define the model’s generation behavior (technically, the function that the model is computing). Moreover, the generative Al system itself is not just the LLM (contrary to what many popular discussions would lead one to believe). As illustrated in the inner box in Figure 2, the generative Al system 205 can apply the massive LLM many times — for example, one time for each output token (think for each word), and the input to the LLM is different each time, as the generated output is added to the input for each subsequent word — a process which can be complicated. The exemplary systems, methods and computer-accessible medium according to exemplary embodiments of the present disclosure can provide explanations that arch over the whole process, and are necessary because there is no visibility' as to what is going on under the hood. Furthermore, 1.6 billion parameters is a small model. State-of-the-art LLMs have over a trillion parameters.
[0075] As a result of the inability to understand and predict accurately the output of a system for a given input prompt, an entire new area of "prompt engineering’" has arisen — figuring out how to get a satisfactory output from a genAI system. In addition to their other dimensions of value, counterfactual characteristic explanations are a key tool for prompt engineering, because they focus the prompt engineer in on the particular aspects of the prompt that are responsible for particular characteristics of the generated text.Exemplary Beat-the-Machine Red Teamins: Understandins Exemplary System’s Possible Behavior
[0076] While the exemplary embodiments of the present disclosure can examine an Al system’s past behavior, they are usable to understand the Al system’s potential behavior. This can be especially the case for objectionable characteristics, like the examples of, e.g., hateful content and biased content. The exemplary architecture and information shown in Figure 2 facilitates an exemplary application of certain methods, such as, e.g., Beat the Machine (see, e.g., Attenberg et al. 2015). For example, humans can be challenged to come up with prompts that can produce GenAI system output that exhibits certain characteristics.Exemplary Extension: Generative Al as the red team!
[0077] For example, adding in the explicit characteristic classification as part of the exemplary7system can facilitate the development of exemplary7Al systems (e.g., exemplary7generative Al systems) to act as red team members, according to further exemplary7embodiments of the present disclosure, as shown in the exemplary architecture of Figure 5. For example, the AI_RT system can generate its own prompts 590, trying to "‘beat the machine”, i.e., finding prompts that lead the GenAI system to produce characteristic outputs. As an example, the AI RT system may be attempting to find prompts that can lead the system to generate biased or hateful output.Additional Discussion
[0078] Provided herein are further discussion for the exemplary embodiments of the present disclosure.
[0079] It is possible to include humans in the CCM module, at the extreme using only humans, e.g., via a system that submits generated output to one or more members of a team (or ■■crowd") of humans and returning the humans’ judgement or score. Further, one can train models to approximate humans’ responses, and then use those models in the CCM module.
[0080] It is also possible for the search procedure to consider the computational expense of actually running the system, possibly choosing new candidate prompts accordingly. It makes for an additional exemplary architecture, such as for making the process less expensive and determining approximate answers.
[0081] Exemplary and non-limiting listing of a (e.g., binary) characteristic of the generated text according to various exemplary7embodiments of the present disclosure:
[0001] Figure 6 shows a block diagram of an exemplar}' embodiment of a system according to the present disclosure which is configured or programmed to be used with the exemplary methods, architectures and processes described herein. For example, exemplary' procedures and architectures in accordance with the present disclosure described herein can be performed by a processing arrangement and / or a computing arrangement (e.g., computer hardware arrangement) 605. Such processing / computing arrangement 605 can be, for example entirely or a part of, or include, but not limited to, a computer / processor 610 that can include, for example one or more microprocessors, and use instructions stored on a computer- accessible medium (e.g.. RAM, ROM. hard drive, or other storage device).
[0002] As shown in Figure 6, for example a computer-accessible medium 615 (e.g., as described herein above, a storage device such as a hard disk, floppy disk, memory stick, CD- ROM, RAM. ROM, etc., or a collection thereof) can be provided (e.g., in communication with the processing arrangement 605). The computer-accessible medium 615 can contain executable instructions 620 thereon. In addition or alternatively, a storage arrangement 625 can be provided separately from the computer-accessible medium 615, which can provide the instructions to the processing arrangement 605 so as to configure the processing arrangement to execute certain exemplary procedures, processes, and methods, as described herein above, for example.
[0082] Further, the exemplary processing arrangement 605 can be provided with or include an input / output ports 635, which can include, for example a wired netw ork, a wireless network, the internet, an intranet, a data collection probe, a sensor, etc. As shown in Figure 6, the exemplary processing arrangement 605 can be in communication with an exemplary display arrangement 630, which, according to certain exemplary embodiments of the present disclosure, can be a touch-screen configured for inputting information to the processing arrangement in addition to outputting information from the processing arrangement, for example. Further, the exemplary display arrangement 630 and / or a storage arrangement 625 can be used to display and / or store data in a user-accessible format and / or user-readable format.
[0083] Throughout the disclosure, the following terms take at least the meanings explicitly associated herein, unless the context clearly dictates otherwise. The term "or" is intended to mean an inclusive "or.” Further, the terms “a,” "an,” and "the” are intended to mean one or more unless specified otherwise or clear from the context to be directed to a singular form.
[0084] In this description, numerous specific details have been set forth. It is to be understood, however, that implementations of the disclosed technology can be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description. References to “some examples,” “other examples,” “one example,” “an example,” “various examples,” “one embodiment,” “an embodiment,” “some embodiments,” “example embodiment,” “various embodiments,” “one implementation,” “an implementation,” “example implementation,” “various implementations,” “some implementations,” etc., indicate that the implementation(s) of the disclosed technology so described may include a particular feature, structure, or characteristic, but not every implementation necessarily includes the particular feature, structure, or characteristic.Further, repeated use of the phrases “in one example,” “in one exemplary embodiment,” or “in one implementation” does not necessarily refer to the same example, exemplary embodiment, or implementation, although it may.
[0085] As used herein, unless otherwise specified the use of the ordinal adjectives “first,” “second,” “third,” etc., to describe a common object, merely indicate that different instances of like objects are being referred to, and are not intended to imply that the objects so described must be in a given sequence, either temporally, spatially, in ranking, or in any other manner.
[0086] While certain implementations of the disclosed technology have been described in connection with what is presently considered to be the most practical and various implementations, it is to be understood that the disclosed technology is not to be limited to the disclosed implementations, but on the contrary, is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
[0087] This written description uses examples to disclose certain implementations of the disclosed technology7, including the best mode, and also to enable any person skilled in the art to practice certain implementations of the disclosed technology, including making and using any devices or systems and performing any incorporated methods. The patentable scope of certain implementations of the disclosed technology is defined in the paragraphs, and may include other examples that occur to those skilled in the art. Such other examples are intended to be w ithin the scope of the paragraphs if they have structural elements that do notdiffer from the literal language of the paragraphs, or if they include equivalent structural elements with insubstantial differences from the literal language of the paragraphs.
Claims
WHAT IS CLAIMED IS;1. A method for explaining generative Al (“genAI"’) system behavior, comprising: supplying a prompt to a genAI system; determining at least one characteristic of an output generated by the genAI system; and determining at least one unit within the prompt that caused the characteristic to be present in the output.
2. The method of claim 1, wherein the at least one unit compnses one or more of a token, a word, a sentence, an n-gram, and a natural language parsing object.
3. The method of claim 1, further comprising assigning a score to the at least one characteristic, wherein the score represents a degree to which the output exhibits the at least one characteristic.
4. The method of claim 3, wherein the score is used as an input to a search procedure, further comprising, with the search procedure, determining at least one replacement unit to generate a subsequent prompt.
5. The method of claim 4, further comprising: providing the subsequent prompt to the genAI system; and assigning a subsequent score to the at least one characteristic in a subsequent output generated by the genAI system on the subsequent prompt.
6. The method of claim 5, further comprising comparing the subsequent score with the score to determine whether the at least one replacement unit has effected the at least one characteristic.
7. The method of claim 6, further comprising iterating the determining at least one replacement unit, creating a subsequent prompt with the at least one replacement unit, supplying the subsequent prompt to the gen Al system, assigning a subsequent score, and comparing scores until a target score is reached.
8. The method of claim 5, wherein the search procedure outputs at least one explanation for the at least one characteristic.
9. The method of claim 8, wherein the at least one explanation comprises the at least one input unit.
10. A system for explaining generative Al ("genAI") system behavior, comprising: one or more computer processors configured to: supply a prompt to a genAI system; determine at least one characteristic of an output generated by the genAI system; and determine at least one unit within the prompt that caused the characteristic to be present in the output.
11. The system of claim 10, wherein the at least one unit comprises one or more of a token, a word, a sentence, an n-gram, and a natural language parsing object.
12. The system of claim 10, further comprising assigning a score to the at least one characteristic, wherein the score represents a degree to which the output exhibits the at least one characteristic.
13. The system of claim 12, wherein the score is used as an input to a search procedure, further comprising, with the search procedure, determining at least one replacement unit to generate a subsequent prompt.
14. The system of claim 13, further comprising: providing the subsequent prompt to the genAI system; and assigning a subsequent score to the at least one characteristic in a subsequent output generated by the genAI system on the subsequent prompt.
15. The system of claim 14, further comprising comparing the subsequent score with the score to determine whether the at least one replacement unit has effected the at least one characteristic.
16. The system of claim 15, further comprising iterating the determining at least one replacement unit, creating a subsequent prompt with the at least one replacement unit,supplying the subsequent prompt to the gen Al system, assigning a subsequent score, and comparing scores until a target score is reached.
17. The system of claim 14, wherein the search procedure outputs at least one explanation for the at least one characteristic.
18. The system of claim 17, wherein the at least one explanation comprises the at least one input unit.
19. A non-transitory computer accessible medium which includes software thereon for explaining generative Al (“genAI”) system behavior, wherein, when at least one computer processor executes the software, the computer processor is configured to perform the procedures, comprising: supplying a prompt to a genAI system; determining at least one characteristic of an output generated by the genAI system; and determining at least one unit within the prompt that caused the characteristic to be present in the output.
20. The non-transitory computer accessible medium of claim 19, wherein the at least one unit comprises one or more of a token, a word, a sentence, an n-gram, and a natural language parsing object.
21. The non-transitory computer accessible medium of claim 19, further comprising assigning a score to the at least one characteristic, wherein the score represents a degree to which the output exhibits the at least one characteristic.
22. The non-transitory computer accessible medium of claim 21, wherein the score is used as an input to a search procedure, further comprising, with the search procedure, determining at least one replacement unit to generate a subsequent prompt.
23. The non-transitory computer accessible medium of claim 22. further comprising: providing the subsequent prompt to the genAI system; andassigning a subsequent score to the at least one characteristic in a subsequent output generated by the genAI system on the subsequent prompt.
24. The non-transitory computer accessible medium of claim 23. further comprising comparing the subsequent score with the score to determine whether the at least one replacement unit has effected the at least one characteristic.
25. The non-transitory computer accessible medium of claim 24, further comprising iterating the determining at least one replacement unit, creating a subsequent prompt with the at least one replacement unit, supplying the subsequent prompt to the gen Al system, assigning a subsequent score, and comparing scores until a target score is reached.
26. The non-transitory computer accessible medium of claim 23, wherein the search procedure outputs at least one explanation for the at least one characteristic.
27. The non-transitory computer accessible medium of claim 26, wherein the at least one explanation comprises the at least one input unit.
28. The method of claim 9, wherein the genAI system is reprogrammed based on the at least one explanation.
29. A method for explaining generative Al (“genAI”) system behavior, comprising: supplying a prompt to a genAI system; determining at least one characteristic of an output generated by the genAI system; and determining at least one unit within the prompt whereby replacing the at least one unit leads the at least one characteristic not to be present in the output.
30. A system for explaining generative Al (“genAI”) system behavior, comprising: one or more computer processors configured to: supply a prompt to a genAI system; determine at least one characteristic of an output generated by the genAI system; anddetermine at least one unit within the prompt whereby replacing the at least one unit leads the at least one characteristic not to be present in the output.
31. A non-transitory computer accessible medium which includes software thereon for explaining generative Al (“genAI”) system behavior, wherein, when at least one computer processor executes the software, the computer processor is configured to perform the procedures, comprising: supplying a prompt to a genAI system; determining at least one characteristic of an output generated by the genAI system; and determining at least one unit within the prompt whereby replacing the at least one unit leads the at least one characteristic not to be present in the output.
Citation Information
Patent Citations
Automatically identifying and minimizing potentially indirect meanings in electronic communications
US20210011976A1
Determining an explanation of a classification
US20210326661A1
System and method for machine learning architecture for out-of-distribution data detection
US20220245422A1