Method for identifying and remediating gaps in artificial intelligence use cases using one or more artificial intelligence models, computer system and non-transitory computer-readable storage medium

TWI931837BActive Publication Date: 2026-07-11CITIBANK N A
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TW113136158
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-09-18
Filing Date
2024-09-24
Publication Date
2026-07-11
Estimated Expiration
2044-09-23

Smart Images

  • Figure IMG-2_DRAW_113136158-A0304-14-0001-1
    Figure IMG-2_DRAW_113136158-A0304-14-0001-1
  • Figure IMG-2_DRAW_113136158-A0304-14-0002-2
    Figure IMG-2_DRAW_113136158-A0304-14-0002-2
  • Figure IMG-2_DRAW_113136158-A0304-14-0003-3
    Figure IMG-2_DRAW_113136158-A0304-14-0003-3
Patent Text Reader

Abstract

The system and method disclosed herein receive literal characters defining the operational boundaries of expected model use cases and operational data. These expected model use cases share common attributes, which are used by a first AI model to construct observed model use cases from the operational data. Each observed model use case includes features such as a text-based description, expected inputs and outputs, several AI models that generate the expected outputs from the inputs, and / or data supporting those AI models. For each observed model use case, a second AI model maps the literal characters and features to one of several risk categories selected based on a risk level associated with those features. The system identifies criteria for the observed model use case within the literal characters and generates discrepancies by comparing those criteria with the features of the observed model use case.
Need to check novelty before this filing date? Find Prior Art

Description

Prior Technology

[0001] Artificial intelligence (AI) models typically operate based on extensive and large-scale training models. These models contain multiple inputs and instructions on how to process each input. When a model receives a new input, it produces an output based on patterns determined from the data it has been trained on. A large language model (LLM) is renowned for its ability to achieve general language generation and other natural language processing tasks, such as classification. LLMs can be used to generate text by taking an input text and repeatedly predicting the next symbol or word—a form of generative AI (e.g., GenAI, GAI). LLMs acquire this ability by learning statistical relationships from text documents during computationally intensive self-supervised and semi-supervised training procedures. The use and applicability of generative AI models (such as LLMs) are increasing over time.

[0002] Generally speaking, organizations are required to comply with compliance requirements set by governments and various regulatory bodies. Different types of organizations must comply with various forms of regulation from different regulatory bodies. Increased compliance requirements for an organization create a more challenging operating environment. Regulators are taking stronger action against non-compliance by imposing hefty fines and causing potential reputational damage. However, compliance requirements become increasingly challenging when regulations cover broad definitions and multiple topics. Simple Explanation of the Diagram

[0003] Figure 1 illustrates an interpretive environment for evaluating language model prompts and outputs for model selection and validation according to some implementations of the present technology.

[0004] Figure 2 shows a block diagram illustrating some components of at least some of the computer systems and other devices typically incorporated into the disclosed system according to some embodiments of the present technology.

[0005] Figure 3 is a system diagram illustrating an example of a computing environment in which the system disclosed in some embodiments of the present technology operates.

[0006] Figure 4 shows a diagram of one of the artificial intelligence (AI) models according to some implementation schemes of this technology.

[0007] Figure 5 is an illustrative diagram illustrating an example environment of a platform for automatically managing compliance guidelines according to some implementation schemes of this technology.

[0008] Figure 6 is an illustrative diagram illustrating an instance environment of a platform that generates mapped gaps based on guidelines and gaps in the use control of some implementations of this technology.

[0009] Figure 7 is a flowchart illustrating one of the procedures for mapping identified gaps in control to one of the operating standards according to some embodiments of the present technology.

[0010] Figure 8 is an illustrative diagram of an example environment of a platform for self-guided identification of operable items according to some implementation schemes of this technology.

[0011] Figure 9 is a block diagram illustrating an instance environment in which guidelines for using input to a verification engine to determine AI compliance are employed according to some implementations of this technology.

[0012] Figure 10 is a block diagram illustrating an instance environment used to generate verification actions to determine the compliance of an AI model according to some implementations of the present technology.

[0013] Figure 11 is a block diagram illustrating an example environment for automatically performing correction actions on an AI model according to some implementation schemes of the present technology.

[0014] Figure 12 is a block diagram illustrating one instance environment of an implementation of the present technology for using a generative AI model to identify and correct compliance gaps in AI use cases.

[0015] Figure 13 is a block diagram illustrating one instance environment of a use case for continuous monitoring of compliance in AI use cases using a generative AI model, according to some implementations of this technology.

[0016] Figure 14 is a flowchart illustrating one of the procedures for identifying and correcting compliance gaps in AI use cases using a generative AI model according to some implementations of this technology.

[0017] By studying the embodiments in conjunction with the drawings, those skilled in the art should gain a clearer understanding of the techniques described herein. The embodiments illustrating the nature of the invention are illustrated by way of example, and the same element symbols may indicate similar elements. Although the drawings depict various embodiments for illustrative purposes, those skilled in the art should recognize that alternative embodiments can be adopted without departing from the principles of the art. Therefore, although specific embodiments are shown in the drawings, the art is adaptable to various modifications. Implementation

[0018] (Several) cross-references to related applications This application is a continuation-into-file of U.S. Patent Application No. 18 / 889,371, filed September 18, 2024, entitled “IDENTIFYING AND REMEDIATING GAPS IN ARTIFICIAL INTELLIGENCE USE CASES USING A GENERATIVE ARTIFICIAL INTELLIGENCE MODEL”, which is a partial continuation-into-file of U.S. Patent Application No. 18 / 653,858, filed May 2, 2024, entitled “VALIDATING VECTOR CONSTRAINTS OF OUTPUTS GENERATED BY MACHINE LEARNING MODELS”, which is a partial continuation-into-file of U.S. Patent Application No. 18 / 637,362, filed April 16, 2024, entitled “DYNAMICALLY VALIDATING AI APPLICATIONS FOR COMPLIANCE”.This application is further a continuation-into-part of U.S. Patent Application No. 18 / 782,019, filed July 23, 2024, entitled "Identifying and Analyzing Actions from Vector Representations of Alphamium Characters Using a Large Language Model," which is a continuation-into-part of U.S. Patent Application No. 18 / 771,876, filed July 12, 2024, entitled "Mapping Identified Gaps in Controls to Operative Standards Using a Generic Artificial Intelligence Model," which is a continuation-into-part of U.S. Patent Application No. 18 / 771,876, filed July 12, 2024, entitled "Dynamic Input-Senate Value of Machine Learning Model Outlets and Methods and Systems of the This is a continuation-in-part of U.S. Patent Application No. 18 / 661,532, filed May 10, 2024, entitled "DYNAMIC, RESOURCE-SENSITIVE MODEL SELECTION AND OUTPUT GENERATION AND METHODS AND SYSTEMS OF THE SAME," and a continuation-in-part of U.S. Patent Application No. 18 / 661,519, filed May 10, 2024, entitled "DYNAMIC EVALUATION OF LANGUAGE MODEL PROMPTS FOR MODEL SELECTION AND OUTPUT VALIDATION AND METHODS AND SYSTEMS OF THE SAME," filed April 11, 2024. The entire contents of the aforementioned applications are incorporated herein by reference.

[0019] Pre-existing LLM and other generative machine learning models promise to be used in a variety of natural language processing and generation applications. Beyond generating human-readable spoken output, pre-existing systems can also leverage LLM to generate technical content, including user-hinted software code, architecture, or code patching, in scenarios such as data analysis or software development pipelines. Based on the specific model architecture and training data used to generate or tune the LLM, these models can exhibit different performance characteristics, specializations, performance behaviors, and attributes.

[0020] However, users or services of pre-existing software development systems (e.g., data pipelines for data processing and model or application development) lack an intuitive, consistent, or reliable method for selecting a specific LLM model and / or designing associated prompts to solve a given problem (e.g., generating desired code associated with a specific software application). Therefore, pre-existing systems risk selecting suboptimal (e.g., relatively inefficient and / or insecure) generative machine learning models. Furthermore, pre-existing software development systems do not control access to various system resources or models. Moreover, pre-existing development pipelines do not validate LLM outputs for security vulnerabilities in a context-dependent and flexible manner. Code generated by an LLM may contain errors or programming errors that can cause system instability (e.g., by loading incorrect dependencies). Some generated outputs may be misleading or unreliable (e.g., attributed to model illusions or outdated training data). Alternatively, some generated data (e.g., associated with natural language text) is not associated with security risks of the same severity. Therefore, pre-existing software development pipelines may require the manual application of rules or policies for output verification based on the precise nature of the output, thereby leading to inefficiencies in data processing and application development.

[0021] The data generation platform disclosed herein implements dynamic evaluation of machine learning prompts for model selection and verification of the resulting output to improve the security, reliability, and modularity of data pipelines (e.g., software development systems). The data generation platform receives a prompt from a user (e.g., a human-readable request related to software development, such as code generation) and determines whether the user is authenticated based on an associated authentication token (e.g., provided concurrently with the prompt). Based on the selected model, the data generation platform determines a set of performance metrics (and / or corresponding values) associated with the requested prompt processed by the selected model. In doing so, the data generation platform evaluates the suitability of the selected model (e.g., LLM) for generating an output based on the received input or prompt. The data generation platform can verify and / or modify the user's prompt against a prompt verification model. Based on the results of the prompt verification model, the data generation platform can modify the prompt to meet any associated verification criteria (e.g., by editing sensitive data or other details), thereby mitigating potential security vulnerabilities, inaccuracies, or anti-manipulation effects associated with the user's prompt.

[0022] (Several) Selected models face compliance challenges regarding AI models and a range of vector constraints (e.g., guidelines, regulations, standards) related to ethical or regulatory considerations, such as protection against bias, harmful language, and intellectual property (IP) rights. For example, vector constraints may include requirements that AI applications generate unbiased, harmless, and IP-free output to uphold ethical standards and protect users. Traditional approaches to regulatory compliance typically involve the manual interpretation of regulatory text, followed by special efforts to align the AI ​​system with compliance requirements. However, manual procedures are subjective, lack scalability, and are prone to errors, making this approach increasingly unsustainable in the face of evolving guidelines and the rapid proliferation of AI applications.

[0023] Therefore, the inventors have further developed a system to provide a systematic and automated method for assessing and ensuring compliance with guidelines (e.g., preventing bias, harmful language, and IP infringement). The revealed technology addresses the complexity of compliance for AI applications. In some implementations, the system uses a meta-model consisting of one or more models to analyze different patterns of AI-generated content. For example, one model may be trained to identify certain patterns within the content (e.g., patterns indicating bias) by assessing demographic attributes and characteristics present in the content. By quantifying bias within the training dataset, the system can effectively scan content for disproportionate correlations with demographic attributes and provide insights into potential biases that could affect the fairness and impartiality of the AI ​​application. In some implementations, the system generates actionable verification actions (e.g., test cases) that serve as input to the AI ​​model for assessing AI application compliance. The system evaluates the AI ​​application against this set of verification actions and generates one or more compliance indicators and / or a set of actions based on a comparison between expected and actual results and interpretations. In some implementations, the system may incorporate a correction module that automatically performs correction procedures to remove non-compliant content from the AI ​​model. The correction module adjusts the AI ​​model's parameters and / or updates training data based on the detection model's findings to ensure rapid resolution and mitigation of non-compliant content.

[0024] Unlike manual processes that rely on human interpretation of guidelines and compliance assessments, this system detects subtle nuances that traditional content moderation methods often miss. It parses and analyzes the textual data within AI model responses, identifying subtle nuances in expression, meaning, and cultural references that may indicate bias or harmful content. Furthermore, through standardized validation criteria, the system establishes clear and objective standards for evaluating the content of an AI application, thereby minimizing the impact of personal bias or interpretation. The system processes large volumes of content quickly and consistently, ensuring all content is evaluated against the same set of standards and guidelines, thus reducing the likelihood of discrepancies or inconsistencies in enforcement decisions.

[0025] In cases where non-compliance is detected, the conventional method of mapping gaps (e.g., problems) in controls (e.g., a set of expected actions) to operational standards (e.g., obligations, guidelines, measures, principles, conditions) heavily relies on manually mapping each gap to one or more operational standards. A gap represents a situation where an expected control is missing or cannot function properly, such as the failure to establish a specific framework within an organization. Operational standards contain controls that can be based on publications (e.g., regulations, organizational guidelines, best practice guidelines, etc.). Using manual procedures heavily depends on individual knowledge and therefore carries a significant risk of potential bias. This subjectivity can lead to inconsistent mappings because different individuals can interpret and apply operational standards such as regulatory requirements in various ways. Furthermore, the large number of identified gaps complicates traditional compliance work. Manually managing this large number of gaps is not only labor-intensive but also prone to oversight. Another significant drawback of conventional methods is the static nature of the mapping process. Conventional methods often fail to consider the dynamic and evolving nature of regulatory requirements and organizational controls.

[0026] Therefore, the inventors have further developed a system for mapping gaps in control to corresponding operational standards using generative AI (e.g., GAI, GenAI, generative artificial intelligence) models (such as a large language model (LLM) in the aforementioned data generation platform). The system determines a set of vector representations of alphanumeric characters represented by one or more operational standards, which contain a first set of actions that comply with constraints in the set of vector representations. The system receives an output generation request via a user interface, the output generation request containing a set of gaps associated with a case that fails to meet the operational standards of the set of vector representations. Using the received input, the system constructs a set of prompts for each gap, wherein the set of prompts for a specific gap contains a set of attributes defining the case and the first set of actions of the operational standard. Each prompt can compare the corresponding gap with the operational standard or the first set of actions of the set of vector representations. For each gap, the system maps the gap to one or more operational standards by supplying prompts to the LLM and, in response, receiving a set of gap-specific operational standards from the LLM containing operational standards associated with the specific gap. Compared to familiar methods, the system reduces reliance on individual knowledge, thus minimizing personal bias and resulting in more consistent mappings across different individuals and teams. Furthermore, the system effectively addresses the large gaps faced by organizations, significantly reducing the labor-intensive nature of manual review.

[0027] In another instance, the conventional approach to identifying actionable items in self-guidelines presents several challenges. Typically, conventional approaches involve human reviewers or automated systems that process guidelines in a linear fashion. Conventional linear approaches often result in the identification of a large number of actionable items. Furthermore, conventional approaches lack the ability to dynamically adapt to changes in guidelines over time. When new guidelines are introduced or existing guidelines are updated, conventional systems often simply add new actionable items without reassessing the overall group of actionable items to ensure that the new items are not redundant or contradictory to previously defined actionable items. Conventional approaches further fail to account for subtle shifts in the interpretation of changes in custom or regulatory language, potentially leaving outdated or irrelevant requirements on the list. Therefore, organizations may end up with an overblown and confusing set of actionable items that does not accurately reflect the current context of the guidelines (e.g., the current regulatory environment).

[0028] Therefore, the inventors have further developed a system for self-guided identification of actionable items using generative AI models (such as one of the LLMs in the aforementioned data generation platforms). The system receives an output generation request from a user interface, containing one input for generating an output using an LLM. Guidelines are partitioned into multiple text subsets based on predetermined criteria (such as the length or complexity of each text subset). Using partitioned guidelines, the system constructs a set of prompts for each text subset. Each text subset can be mapped to one or more actions in a first set of actions. Subsequent actions in this second set can be generated based on previous actions. The system generates a third set of actions by summarizing the corresponding actions in the second set for each text subset. Unlike familiar linear procedures that result in a large number of redundant actionable items, the system, through heuristic analysis of guidelines, can identify common actionable items without parsing the guide document word by word. The revealed system reduces the number of identified actionable items to only relevant actionable items. Furthermore, the system's dynamic and contextual awareness allows it to respond to changes in guidance over time by reassessing and mapping changes as actionable items change.

[0029] Furthermore, conventional compliance methods are insufficient given the increasingly complex and widespread regulation of AI applications. Conventional compliance methods typically involve manual procedures, static documentation, and periodic audits. While these methods are feasible for traditional systems with well-defined boundaries and limited complexity, regulations are becoming increasingly complex. For example, conventional compliance methods are particularly challenging due to the broad definitions of "AI systems" or "models" in regulations such as the EU AI Act (which encompass a wide range of automated systems that, due to their size and complexity, cannot be manually evaluated). For instance, the EU AI Act's broad definition of "AI systems" includes machine learning models, expert systems, and even simpler rule-based systems, meaning that many systems previously unconsidered for AI are now subject to regulatory audits. Conversely, for example, California Senate Bill 1047 (California SB-1047) defines "covered models" on or after January 1, 2027, as any artificial intelligence model trained using a certain amount of computing power determined by the Government Operations Authority, where the cost exceeds $100 million when calculated using the average market price of cloud computing at the start of training. This broad definition means that many AI systems—including those previously not considered subject to regulatory scrutiny, such as basic machine learning models or simple decision trees used in commercial operations—can now fall under this regulation. Furthermore, definitions of similar terms vary across different regulations (e.g., the EU AI Act versus California SB-1047). For financial institutions, reclassification can result in increased compliance burdens and the need to update existing documentation and procedures. Even long-term systems that organizations may not have previously perceived as requiring monitoring—such as rule-based engines and pattern recognition tools—can now fall under the new regulations. Additionally, some regulations require the AI ​​used to ensure compliance to comply itself, creating a periodic challenge of requiring compliance tools to also be subject to regulatory oversight.

[0030] Furthermore, the definition of a particular point in time can vary across regulations. For example, financial institutions face significant challenges because they need to maintain extensive documentation for their models within a Model Risk Management (MRM) framework, including procedures for model development, validation, performance monitoring, and control. The EU AI Act potentially introduces additional documentation requirements, such as transparency reports, risk assessments, and compliance checks specifically for AI systems. Therefore, for compliance purposes, the organizational tools and data of any system falling under the definition of "AI system" need to be linked to the regulatory definition rather than a standard technical understanding.

[0031] Furthermore, under specific regulations, AI systems require continuous evaluation throughout their entire lifecycle, and their decisions must be understandable and interpretable by humans, necessitating the continuous storage of decisions and their underlying principles. For example, the EU AI Act categorizes AI systems into four risk levels: unacceptable risk, high risk, limited risk, and minimum risk. High-risk systems (including those used in sensitive areas such as healthcare, transportation, or law enforcement) require strict compliance with regulations, including transparency, documentation, and human oversight. For high-risk AI systems, companies must implement a risk management system that continuously evaluates the AI ​​system throughout its lifecycle, maintain detailed records of data sources and processing methods, provide clear information on the operation and limitations of the AI ​​system, establish agreements for human intervention in key decision-making processes, and implement robust cybersecurity measures. In another example under California SB-1047, organizations using "covered models" as defined in that regulation are required to implement comprehensive security and safety protocols that include the ability to completely shut down the model, and further require the retention of unedited copies of security protocols and audit reports for five years while the model is in use, thus requiring continuous evaluation. Therefore, conventional compliance methods that typically rely on manual, periodic procedures and static documentation are insufficient for the continuous monitoring requirements of AI systems in specific risk categories.

[0032] Furthermore, AI systems face new challenges in meeting compliance requirements across multiple interconnected topics, such as risk management, data control, transparency, human oversight, and cybersecurity measures. Each assessed area requires specialized knowledge, making documentation complex and resource-intensive. Documenting compliance across these different and interconnected areas involves coordinating the work of multiple teams, maintaining up-to-date records, and ensuring that all aspects of the AI ​​system comply with regulatory requirements—a challenging and ongoing task. For example, effective risk management depends on accurate data control to ensure that data used for risk assessment is reliable and compliant with privacy regulations, while robust cybersecurity measures are necessary to protect data from leakage, thereby supporting both risk management and data control efforts. Due to the interconnectedness of these areas, duplication of work wastes valuable resources such as CPU usage, storage capacity, and manpower, as multiple teams redundantly process the same data, run similar compliance checks, and maintain overlapping documentation, leading to inefficiency and increased operating costs.

[0033] Furthermore, to monitor the compliance of AI systems, the decisions made by these systems must be understood and interpreted by humans. This implicitly means that decisions and their underlying principles should be continuously stored and managed. This presents several challenges, especially in the context of complex AI models that can operate as "black boxes," where the decision-making process is not easily decipherable. From the user's perspective, the AI ​​model acts as a "black box," where inputs are fed into the system and output predictions are generated without revealing the underlying logic. The opaque nature of AI systems makes it difficult to track how specific decisions are made (especially in the case of complex documents), thereby complicating the identification of compliance gaps.

[0034] Therefore, the inventors have further developed a system for using generative AI models (such as LLMs in the aforementioned data generation platform) to identify and address gaps in AI use cases. The system (1) uses an existing inventory and / or (2) uses regulatory definitions to build an inventory of model use cases (i.e., observed model use cases) for the organization. A model use case may be a specific application or case in which AI technology is used (or can be used). A model use case may contain a set of features detailing the background, objectives, and requirements of the AI ​​system (e.g., a description of the problem being solved, background of the model use case, data inputs and outputs, the AI ​​model and algorithm used, and expected benefits or improvements). The inventory may include any applicable tools used by the organization, including third-party tools (e.g., from downstream suppliers), and may be identified through a data capture augmentation generation (RAG) search. The system may identify areas requiring additional information and formulate a request for that additional information.

[0035] For each identified AI use case in the inventory, the system uses a rule-based system and / or one or more AI models to examine the AI ​​use cases based on applicable regulations (e.g., the EU AI Act, California SB-1047). Based on regulations, the AI ​​use cases are categorized into a risk category. This categorization can be performed by analyzing input data (i.e., AI use cases and applicable regulations) and mapping the AI ​​use cases to one of the predefined risk categories using a rule-based system and / or neural networks. Using identified compliance requirements, the system identifies a set of compliance gaps for each AI use case. Gaps represent situations where one of the expected compliance requirements of a regulation is missing or not functioning properly. To address the identified gaps, an action plan outlining the tasks to correct them is established. For example, the system can compile the required documentation to correct the gaps and generate a time-stamped compliance report that meets one of the applicable regulations for the AI ​​use case.

[0036] Internally, the system can automatically reassess and identify new gaps, generating an alert based on these gaps (e.g., triggering an anomaly for potential "prohibited" or "high-risk" AI use cases). In some implementations, the system transforms existing documents (e.g., previously compiled documents) into new documents that meet new regulatory requirements. The system can combine textual analysis of existing documents with relevant regulations to categorize risk types (disregarding pre-assigned risk ratings provided by other regulations). Subsequently, the system uses the existing documents to generate new documents that comply with the updated regulatory requirements.

[0037] Unlike conventional methods that rely on manual procedures, static documentation, and periodic audits, the revealed system and methodology generate an inventory that ensures all relevant AI use cases are considered through RAG (Regulatory Aggregate Algorithm) searches of operational data. This reduces compliance burdens and ensures that even previously overlooked AI systems are included in the analysis. Furthermore, unlike conventional methods that struggle to address differing definitions of "AI system" across various regulations, the revealed system uses rule-based systems and / or AI models to examine each identified AI use case against applicable regulations. By categorizing AI use cases into risk classes based on specific regulations, the system ensures that each use case is evaluated according to the specific requirements of the relevant regulatory framework. This dynamic categorization and gap-finding process allows for immediate adjustments and continuous compliance, thereby reducing the risk of regulatory loopholes and associated penalties. Additionally, the system implements a risk management system that continuously evaluates AI systems throughout their lifecycle. By generating time-stamped compliance reports and compiling necessary documentation, the system provides a dynamic and automated solution for continuous evaluation, thereby reducing the risk of non-compliance. The system addresses the challenge of meeting compliance requirements across multiple interconnected topics using a single system of (several) generative AI models across various subjects, reducing the need for repetitive work. The system's ability to aggregate documents and generate compliance reports ensures that all interconnected areas are addressed coherently, allowing organizations to allocate resources more effectively and reduce operational overhead.

[0038] Furthermore, unlike conventional methods that struggle to address the opaque nature of complex AI models, this system integrates compliance requirements and operational data into a set of compliance gaps and / or automatically executed compliance actions. By mapping compliance requirements to the AI ​​system's operational data (such as model inputs and outputs, decision-making procedures, and performance metrics), the system can pinpoint specific areas where compliance gaps exist. For example, if a regulation requires detailed documentation of data sources and processing methods, the system can automatically examine existing documents for this requirement and mark any discrepancies as compliance gaps.

[0039] Compared to traditional methods used in operational models, the methods disclosed in this paper result in a reduction in greenhouse gas emissions. Globally, approximately 40 billion tons of CO2 are emitted annually. Digital technologies account for about 4% of this figure. Furthermore, known user device and application settings can sometimes exacerbate climate change. For example, U.S. power plants consume an average of approximately 500 grams of carbon dioxide per kilowatt-hour of electricity generated. The implementation schemes disclosed in this paper for conserving hardware, software, and network resources can mitigate climate change by reducing and / or preventing additional greenhouse gas emissions into the atmosphere. For example, using generative AI models, as described in this paper, to identify and correct gaps in AI use cases can reduce electricity consumption compared to traditional methods. Specifically, automated computer-based compliance and monitoring tasks reduce the need for extensive manual intervention and redundant procedures (which can consume additional computing power and energy). Continuous compliance monitoring reduces the need for periodic, resource-intensive audits and assessments, which traditionally require significant computing power to process large amounts of data. A surge in electricity consumption can lead to higher greenhouse gas emissions. Power plants may need to rapidly increase production to meet sudden increases in demand and may rely on less efficient and more polluting energy sources. For example, instead of conducting quarterly or annual compliance reviews involving extensive data processing and analysis, the revealed system can perform compliance monitoring on a continuous basis to handle smaller incremental data updates, resulting in more consistent energy consumption.

[0040] Furthermore, in the United States, data centers account for approximately 2% of the country's electricity consumption, and globally, they account for approximately 200 megawatt-hours (TWh). Transmitting 1 GB of data generates approximately 3 kg of CO2. Therefore, each GB of downloaded data results in approximately 3 kg of CO2 or other greenhouse gas emissions. Storing 100 GB of data in the cloud annually generates approximately 0.2 tons of CO2 or other greenhouse gas emissions. The continuous monitoring and dynamic compliance management described in this paper enables the system to detect and resolve compliance issues as they arise, rather than waiting for the next scheduled review. This proactive approach not only ensures more effective compliance maintenance but also reduces the likelihood of significant non-compliance issues requiring extensive corrective action. By addressing potential problems early, the system reduces the need for resource-intensive remediation work, further conserves energy, and eliminates the need for wasteful CO2 emissions. Furthermore, maintaining AI models' compliance with environmental regulations ensures that the AI ​​system itself adheres to standards that reduce its environmental impact. By continuously monitoring and correcting energy efficiency gaps in AI operations, the system helps reduce the carbon footprint associated with data processing and storage. Compliance with environmental regulations can include requirements for energy efficiency, waste reduction, and sustainable resource use. By meeting these requirements, broader environmental goals are achieved, such as reducing greenhouse gas emissions and conserving natural resources. Therefore, compared to conventional internet technologies, the proposed implementation scheme mitigates climate change and its impacts by reducing the amount of data stored and downloaded.

[0041] Significant technical uncertainties arise from the use of existing methods to establish a system for dynamically identifying and addressing compliance gaps in AI use cases. Building such a system requires addressing several unresolved issues inherent in existing methods for AI compliance management, such as how to interpret regulations and apply them to AI use cases. AI regulations vary significantly across different jurisdictions and industries, making it challenging to establish a system that can accurately interpret and apply such complex and variable regulatory standards. Similarly, existing methods for AI compliance management do not provide a means for continuous learning and adaptation to new regulatory changes and updates.

[0042] Traditional methods rely on periodic reviews and audits, which are insufficient for the dynamic nature of AI systems. Given regulations requiring continuous compliance (such as the EU AI Act and California Senate Bill 1047), traditional methods are inadequate because they fail to continuously track compliance status and identify gaps as they arise, rather than during scheduled audits. For example, a traditional system might manually review logs, data processing workflows, and model outputs. The process might involve extracting data from various sources (such as databases, log files, and application programming interfaces (APIs)) and then manually cross-referencing the data against regulatory requirements. This manual process is not only time-consuming and prone to human error, but it also fails to capture compliance issues that can arise between audits. In contrast, the revealed system, by integrating in real-time with AI models and data pipelines, uses APIs and an event-driven architecture to capture data as it is processed, determining how to dynamically meet regulatory requirements. For example, the system can automatically scan and interpret regulatory text, mapping requirements to specific data processing activities and model behaviors. In addition, the system can identify anomalies in data access or processing that indicate a compliance vulnerability, and can automatically execute computer-executable tasks to fix the compliance vulnerability.

[0043] To overcome technological uncertainties, the inventors systematically evaluated various design alternatives. For example, they tested different machine learning algorithms to determine the most effective one for dynamic compliance monitoring given variable data in an AI use case. They also experimented with a rule-based approach where predefined rules are manually encoded to map AI use cases to specific regulatory guidelines. This involves building a broad database of rules corresponding to different regulations and manually updating these rules as new regulations emerge. Additionally, they explored a template-based approach where compliance templates are created for different types of AI applications and used to guide compliance monitoring procedures.

[0044] However, rule-based approaches have proven inflexible and difficult to maintain. As regulations evolve, manually updating rules becomes increasingly cumbersome and error-prone, leading to delays in compliance updates and potential gaps in regulatory coverage. Similarly, due to the variability of regulations, template-based approaches lack the granularity needed to address the specific nuances of different AI use cases. Templates are too general, resulting in either overly broad compliance checks that are likely to produce many false positives or overly narrow checks that may overlook compliance issues.

[0045] Therefore, the inventors experimented with different methods for dynamically identifying and addressing compliance gaps. For example, they tested various machine learning models to analyze regulatory text (e.g., using NLP) and automatically extract relevant compliance criteria to build a system that can adapt to new regulations in real time. The system can use, for example, classification models (such as Support Vector Machines (SVM) and Random Forests) to map the extracted criteria to specific AI use cases to classify AI use cases based on their risk profile and regulatory requirements. The system can use the criteria to continuously monitor AI applications to identify deviations from expected behavior that indicate compliance issues.

[0046] While the current description provides examples related to LLM, those familiar with this technique will understand that the techniques disclosed are applicable to other forms of machine learning or algorithms, including unsupervised, semi-supervised, supervised, and reinforcement learning techniques. For example, the data generation platform disclosed can evaluate model outputs from Support Vector Machines (SVM), k-Nearest Neighbors (KNN), decision-making, linear regression, random forests, naïve Bayes, or logistic regression algorithms and / or other suitable computational models.

[0047] In the following description, for illustrative purposes, numerous specific details are set forth to provide a thorough understanding of one implementation of the present technology. However, those skilled in the art will understand that implementations of the present technology can be practiced without such specific details.

[0048] The phrases "in some embodiments," "in several embodiments," "according to some embodiments," "in the illustrated embodiments," "in other embodiments," and the like generally mean that the specific feature, structure, or characteristic following the phrase is included in at least one embodiment of the present technology and may be included in more than one embodiment. Furthermore, these phrases do not necessarily refer to the same embodiment or different embodiments. [Overview of Data Generation Platforms] [] [ ]

[0049] Figure 1 illustrates an interpretive environment 100 for evaluating machine learning model inputs (e.g., language model hints) and outputs for model selection and validation, according to some embodiments of the present technology. For example, environment 100 includes a data generation platform 102 capable of communicating (e.g., transmitting or receiving data to or from a data node 104 and / or a third-party database 108a to 108n) via a network 150. The data generation platform 102 may comprise software, hardware, or a combination of both, and may reside on a physical server or a virtual server running on a physical computer system (e.g., as described in Figure 3). For example, the data generation platform 102 may be distributed across various nodes, devices, or virtual machines (e.g., in a distributed cloud server). In some embodiments, the data generation platform 102 may be configured on a user device (e.g., a laptop, smartphone, desktop computer, tablet computer, or other user-suitable device). In addition, the data generation platform 102 may reside on a server or node and / or may directly or indirectly interface with third-party databases 108a to 108n.

[0050] Data node 104 can store various types of data, including one or more machine learning models, prompt verification models, associated training data, user data, performance metrics and corresponding values, verification criteria, and / or other suitable data. For example, data node 104 may include one or more databases, such as an event database (e.g., a database for storing records, logs, or other information associated with LLM-related user actions), a vector database, an authentication database (e.g., a database for storing authentication tokens associated with users of data generation platform 102), a confidential database, a sensitive token database, and / or a deployment database.

[0051] An event database may contain data associated with events related to the data generation platform 102. For example, the event database stores records associated with user input or prompts used to generate an associated natural language output (e.g., a prompt indicating intent to use an LLM process). The event database may store timestamps and associated user requests or prompts. In some embodiments, the event database may receive records from the data generation platform 102, including model selection / determination, prompt verification information, user authentication information, and / or other suitable information. For example, the event database stores platform-level metrics (e.g., bandwidth data, central processing unit (CPU) utilization metrics, and / or memory utilization associated with devices or servers associated with the data generation platform 102). By doing so, the data generation platform 102 can store and track information related to performance, errors, and troubleshooting. The data generation platform 102 may include one or more subsystems or subcomponents. For example, the data generation platform 102 includes a communication engine 112, an access control engine 114, a leakage mitigation engine 116, a performance engine 118, and / or a generative model engine 120.

[0052] A vector database may contain data associated with vector embeddings of data. For example, a vector database may contain numerical representations (e.g., arrays of values) that represent the semantic meaning of unstructured data (e.g., text, audio, or other similar data). For example, data generation platform 102 receives input such as unstructured data (containing text, such as a prompt) and uses a vector coding model (e.g., using a transformer or neural network architecture) to generate vectors representing the meaning of data objects (e.g., words in a document) within a vector space. By storing information in a vector database, data generation platform 102 can represent input, output, and other data in a processable format (e.g., using an associated LLM), thereby improving the efficiency and accuracy of data processing.

[0053] An authentication database may contain information associated with user or device authentication. For example, the authentication database may contain stored tokens associated with registered users or devices on the data generation platform 102 or its associated development pipeline. For example, the authentication database stores keys (e.g., public keys that match private keys linked to users and / or devices). The authentication database may contain other user or device information (e.g., user identifiers such as user names, or device identifiers such as Media Access Control (MAC) addresses). In some implementations, the authentication database may contain user information and / or restrictions associated with such users.

[0054] A sensitive symbol (e.g., confidential) database may contain data associated with confidential or other sensitive information. For example, confidential information may contain sensitive information such as application programming interface (API) keys, passwords, credentials, or other such information. Sensitive information may include personally identifiable information (PII), such as names, identity card numbers, or biometric information. By storing confidential or other sensitive information, the data generation platform 102 can evaluate prompts and / or outputs to prevent the leakage or disclosure of this sensitive information.

[0055] A deployment repository may contain data associated with the deployment, use, or viewing of results associated with the data generation platform 102. For example, a deployment repository may contain server systems (e.g., physical or virtual) storing verified outputs or results from one or more LLMs, where such results are accessible to requesting users.

[0056] Data generation platform 102 can receive input (e.g., prompts), training data, validation criteria, and / or other suitable data from one or more devices, servers, or systems. Data generation platform 102 can receive such data using communication engine 112, which may include software components, hardware components, or a combination of both. For example, communication engine 112 includes or interfaces with a network card (e.g., a wireless network card and / or a wired network card), which is associated with software used to drive the card and enables communication with network 150. In some embodiments, communication engine 112 can also receive data from data node 104 or another computing device and / or communicate with data node 104 or another computing device. Communication engine 112 can communicate with access control engine 114, leakage mitigation engine 116, performance engine 118, and generative model engine 120.

[0057] In some implementations, data generation platform 102 may include access control engine 114. Access control engine 114 may perform tasks related to user / device authentication, control, and / or permissions. For example, access control engine 114 receives credential information, such as authentication tokens associated with a requesting device and / or user. In some implementations, access control engine 114 may retrieve associated stored credentials (e.g., stored authentication tokens) from an authentication database (e.g., stored within data node 104). Access control engine 114 may include software components, hardware components, or a combination of both. For example, access control engine 114 includes one or more hardware components (e.g., processors) capable of performing operations for authenticating a user, device, or other entity (e.g., service) requesting access to one of the LLMs associated with data generation platform 102. Access control engine 114 may directly or indirectly access data, systems, or nodes associated with third-party databases 108a to 108n and may transfer data to such nodes. Alternatively, the access control engine 114 may receive data from and / or send data to the following components: the communication engine 112, the leakage mitigation engine 116, the performance engine 118, and / or the generative modeling engine 120.

[0058] The leakage mitigation engine 116 can perform tasks related to the verification of inputs and outputs associated with the LLM. For example, the leakage mitigation engine 116 verifies inputs (e.g., prompts) to prevent the leakage of sensitive information or malicious manipulation of the LLM, and verifies the security or safety of the resulting outputs. The leakage mitigation engine 116 may include software components (e.g., modules / virtual machines containing prompt verification models, performance criteria, and / or other suitable data or programs), hardware components, or a combination of both. As an illustrative example, the leakage mitigation engine 116 monitors prompts containing sensitive information (e.g., PII) or other prohibited text to prevent information from leaking from the data generation platform 102 to entities associated with the target LLM. The leakage mitigation engine 116 can communicate with the communication engine 112, access control engine 114, performance engine 118, generative modeling engine 120, and / or other components associated with the network 150 (e.g., data nodes 104 and / or third-party databases 108a to 108n).

[0059] Performance engine 118 can perform performance-related tasks related to monitoring and controlling the data generation platform 102 (e.g., or associated development pipelines). For example, performance engine 118 includes software components (e.g., a performance monitoring module), hardware components, or a combination thereof. For illustration, performance engine 118 can estimate performance metrics (e.g., estimated cost or memory usage) associated with processing a given prompt using a selected LLM. In doing so, performance engine 118 can determine whether to allow user access to a given LLM based on the output of a user request and the associated estimated system effects. Performance engine 118 can communicate with communication engine 112, access control engine 114, performance engine 118, generative modeling engine 120, and / or other components associated with network 150 (e.g., data nodes 104 and / or third-party databases 108a to 108n).

[0060] Generative model engine 120 can perform tasks related to machine learning inference (e.g., natural language generation based on a generative machine learning model (such as an LLM)). Generative model engine 120 may include software components (e.g., one or more LLMs, and / or API calls to devices associated with such LLMs), hardware components, and / or combinations thereof. For illustration, generative model engine 120 may provide user prompts to a requested, selected, or determined model (e.g., an LLM) to generate an output (e.g., a user query within the prompt). Thus, generative model engine 120 enables flexible and configurable generation of data (e.g., text, code, or other suitable information) based on user input, thereby improving the flexibility of software development or other such tasks. Generative model engine 120 can communicate with communication engine 112, access control engine 114, performance engine 118, generative model engine 12 and / or other components associated with network 150 (e.g., data node 104 and / or third-party databases 108a to 108n).

[0061] The engines, subsystems, or other components of data generation platform 102 are illustrative. Therefore, the operation, subcomponents, or other characteristics of a particular subsystem of data generation platform 102 can be distributed, varied, or modified across other engines. In some implementations, specific engines can be deprecated, added, or removed. For example, operations associated with leakage mitigation may be performed at performance engine 118 rather than at leakage mitigation engine 116. [Suitable computing environment for data generation platform] [ ]

[0062] Figure 2 shows a block diagram illustrating some components of at least some of the computer systems and other devices 200 typically incorporated into the disclosed system (e.g., data generation platform 102) operating on some embodiments of the present technology. In various embodiments, such computer systems and (some) other devices 200 may include server computer systems, desktop computer systems, laptop systems, laptops, mobile phones, personal digital assistants, televisions, cameras, car computers, electronic media players, internet services, mobile devices, watches, wearable devices, glasses, smartphones, tablets, smart displays, virtual reality devices, augmented reality devices, etc. In various implementations, the computer system and device includes one or more of the following: input components 204, including a keyboard, microphone, image sensor, touch screen, buttons, trackpad, mouse, optical disc (CD) drive, digital video disc (DVD) drive, 3.5 mm input jack, high-definition multimedia interface (HDMI) input connection, video graphics array (VGA) input connection, universal serial bus (USB) input connection, or other computational input components; output components 206, including a display screen (e.g., liquid crystal display (LCD), organic light-emitting diode (OLED), cathode ray tube (CRT), etc.), speakers, 3.5 mm input jack, microphone, image sensor, touch screen, buttons, trackpad, mouse, optical disc (CD) drive, digital video disc (DVD) drive, 3.5 mm input jack, high-definition multimedia interface (HDMI) input connection, video graphics array (VGA) input connection, universal serial bus (USB) input connection, or other computational input components; mm output jacks, lamps, light-emitting diodes (LEDs), haptic motors, or other output-related components; (a plurality of) processors 208, including a CPU for executing computer programs and a GPU for executing computer graphics programs and processing computing graphics elements; (a plurality of) storage devices 210, including at least one computer memory for storing programs (e.g., (a plurality of) applications 212, (a plurality of) models 214, and other programs) and data when they are used, including facility and associated data, including a core operating system and device drivers; (a plurality of) network connectivity components 216 for enabling the computer system to communicate with other computer systems and send and / or receive data, such as via the Internet or another network and their networking hardware (such as switches, routers, repeaters, cables and fiber optics, optical transmitters and receivers, radio transmitters and receivers, and the like); (a plurality of) permanent storage devices 218, such as a hard disk drive or flash drive for permanent storage of programs and data; and computer-readable media disk drive 220 (For example, at least one non-transitory computer-readable medium), which does not contain a tangible storage component (such as a floppy disk, CD-ROM, or DVD drive) for reading programs and data stored on a computer-readable medium. Although computer systems configured as described above are typically used to support the operation of facilities, those skilled in the art will understand that facilities can be implemented using devices of various types and configurations with various components.

[0063] Figure 3 is a system diagram illustrating one example of a computing environment 300 in which the system disclosed in some embodiments of the present technology operates. In some embodiments, environment 300 includes one or more client computing devices 302a to 302d, instances of which may host a graphical user interface associated with the client device. For example, one or more of client computing devices 302a to 302d include a user device and / or a device associated with a service requesting a response to a query from an LLM. Client computing device 302 operates in a networked environment using a logical connection via network 304 (e.g., network 150) to one or more remote computers (such as a server computing device (e.g., a server system housing the data generation platform 102 of Figure 1)). In some embodiments, client computing device 302 may correspond to device 200 (Figure 2).

[0064] In some embodiments, server computing unit 306 is an edge server that receives client requests and coordinates the fulfillment of those requests through other servers (such as server computing units 310a to 310c). In some embodiments, server computing units 306 and 310 include computing systems. Although each server computing unit 306 and 310 is logically presented as a single server, each server computing unit may be part of a distributed computing environment encompassing multiple computing units located in the same or geographically disparate physical locations. In some embodiments, each server computing unit 310 corresponds to a server group.

[0065] Client computing unit 302 and server computing units 306 and 310 can each be used as a server or client among other servers or client devices. In some embodiments, the server computing units (306, 310a to 310c) are connected to a corresponding database (308, 312a to 312c). For example, the corresponding database includes a database stored in data node 104 (e.g., a sensitive symbol database, an event database, or another suitable database). As discussed above, each server computing unit 310 may correspond to a server group, and each of these servers may share a database or may have its own database (and / or interface with external databases (such as third-party databases 108a to 108n)). In addition to the information described for data node 104 in Figure 1, databases 308 and 312 may also store (e.g., store) other suitable information, such as sensitive or prohibited symbols, user credentials, authentication data, graphical representations, code samples, system policies or other policies, templates, computational languages, data structures, software application identifiers, visual layouts, computational language identifiers, mathematical formulas (e.g., weighted averages, weighted sums or other mathematical formulas), graphical elements (e.g., colors, shapes, text, images, multimedia), system protection mechanisms (e.g., prompts to verify model parameters or criteria), software development or data processing architectures, machine learning models, AI models, training data for AI / machine learning models, historical information or other information.

[0066] Although databases 308 and 312 are logically presented as a single unit, databases 308 and 312 may each encompass a distributed computing environment with multiple computing devices located within their respective servers or at the same or geographically different physical locations.

[0067] Network 304 (e.g., corresponding to network 150) can be a local area network (LAN) or a wide area network (WAN), but can also be other wired or wireless networks. In some embodiments, network 304 is the Internet or some other public or private network. Client computing device 302 is connected to network 304 via a network interface, such as by wired or wireless communication. Although the connection between server computing device 306 and server computing device 310 is shown as a separate connection, such connection can be any type of LAN, WAN, wired network, or wireless network, including network 304 or a separate public or private network. [Example Implementation Plan of the Model in the Data Generation Platform] [ ]

[0068] Figure 4 illustrates one of the AI ​​models according to some embodiments of the present technology. AI model 400 is shown. In some embodiments, AI model 400 can be any AI model. In some embodiments, AI model 400 can be part of or operate in conjunction with server computing device 306 (Figure 3). For example, server computing device 306 can store a computer program that can use information obtained from AI model 400 to provide information to AI model 400 or communicate with AI model 400. In other embodiments, according to some embodiments of the present technology, AI model 400 can be stored in database 308 and can be retrieved by server computing device 306 to perform / process information related to AI model 400.

[0069] In some implementations, AI model 400 may be a machine learning model 402. Machine learning model 402 may comprise one or more neural networks or other machine learning models. As an example, a neural network may be based on a large ensemble of neural units (or artificial neurons). A neural network may loosely mimic the way a biological brain works (e.g., via a large ensemble of biological neurons connected by axons). Each neural unit of a neural network may be connected to many other neural units in the network. These connections may have a coercive or inhibitory effect on the activation state of the connected neural units. In some implementations, individual neural units may have a summing function that combines all the values ​​of their inputs. In some implementations, each connection (or the neural unit itself) may have a limiting function such that a signal must exceed a threshold value before it propagates to other neural units. Such neural network systems can learn and be trained independently, rather than through explicit programming, and may perform significantly better than traditional computer programs in certain areas of problem-solving. In some implementations, the neural network may comprise multiple layers (e.g., a signaling path traverses from a preceding layer to a subsequent layer). In some implementations, backpropagation techniques may be utilized by the neural network, where positive stimulation is used to reset weights on "pre-" neurons. In some implementations, stimulation and inhibition of the neural network can flow more freely, with connections interacting in a more chaotic and complex manner.

[0070] As an instance, with respect to Figure 4 , the machine learning model 402 may acquire input 404 and provide output 406 . In one use case, the output 406 may be fed back to the machine learning model 402 as an input for training the machine learning model 402 (e.g., a user indication, marker associated with the input, or other reference feedback information, alone or in combination with the accuracy of the output 406). In another use case, the machine learning model 402 may update its configuration (e.g., weights, biases, or other parameters) based on its evaluation of its predictions (e.g., output 406) and reference feedback information (e.g., user instructions for accuracy, reference markers, or other information). In another use case, in the case where the machine learning model 402 is a neural network, the connection weights can be adjusted to coordinate the difference between the neural network's prediction and reference feedback. In a further use case, one or more neurons (or nodes) of the neural network may require their and the like respective errors to be sent back to them, etc. via the neural network to facilitate an update procedure (e.g., the backpropagation of the error). For example, an update to the connection weights may be reflected in the magnitude of the error in the backpropagation after a forward traversal has been completed. In this way, for example, the machine learning model 402 may be trained to produce better predictions.

[0071] As an instance, in the case where the prediction model includes one neural network, the neural network may contain one or more input layers, hidden layers, and output layers. The input and output layers may contain one or more nodes, respectively, and the hidden layer may each contain a complex number of nodes. When an entire neural network contains multiple parts trained for different objectives, there may or may not be an input layer or an output layer between the different parts. The neural network may also contain different input layers used to receive various input data. Also, in different instances, the data can be input to the input layer in various forms and to the respective nodes of the input layer of the neural network in various dimensional forms. For example, in a neural network, nodes of a layer other than the output layer are connected via links to a node of a subsequent layer used to transmit output signals or information from the current layer to a subsequent layer. The number of links may correspond to the number of nodes included in subsequent layers. For example, in adjacent fully connected layers, at one point each node of the current layer may have a respective link to each node of the subsequent layer, it should be noted that in some instances such fully connected may be subsequently pruned or minimized during training or optimization. In a one-recursive structure, one node of a layer can be input to the same node or layer again at a subsequent time, whereas in a bidirectional structure, both forward and reverse connections can be provided. Links are also referred to as connection or connection weights, referring to the corresponding “connection weights” that the hardware implements to connect to or provided by such connections of the neural network. During training and implementation, these connections and connection weights may optionally be implemented, removed, and varied to generate or obtain a resulting neural network that is thereby trained and may be implemented accordingly for the trained target (such as target recognition for any of the above instances). [Use the data generation platform to map control gaps to operating standards] [ ]

[0072] Figure 5 is an illustrative diagram of an instance environment 500 of a platform for automatically managing compliance guidelines according to some embodiments of the present technology. Environment 500 includes a user 502, a platform 504, a data provider 506, an AI model agent 508, an LLM 510, a data cache 512, a prompt storage 514, and an execution storage log 516. Platform 504 is implemented using components of instance device 200 and computing device 302, which are illustrated and described in more detail with reference to Figures 2 and 3, respectively. Similarly, embodiments of instance environment 500 may include different and / or additional components, or may be connected in different ways.

[0073] User 502 interacts with platform 504 via, for example, a user interface. Platform 504 may be the same as or similar to the data generation platform 102 shown in Figure 1. User 502 can input data, configure compliance parameters, and manage compliance performance through an intuitive interface provided by the platform. Platform 504 can perform various compliance management tasks, such as compliance checks and regulatory analysis.

[0074] Data provider 506 provides platform 504 with data used in management, which may include regulatory guidance, compliance requirements, organizational guidelines, and other relevant information. The data provided by data provider 506 can be accessed via an application programming interface (API) or database containing policies, obligations, and / or controls within the operating standards. In some implementations, the data supplied by data provider 506 includes publications (e.g., regulatory guidance, compliance requirements, organizational guidelines) themselves. Data provider 506's structured repository allows platform 504 to efficiently retrieve and use data across different management processes. In some implementations, data provider 506 includes existing mappings associated with operating standards. For example, pre-established mappings may exist between operating standards and gaps (e.g., issues). In another instance, pre-established mappings may exist between operating standards and publications. Using existing relationships, platform 504 can more effectively map specific identified gaps to relevant operating standards. For example, if a newly identified gap is similar to or the same as one of the previously identified gaps in a pre-existing mapping (e.g., sharing similar case attributes or meta data labels), then platform 504 can use the pre-existing mapping of the previously identified gap to more easily identify the mapping of the newly identified gap.

[0075] AI model agent 508 acts as an intermediary between the platform 504 and the Large Language Model (LLM) 510. AI model agent 508 facilitates communication and data exchange between platform 504 and LLM 510. In some implementations, AI model agent 508 operates as a plug-in to interconnect platform 504 and LLM 510. In some implementations, AI model agent 508 includes dissimilar modules such as data interception, detection, or action execution. In some implementations, containerization methods such as Docker are used within AI model agent 508 to ensure consistent deployment across environments and minimize dependencies. LLM 510 analyzes data input by user 502 and data obtained from data provider 506 to identify patterns and generate compliance-related outputs. In some implementations, AI model agent 508 enforces access control policies to protect sensitive data and functionality exposed to LLM 510. For example, AI model agent 508 can use encryption standards, token-based authentication, and / or role-based access control (RBAC) to sanitize data received from platform 504 to protect sensitive information. Received data can be encrypted to ensure that all sensitive information is transformed into an unreadable format that can only be accessed through decryption using the appropriate key. Token-based authentication is used by generating a unique token for each user's work phase or transaction. The token serves as a digital identifier by verifying the user's identity and granting access to specific data or functions within the system. Additionally, RBAC can restrict data access based on a user's role within the organization. Specific permissions can be assigned to each role to ensure that users only access data relevant to their responsibilities.

[0076] In some implementations, AI model agent 508 employs content analysis to distinguish between sensitive and non-sensitive information by identifying specific patterns, keywords, or formats that indicate sensitive information. In some implementations, the list of indicators for sensitive information is generated by an internal generative AI model within platform 504 (e.g., using a set of commands similar to "generate multiple instances of PII"). The generative AI model can be trained on a dataset containing instances of sensitive data elements, such as personally identifiable information (PII), financial records, or other confidential information. Once the AI ​​model has been trained, it can generate indicators for sensitive information (e.g., specific patterns, keywords, or formats) based on the learned associations of the model. For example, gap data containing sensitive financial information (such as account numbers, transaction details, and personal information of interested parties) can be identified and subsequently removed and / or masked.

[0077] Data cache 512 can store data for a period of time to reduce the time required to access frequently used information. Data cache 512 ensures that the system can quickly retrieve necessary data without repeatedly querying data provider 506, thus improving the overall efficiency of platform 504. In some implementations, a caching strategy is implemented that includes a cache eviction policy, such as Least Recently Used (LRU) or time-based expiration, to ensure that the cache remains up-to-date and responsive while optimizing memory utilization. LRU allows data cache 312 to track which data items have been recently accessed. When data cache 312 reaches its maximum capacity and an item (e.g., a data packet) needs to be evicted to make room for new data, data cache 312 will remove the least recently used item. Time-based expiration involves setting a specific duration for which a data item is considered valid in data cache 312. Once this duration expires, the data item is automatically invalidated and removed from data cache 312 to preserve space in data cache 312.

[0078] Hint store 514 contains predefined hints that guide LLM 510 in processing data and producing output. Hint store 514 is a repository for storing pre-existing hints in a structured and accessible format (e.g., using a distributed database or NoSQL store), which allows AI model 506 to efficiently retrieve and utilize them. In some embodiments, hints are preprocessed to remove any irrelevant information, standardize the format, and / or organize the hints into a structured database scheme. In some embodiments, hint store 514 is a vector store, where hints are vectorized and stored in a vector space model, and each hint is mapped to a high-dimensional vector representing the semantic features of the hint and its relationships with other hints. In some embodiments, a graph database (such as Neo4j™ or Amazon Neptune™) is used to store hints. Graph databases represent data as nodes and edges, allowing the modeling of relationships between hints to demonstrate interdependencies. In some embodiments, hints are stored in a distributed file system (such as Apache Hadoop™ or Google Cloud Storage™). These systems offer scalable storage for large amounts of data and support parallel processing and distributed computing. Tips stored in a distributed file system can be accessed and processed simultaneously by multiple nodes, allowing for faster retrieval and analysis. For example, details of a specific gap (such as relevant metrics, severity levels, and / or specific publication references) can be structured for use in one of the LLM 510 tips by inserting these details into appropriate locations within a predefined tip.

[0079] Execution storage log 516 records some or all of the actions and procedures performed by platform 504. Execution storage log 516 can be used as audit evidence, providing a history of compliance activities and decisions made by platform 504. The recorded items in execution storage log 515 may include details such as timestamps, user identifiers, specific actions performed, and related contextual information. In some implementations, execution storage log 516 can be accessed via an API through platform 504.

[0080] Figure 6 is an illustrative diagram illustrating an example environment 600 of a platform that generates mapped gaps based on guidelines and gaps in the use controls of some embodiments of the present technology. Environment 600 includes guidelines 602, operating standards 604, gaps 606, platform 608, and mapped gaps 610. Platform 608 is the same as or similar to platform 504 with reference to Figure 5. Embodiments of example environment 500 may include different and / or additional components, or may be connected in different ways.

[0081] Guideline 602 may include publications on regulations, standards, and policies that an organization complies with. Guideline 602 serves as a benchmark for measuring compliance. Guideline 602 may include publications such as jurisdictional guidelines and organizational guidelines. Jurisdictional guidelines (e.g., government regulations) may include guidance collected from authoritative sources such as government websites, legislatures, and regulatory bodies. Jurisdictional guidelines may be published in legal documents or official publications and cover aspects related to the development, deployment, and use of AI technologies within a specific jurisdiction. For example, the California Consumer Privacy Act (CCPA) mandates cybersecurity measures (such as encryption, access controls, and data breach notification requirements) to protect personal data. Therefore, AI developers must implement cybersecurity measures (such as encryption) within the AI ​​models they design and build to ensure the protection of sensitive user data and compliance with regulations. Organizational guidelines include internal policies, procedures, and guidelines established by an organization to manage its operations. Organizational guidelines can be developed that align with industry standards, legal requirements, best practices, and organizational objectives. For example, organizational guidelines may require AI models to include certain access controls to restrict unauthorized access to the model's API or data and / or to have a specific level of flexibility before deployment.

[0082] In some implementations, guidance 602 can be any of text, image, audio, video, or other computer-ingestible formats. For non-text guidance 602 (e.g., image, audio, and / or video), guidance 602 can first be converted into text. Optical character recognition (OCR) can be used for images containing text, and speech-to-text algorithms can be used for audio input. For example, a speech-to-text engine can be used to convert an audio recording detailing financial guidance into text, allowing the system to parse the text output and integrate it into existing guidance 602. Similarly, a video demonstrating a specific procedure or protocol can be processed to extract text information (e.g., extract subtitles).

[0083] In some implementations, where text conversion is not feasible or desired, the system can use vector comparison to directly process non-textual input. For example, image and audio files can be converted into numerical vectors using feature extraction techniques (e.g., using convolutional neural networks (CNNs) for images and Mel-frequency cepstral coefficients (MFCCs) for audio files). The vectors represent corresponding characteristics of the input data (e.g., edges, texture, or shape of an image, or spectral features of an audio file).

[0084] In some implementations, guidance 602 may be stored in a vector store. The vector store stores guidance 602 in a structured and accessible format (e.g., using a distributed database or NoSQL store), allowing platform 608 to efficiently retrieve and utilize it. In some implementations, guidance 602 is preprocessed to remove any irrelevant information, standardize the format, and / or organize guidance 602 into a structured database scheme. Once guidance 602 is ready, it can be stored in a vector store using a distributed database or NoSQL store. To store guidance 602 in a vector store, guidance 602 can be encoded into a vector representation. The textual data of guidance 602 is transformed into numerical vectors that capture the semantic meaning and relationships between words or phrases in guidance 602. For example, word embeddings and / or TF-IDF encoding are used to encode the text into vectors. Word embeddings (such as Word2Vec or GloVe) learn the vector representation of words based on their contextual usage in a large textual corpus. Each word is represented by a vector in a high-dimensional space, where similar words have similar vector representations. TF-IDF (Term Frequency Inverse Document Frequency) encoding calculates the importance of a word in a guide relative to its frequency across the entire corpus of guide 602. For example, the system can assign higher weights to words that are more unique to a specific document and less common across the entire corpus.

[0085] In some implementations, the pointers 602 are stored using a graph database (such as Neo4j™ or Amazon Neptune™). The graph database represents the data as nodes and edges, allowing the relationships between the pointers 602 to demonstrate interdependencies. In some implementations, the pointers 602 are stored in a distributed file system (such as Apache Hadoop™ or Google Cloud Storage™). These systems provide scalable storage for large amounts of data and support parallel processing and distributed computing.

[0086] Vector storage can be stored in a cloud environment hosted by a cloud provider or in a self-hosted environment. In a cloud environment, vector storage offers scalability through cloud services provided by a platform such as AWS™ or Azure™. Storing vector storage in a cloud environment requires selecting cloud services, dynamically deploying resources through the provider's interface or API, and configuring networking components for secure communication. Cloud environments allow vector storage to expand its storage capacity without manual intervention. As storage demand grows, additional resources can be automatically deployed to meet the increased workload. Furthermore, cloud-based caching modules can be accessed from anywhere with an internet connection, providing convenient access to historical data for users across different locations or devices. Conversely, in a self-hosted environment, vector storage is stored on a private network server. Deploying vector storage in a self-hosted environment requires setting up the server using the necessary hardware or virtual machines, installing an operating system, and storing the vector storage. In a self-hosted environment, an organization has complete control over vector storage, allowing it to implement financial measures and compliance policies tailored to its specific needs. For example, organizations in industries with stringent data privacy and financial regulations (such as financial institutions) can mitigate security risks by storing vector storage in a self-hosted environment.

[0087] Operating Standard 604 can be a specific obligation derived from a guideline to comply with that guideline, and can encompass both specific operational guidelines and general principles. In some instances, Operating Standard 604 can serve as an operational guideline that an organization must follow to meet the requirements set forth in regulatory guidelines or industry best practices (e.g., Guideline 602). For example, an operating standard derived from a data protection guideline may mandate the adoption of a specific framework (e.g., the General Data Protection Regulation (GDPR)) for the handling of personal data, outlining data access procedures, encryption standards, and breach notification agreements. In another instance, an operating standard may include prohibitions on a specific action, such as transmitting confidential information to external sources. In further instances, Operating Standard 604 encompasses fundamental principles or benchmarks derived from guidelines that guide organizational practices and behaviors toward achieving desired results. For example, within the context of an organization's ethical standards, an operating standard may include principles such as integrity, transparency, and accountability.

[0088] Gap 606 is an example of a situation where current controls or procedures do not meet operational standards. Gap 606 can be attributed to a lack of required controls or inadequacy of existing controls. For example, in the context of data security, a gap can be identified if a company lacks a comprehensive data encryption policy, despite regulatory requirements specifying encryption standards for sensitive information. In another instance, a gap can be identified if an organization, despite having implemented access controls for sensitive systems, fails to regularly review and update user permissions as required by industry best practices, thereby allowing potential vulnerabilities to go unaddressed.

[0089] Gap 606 can be managed through a systematic approach that incorporates self-reporting and comprehensive storage of attributes tailored to each case associated with Gap 606. A case of Gap 606 refers to a specific instance or situation within an organization where current controls or procedures do not meet established operating standards 604. Each case associated with a Gap 606 represents a distinct use case. For example, a case may include a cybersecurity vulnerability attributable to insufficient data encryption practices, or a compliance issue related to incomplete documentation of financial transactions. Each identified Gap 606 can be recorded using case attributes (e.g., meta-data, tags) (such as a descriptive title, severity rating (e.g., on a scale of 1 to 5, where 1 represents severe and 5 represents negligible) and / or tags linking Gap 606 to specific business units or regulatory requirements). Case attributes provide a clear understanding of the impact and context of the gap. In some implementations, platform 608 includes a user interface that allows users to input and edit the case attributes of each gap 606.

[0090] Platform 608 receives instructions, operating standards, and / or identified gaps, and generates a mapped gap 610. The mapped gap correlates the identified gap with a specific operating standard that the identified gap failed to meet. The method of mapping identified gaps using specific operating standards is further discussed in Figure 7.

[0091] Figure 7 is a flowchart illustrating a procedure 700 for mapping identified gaps in control to operational criteria according to some embodiments of the present technology. In some embodiments, procedure 700 is performed by components of the example device 200 and the computing device 302, which are illustrated and described in more detail with reference to Figures 2 and 3, respectively. A specific entity, such as LLM 510, is illustrated and described in more detail with reference to Figure 5. Similarly, embodiments may include different and / or additional steps or may perform steps in a different order.

[0092] In action 702, the system determines a set of vector representations of alphanumeric characters represented by one or more operational standards, which contain a first set of actions configured to comply with the constraints in that set of vector representations. This set of vector representations of alphanumeric characters is the same as or similar to the publication of guideline 602 described with reference to Figure 6. Figure 6 describes a method for transforming different modes of the guideline (e.g., text, image, audio, video) into vector representations.

[0093] In some implementations, the system receives an indicator of a type of operation associated with a vector representation. The system identifies a set of associated operation criteria associated with the type of vector representation. The system obtains this set of associated operation criteria via an application programming interface (API). For example, the system includes an input channel or interface capable of receiving signals or data tags indicating the type (e.g., nature or purpose) of the vector representation being processed. The system can use an API to retrieve this set of associated operation criteria by implementing an API endpoint or integration point that connects the system to a centralized repository or database of operation criteria, which may use associated post-data tags related to the type of vector representation.

[0094] In some implementations, the AI ​​model is a first AI model. The system can supply the original publication of the set of vector representations or guidelines to a second AI model. In response, the system can receive from the second AI model a set of summaries of the set of vector representations, wherein at least one of the prompts in the set contains one or more of the summaries in the set. The set of summaries is a representation of the set of vector representations. In some implementations, the set of summaries serves as a refined and coherent representation of the text content derived from the set of vector representations. The set of summaries encompasses key themes, sentiments, or relevant information embedded in the guidelines. The summarization process not only captures the essence of the user's sentiment but also allows for effective understanding and analysis. By compressing large amounts of text content into a compressed summary (e.g., the set of summaries), the system allows the user to obtain a comprehensive and accessible understanding of the guidelines. For example, prompts input to the second AI model may request a summary of the provided text or guidelines by including prompts such as "Summarize the following text as key points" or "Provide a concise summary capturing the main themes and most important information." Additionally, the prompts may include background information or specific aspects of the issue, such as "providing key regulatory requirements and their meanings." The prompts may also include definitions of specific terms, such as operating standards or controls.

[0095] In action 704, the system receives an output generation request via a user interface, which includes inputs for generating an output using a Large Language Model (LLM). The input contains a set of gaps associated with one or more cases that fail to meet one or more operational criteria of that set of vector representations. An example of a gap is illustrated with reference to gap 606 in Figure 6. Each case is associated with a unique identifier and a corresponding metric, indicating one or more missing actions from the first set of actions in the case. An example of a case is illustrated with reference to Figure 6. Each gap in this set of gaps contains a set of attributes defining the case, including the case's unique identifier, the case's corresponding metric, the corresponding vector representation associated with the case, a case title, a case overview, and / or a case severity level.

[0096] In some implementations, the set of attributes defining a case includes a binary indicator of the case's severity level, a category for the case's severity level, and / or a probability associated with the case's severity level. For example, the binary indicator could be set to "1" for severe (indicating an issue requiring immediate attention) or "0" for non-severe (where the issue is less urgent but still requires a solution). In another instance, the category could range from "low" to "high" severity, helping to prioritize remedial actions based on the potential impact and risks associated with each case. In a further example, a high probability value could indicate that, if not addressed promptly, the compliance gap is highly likely to result in regulatory penalties or data breaches.

[0097] In action 706, using the received input, the system constructs a set of prompts for each gap in the set of gaps. The set of prompts for a specific gap includes the set of attributes defining the case, such as a case identifier, severity assessment (e.g., criticality level), an overview of the compliance issue, a first set of actions for one or more operational criteria (e.g., actionable guidelines or general principles in Figure 6), and / or the set of vector representations. In some implementations, the set of prompts for each gap in the set of gaps includes a set of preloaded query context defining one or more sets of alphanumeric characters associated with the set of vector representations. The preloaded query context includes predefined templates, rules, or configurations specifying criteria for mapping gaps to operational criteria. For example, the preloaded query context may include definitions of terms such as operational criteria and / or gaps. The prompts are used as input to a large language model (LLM), which is designed to process natural language input and produce structured output based on learned patterns and data.

[0098] In action 708, for each gap in the set of gaps, the system maps the gap to one or more operational standards in the set of vector representations. The system provides a prompt for the specific gap to the LLM. In response to the input prompt, the system receives from the LLM a set of gap-specific operational standards containing one or more operational standards associated with the specific gap. In some implementations, the system may generate an interpretation of how each gap is mapped to one or more operational standards for each gap-specific operational standard in the set of gap-specific operational standards. The output of the LLM may be presented in the form of alphanumeric characters. In some implementations, in response to the input prompt, the system receives the set of gap-specific operational standards and the corresponding set of vector representations from the AI ​​model.

[0099] In some implementations, the prompts in the LLM include guidance on providing a first interpretation of why a particular gap should be mapped to a particular operating standard and a second interpretation of why a particular gap should not be mapped to a particular operating standard. The prompts may further include guidance on why the first or second interpretation is more weighted (e.g., why a particular mapping occurs). In some implementations, an individual may approve or disapprove the mapping based on the first and / or second interpretation. Allowing an individual in the loop (HITL) and generating both a first and second interpretation provides platform users with transparency regarding the generated mappings.

[0100] In action 710, the system generates a graphical representation indicating one of the specific operational standards for the set of gaps to be displayed at the user interface. The graphical representation includes a first representation of each gap in the set and a second representation of one of the corresponding operational standards for that set of gaps. In some implementations, each gap is visually represented to highlight its specific attributes, such as severity level, case identifier, and an overview of the detailed gap. The graphical representation may use integrated color coding, diagrams, or annotations to represent severity levels, compliance progress, or overdue actions using charts, graphs, or visual frameworks. Annotations within the graphical representation may provide additional background information or explanations about each gap and its consistency with the operational standards. Overlays may be used to indicate overdue actions, completed mappings, and / or compliance deadlines.

[0101] In action 712, using the set of gap-specific operational criteria, the system generates a second set of actions for each gap in the set, containing one or more actions from a first set of actions indicated by the corresponding gap-specific operational criteria. The second set of actions can modify a portion of a case within the corresponding gap to satisfy one or more operational criteria represented by the vector. For example, actions may involve updating policies, enhancing security measures, implementing new agreements, and / or conducting training phases to improve organizational practices and mitigate risks. Each action can be directly linked to the corresponding gap and its associated operational criteria.

[0102] In some implementations, the set of prompts is a first set of prompts, and the specific operational criteria for the set of gaps is a first set of operational criteria. Using the received input, the system constructs a second set of prompts for each gap in the set of gaps. The second set of prompts for a specific gap includes the set of attributes defining the case and the set of vector representations. Using the second set of prompts, the system receives a second set of operational criteria from the LLM for each gap in the set of gaps. Using the second set of operational criteria, the system constructs a third set of prompts for each gap in the set of gaps. The third set of prompts for a specific gap includes the set of attributes defining the case and a first set of actions for one or more operational criteria. Using the third set of prompts, the system receives a third set of operational criteria from the LLM for each gap in the set of gaps. This iterative approach of using multiple sets of prompts with the LLM enhances the system's ability to dynamically adapt and respond to previously generated mappings and thus facilitates a continuous improvement process, where insights gained from each interaction cycle lead to more refined strategies for achieving consistency between organizational and operational criteria.

[0103] In some implementations, this set of prompts is a first set of prompts. For each vector representation in the received vector representation set, the system identifies a set of text content representing that set of vector representations. The system partitions this set of text content into multiple subsets of text content based on predetermined criteria. The predetermined criteria may include the length and / or complexity of each subset of text content. For example, the predetermined criteria may be a token count or character limit to ensure the consistency and coherence of the segmentation process. Blocking the text content breaks down large amounts of text content into manageable units. For token-based partitions, the system calculates the number of linguistic units or tokens within the text content. In some implementations, depending on the specific language analysis used, these tokens cover individual words, phrases, or even characters. Predetermined token counting criteria set a quantitative guideline specifying the number of linguistic units covered within each block. In some implementations, when a single-character constraint criterion is used, the system focuses on the total number of characters within the character constraint criterion for the text content. In other implementations, it involves evaluating numeric characters and spaces, providing a more granular measure of the structural complexity of the content. A predefined character constraint establishes an upper limit value, guiding the system to build segments that adhere to the predefined character constraint.

[0104] The system can receive user feedback related to deviations between a set of gap-specific operational standards and a set of desired operational standards. The system can iteratively adjust the set of prompts to modify the gap-specific operational standards to the desired operational standards. The system can generate action plans, update compliance strategies, and / or refine operational practices to enhance consistency with the set of vector representations. The system can generate a set of actions (e.g., a modification plan) to adjust a case's current attributes to a set of desired attributes for that case. The system can identify the root cause of the discrepancy between a case's attributes and its desired attributes. For example, the desired attributes for a case might include a specific action (e.g., an anonymization procedure) not found in the case's current attributes. Actions (e.g., anonymization procedures) can be pre-loaded into the system. [The platform uses data generation to automatically guide the creation of actionable projects] [ ]

[0105] Figure 8 is an illustrative diagram of an example environment 800 of a platform according to some embodiments of the present technology, in which a self-guided instruction 802 identifies operable items 810a to 810n. Environment 800 includes instruction 802, platform 804, text subsets 806a to 806n, prompts 808a to 808n, and operable items 810a to 810n. Instruction 802 is the same as or similar to instruction 602 with reference to Figure 6. Platform 804 is the same as or similar to platform 504 with reference to Figure 5. Embodiments of example environment 800 may include different and / or additional components, or may be connected in different ways.

[0106] Platform 804 can be a web-based application that hosts various use cases (such as compliance) and allows users to interact via a front-end interface. Input to platform 804 can be instructions 802 in various formats (e.g., text, Excel). Further examples of platform 804 are discussed with reference to platform 504 in Figure 5. The backend of platform 804 can segment (e.g., partition) the instructions into text subsets 806a to 806n and vectorize these text subsets. The vectorized representations of text subsets 806a to 806n can be stored in a database accessible by platform 804. Platform 804 can use an API call to send prompts to an AI model (such as an LLM), as further described in Figure 5. The AI ​​model processes the prompts and returns the output of actionable items to the backend of platform 804, which can format the output into a user-friendly structure.

[0107] Text subsets 806a to 806n refer to portions of Guideline 802 that have been extracted or segmented into smaller pieces (e.g., based on specific criteria). Each text subset 806a to 806n can be categorized by topic, section, or other relevant factors. By breaking down large amounts of text into subsets, the platform can focus on specific parts of the guideline. This structured approach further allows the platform to efficiently handle large volumes of regulatory text.

[0108] Hints 808a to 808n are specific queries or instructions generated from text subsets 806a to 806n, which are designed to guide the behavior and output of an AI model, such as identifying actionable items from text subsets 806a to 806n of regulatory guidance 802. For example, a corresponding hint 808a is constructed for text subset 806a. In some implementations, a hint may contain multiple text subsets. In some implementations, a single text subset may be associated with multiple hints. Hints 808a to 808n cause the AI ​​model to identify specific attributes (such as regulatory obligations or compliance requirements) of text subsets 806a to 806n to dynamically generate meaningful outputs (e.g., actionable items). In some implementations, a second AI model may be used to generate hints 808a to 808n. The second AI model can directly analyze text subsets 806a to 806n or guidelines 802 to identify features of the text subsets, such as background content, relationships between entities and features, by, for example, breaking down the input into smaller components and / or tagging predefined keywords. The second AI model can use the identified features to construct context-relevant prompts. For example, if the input is related to compliance guidelines, the second AI model can identify portions of the guidelines and highlight prompts that emphasize the most relevant information (e.g., information related to compliance guidelines). Prompts may include specific questions or statements that guide the first AI model to focus on a particular type of situation (e.g., "What are the key data protection compliance requirements in this guideline?").

[0109] In some implementations, the second AI model may employ query expansion. Query expansion is a process of enhancing the original query to improve the comprehensiveness of the response by including synonyms, related concepts, and / or additional context-relevant terms. For example, if the initial prompt is "identify key actionable items for data protection," the second AI model may expand the query by including keywords such as "privacy regulations," "data security measures," and "information control." In some implementations, the second AI model may refer to domain-specific thesaurus and / or pre-trained word embeddings to find synonyms and related terms for the identified elements.

[0110] Tips 808a to 808n may include definitions, keywords, and instructions that guide AI models to identify relevant actionable items. For example, a definition may clarify what constitutes an "actionable item" or "obligation." Furthermore, tips 808a to 808n may specify keywords such as "must," "should," or "required." Keywords may indicate mandatory actions or prohibitions that need to be identified as actionable items. For example, a tip may instruct the AI ​​model to flag any sentence containing the word "must," as it may represent a regulatory requirement. In another instance, tips 808a to 808n may guide the AI ​​model to extract all examples of deadlines for compliance actions, descriptions of required documents, or procedures for reporting to regulatory agencies. Instructions may also include formatting guidelines to ensure that extracted actionable items are presented in a consistent and usable format.

[0111] Actionable items 810a to 810n (e.g., guidelines, actions) are specific tasks or requirements identified by the AI ​​model from the guidelines based on the analysis of text subsets 806a to 806n and prompts 808a to 808n. In some implementations, actionable items 810a to 810n may be distilled into comprehensive instructions defining specific measures or procedures to be implemented, rather than merely extracts from text subsets 806a to 806n. For example, an actionable item may outline the frequency and format of required compliance reports, specify the information to be included, and designate the department responsible for submission. Actionable items 810a to 810n are designed to translate regulatory text into actionable steps that the organization can directly implement. Actionable items 810a to 810n may include tasks such as reporting, record keeping, compliance checks, and other regulatory actions.

[0112] Each actionable item may include meta-data, such as the types of responsible parties within the organization, affected customers or stakeholders, and / or other relevant identifiers. An AI model may use Natural Language Processing (NLP) algorithms to parse subsets of text 806a to 806n to identify relevant phrases, keywords, and semantic structures (e.g., as indicated by prompts 808a to 808n) within instruction 802. Prompts 808a to 808n may guide the AI ​​model by providing contextual prompts and specific queries that guide the AI ​​model to focus on a particular instruction or state of instruction within instruction 802. [An Example Implementation Plan for a Verification Engine, One of the Data Generation Platforms] [ ]

[0113] Figure 9 is a block diagram illustrating one instance environment 900 used to determine AI compliance by using guidelines input to a verification engine according to some embodiments of the present technology. Environment 900 includes guidelines 902 (e.g., governing regulations 904, organizational regulations 906, AI application-specific regulations 908), a vector storage 910, and a verification engine 912. The verification engine may be the same as or similar to the generative model engine 120 in the data generation platform 102 discussed with reference to Figure 1. The vector storage 910 and verification engine 912 are implemented using components of the instance device 200 and computing device 302, which are illustrated and described in more detail with reference to Figures 2 and 3, respectively. Similarly, embodiments of instance environment 900 may include different and / or additional components, or may be connected in different ways.

[0114] Guideline 902 may include various elements, such as governing regulations 904, organizational regulations 906, and AI application-specific regulations 908 (e.g., unsupervised learning, natural language processing (NLP), generative AI). Governing regulations 904 (e.g., government regulations) may include regulations collected from authoritative sources such as government websites, legislative bodies, and regulatory agencies. Governing regulations 904 may be published in legal documents or official publications and cover the development, deployment, and use of AI technologies within a specific jurisdiction. Organizational regulations 906 include internal policies, procedures, and guidelines established by an organization to manage AI-related activities within its operations. Organizational regulations 906 may be developed in accordance with industry standards, legal requirements, and organizational objectives. AI application-specific regulations 908 include regulations related to specific types of AI applications, such as unsupervised learning, natural language processing (NLP), and generative AI. Each type of AI application presents unique challenges and considerations regarding compliance, ethical use, and / or regulatory compliance. For example, unsupervised learning algorithms where models learn from unlabeled input data can withstand regulations that prevent bias and discrimination in unsupervised learning models. Natural Language Processing (NLP) technologies that enable computers to understand, interpret, and generate human language can withstand specific regulations intended to protect user privacy. Generative AI that autonomously creates new content can focus on intellectual property rights, content moderation, and ethical use cases. AI developers may need to incorporate additional mechanisms for copyright protection, content filtering, and / or user consent management to comply with regulations related to generative AI technologies.

[0115] Instruction 902 is stored in a vector storage 910. Vector storage 910 stores instruction 902 in a structured and accessible format (e.g., using a distributed database or NoSQL storage), which allows verification engine 912 to efficiently retrieve and utilize it. In some implementations, instruction 902 is preprocessed to remove any irrelevant information, standardize the format, and / or organize instruction 902 into a structured database scheme. Once instruction 902 is ready, it can be stored in vector storage 910 using a distributed database or NoSQL storage.

[0116] To store guide 902 in vector storage 910, guide 902 can be encoded into a vector representation for subsequent retrieval by verification engine 912. The textual data of guide 902 is transformed into numerical vectors that capture the semantic meaning and relationships between words or phrases within guide 902. For example, word embeddings and / or TF-IDF encoding are used to encode the text into vectors. Word embeddings (such as Word2Vec or GloVe) learn the vector representation of words based on their contextual usage within a large textual corpus. Each word is represented by a vector in a high-dimensional space, where similar words have similar vector representations. TF-IDF (Term Frequency Inverse Document Frequency) encoding calculates the importance of a word in a guide relative to its frequency across the entire corpus of guide 902. For example, the system can assign higher weights to words that are more unique to a specific document and less common across the entire corpus.

[0117] In some implementations, the guide 902 is stored using a graph database (such as Neo4j™ or Amazon Neptune™). The graph database represents the data as nodes and edges, allowing the relationships between the guides 902 to demonstrate interdependencies. In some implementations, the guide 902 is stored in a distributed file system (such as Apache Hadoop™ or Google Cloud Storage™). These systems provide scalable storage for large amounts of data and support parallel processing and distributed computing. Guides 902 stored in a distributed file system can be accessed and processed simultaneously by multiple nodes, allowing for faster retrieval and analysis via a verification engine.

[0118] The Vector Storage 910 can be stored in a cloud environment hosted by a cloud provider or in a self-hosted environment. In a cloud environment, the Vector Storage 910 offers scalability through cloud services provided by a platform such as AWS™ or Azure™. Storing the Vector Storage 910 in a cloud environment requires selecting cloud services, dynamically deploying resources through the provider's interface or API, and configuring networking components for secure communication. The cloud environment allows the Vector Storage 910 to expand its storage capacity without manual intervention. As storage demand grows, additional resources can be automatically deployed to meet the increased workload. Furthermore, the cloud-based caching module can be accessed from anywhere with an internet connection, providing convenient access to historical data for users across different locations or devices.

[0119] Conversely, in a self-hosted environment, the vector storage 910 is stored on a private network server. Deploying the vector storage 910 in a self-hosted environment requires setting up the server using the necessary hardware or virtual machines, installing an operating system, and storing the vector storage 910. In a self-hosted environment, the organization has complete control over the vector storage 910, allowing it to implement customized security measures and compliance policies tailored to its specific needs. For example, organizations in industries with stringent data privacy and security regulations (such as financial institutions) can mitigate security risks by storing the vector storage 910 in a self-hosted environment.

[0120] Verification engine 912 accesses guide 902 from vector storage 910 to initiate a compliance assessment. Verification engine 912 can establish a connection to one of the vector storage devices 910 using an appropriate API or database driver. The connection allows verification engine 912 to query vector storage 910 and retrieve relevant guides for the AI ​​application being assessed. Frequently accessed guides 902 are stored in memory, which allows verification engine 912 to reduce latency and improve response time for compliance assessment tasks. In some implementations, relevant guides are retrieved only based on the specific AI application being assessed. For example, data tags, categories, or keywords associated with the AI ​​application can be used to filter guides 902.

[0121] Verification Engine 912 assesses the compliance of AI applications with Guideline 902 (e.g., the use of semantic search, pattern recognition, and machine learning techniques). For example, Verification Engine 912 compares vector representations of different interpretations and results by calculating the cosine of the angle between two vectors indicating directional similarity. Similarly, to compare interpretations, Verification Engine 912 can measure the intersection within the union of expected and case-specific word groups.

[0122] Figure 10 is a block diagram illustrating one instance environment 1000 used to generate verification actions to determine the compliance of an AI model according to some embodiments of the present technology. Environment 1000 includes training data 1002, a meta-model 1010, verification actions 1012, a cache 1014, and a vector storage 1016. The meta-model 1010 is the same as or similar to the meta-model 402 illustrated and described in more detail with reference to Figure 4. The meta-model 1010 is implemented using components of the instance device 200 and the computing device 302 illustrated and described in more detail with reference to Figures 2 and 3, respectively. Similarly, embodiments of instance environment 1000 may include different and / or additional components, or may be connected in different ways.

[0123] The training data includes information from sources such as enterprise application 1004, other AI applications 1006, and / or an internal document search AI 1008. Enterprise application 1004 refers to various types of software tools or systems used to facilitate business operations and may contain data related to, for example, loan transaction history, customer financial profiles, credit scores, and income verification documents. For example, data from a banking application can provide insights into an applicant's banking behavior, such as average account balance, transaction frequency, and bill payment history. Other AI applications 1006 may include, for example, credit scoring models, fraud detection algorithms, and risk assessment systems that can be used by lenders to evaluate loan applications. Data from AI application 1006 refers to various software systems that utilize artificial intelligence (AI) technology to perform specific tasks or functions. The data may include credit risk scores and fraud risk indicators. For example, an AI-driven credit scoring model can provide a risk assessment score based on an applicant's credit history, debt-to-income ratio, and other financial factors. Internal Document Search AI 1008 is an AI system customized for searching and extracting information from internal documents within an organization. For example, Internal Document Search AI 1008 can be used to retrieve and analyze relevant documents such as loan agreements, regulatory compliance documents, and internal policies. Information from internal documents may include, for example, legal disclosures, loan terms and conditions, and compliance guidelines. For instance, the AI ​​system can flag loan applications that contain discrepancies or inconsistencies with regulatory guidelines or internal policies.

[0124] Training data 1002 is fed into meta-model 1010 to train it, enabling it to learn patterns and characteristics associated with compliant and non-compliant AI behavior. Further discussion of the artificial intelligence and training methods is illustrated in Figure 7. Meta-model 1010 uses the learned patterns and characteristics to generate verification actions 1012, which serve as potential use cases designed to evaluate the compliance of the AI ​​model. Verification actions 1012 can encompass various cases and use cases related to the specific application areas of the AI ​​model being evaluated. Further methods for establishing verification actions are illustrated in Figures 12 through 14.

[0125] In some implementations, the generated verification action 1012 may be stored in a cache 1014 and / or a vector store 1016. The cache 1014 serves as a temporary storage location for recently accessed or frequently used verification actions, facilitating efficient retrieval when needed. On the other hand, the vector store 1016 provides a vector representation for storing verification actions, enabling a structured repository for efficient storage and retrieval based on similarity or other criteria. The vector store 1016 stores the generated verification action 1012 in a structured and accessible format (e.g., using a distributed database or NoSQL store), allowing the meta-model 1010 to perform efficient retrieval and utilization. In some implementations, the generated verification action 1012 is preprocessed to remove any irrelevant information, standardized in format, and / or organized into a structured repository scheme. Once the generated verification action 1012 is ready, it can be stored in a vector storage 1016 using a distributed database or NoSQL storage.

[0126] In some implementations, the generated verification actions 1012 are stored using a graph database (such as Neo4j™ or Amazon Neptune™). The graph database represents data as nodes and edges, allowing the relationships between the generated verification actions 1012 to demonstrate interdependencies. In some implementations, the generated verification actions 1012 are stored in a distributed file system (such as Apache Hadoop™ or Google Cloud Storage™). The system provides scalable storage for large amounts of data and supports parallel processing and distributed computing. The generated verification actions 1012 stored in a distributed file system can be accessed and processed simultaneously by multiple nodes, allowing for faster retrieval and analysis via the meta-model 1010.

[0127] The Vector Storage 1016 can be stored in a cloud environment hosted by a cloud provider or in a self-hosted environment. In a cloud environment, the Vector Storage 1016 offers scalability through cloud services provided by a platform (e.g., AWS™, Azure™). Storing the Vector Storage 1016 in a cloud environment requires selecting cloud services, dynamically deploying resources through the provider's interface or API, and configuring networking components for secure communication. The cloud environment allows the Vector Storage 1016 to expand its storage capacity without manual intervention. As storage demand grows, additional resources can be automatically deployed to meet the increased workload. Furthermore, the cloud-based caching module can be accessed from anywhere with an internet connection, providing convenient access to historical data for users across different locations or devices.

[0128] Conversely, in a self-hosted environment, the vector storage 1016 is stored on a private network server. Deploying the vector storage 1016 in a self-hosted environment requires configuring the server using the necessary hardware or virtual machines, installing an operating system, and storing the vector storage 1016. In a self-hosted environment, the organization has complete control over the vector storage 1016, allowing the organization to implement customized security measures and compliance policies tailored to its specific needs. For example, organizations in industries with stringent data privacy and security regulations (such as financial institutions) can mitigate security risks by storing the vector storage 1016 in a self-hosted environment.

[0129] The metamodel 1010 accesses the vector storage 1016 to generate verification actions 1012 to initiate a compliance assessment. The system can establish a connection to one of the vector storage 1016 using an appropriate API or database driver. The connection allows the metamodel 1010 to query the vector storage 1016 and retrieve the relevant vector constraints of the AI ​​application being evaluated. Frequently accessed verification actions 1012 are stored in memory, which allows the system to reduce latency and improve the response time of compliance assessment tasks.

[0130] In some implementations, relevant verification actions are retrieved only based on the specific AI application being evaluated. For example, data tags, categories, or keywords associated with the AI ​​application can be used to filter verification action 1012. Relevant verification actions can be specifically selected based on the specific background content and requirements of the AI ​​application being evaluated. For example, the system analyzes data tags, keywords, or categories associated with verification action 1012 stored in the system's database. Using the specific background content and requirements of the AI ​​application, the system filters and retrieves relevant verification actions from the database.

[0131] Various filtering criteria can be used to select relevant verification actions. In some implementations, the system uses Natural Language Processing (NLP) to parse the text of verification action 1012 and identify key terms, phrases, and clauses representing regulatory obligations relevant to the domain of the AI ​​application. Specific terms relevant to the domain of the AI ​​application can be predefined and include, for example, "patient privacy" for a healthcare application. Using specific terms relevant to the domain of the AI ​​application as a filtering criterion, the system can filter out irrelevant verification actions. To identify relevant verification actions from verification action 1012, the system can determine specific terms to be used as filtering criteria by calculating the similarity between vectors representing domain-specific terms (e.g., "healthcare") and vectors representing other domain-related terms (e.g., "patient privacy"). Domain-specific terms can be identified based on the proximity of other terms to known terms of interest. A similarity threshold can be applied to filter out terms that are not sufficiently similar to known domain-specific terms.

[0132] In some implementations, the system can tag relevant verification actions with attributes that help contextualize them. Tags serve as markers to categorize and organize verification actions 1012 based on predefined criteria (such as regulatory themes (e.g., data privacy, fairness, transparency) or jurisdictional relevance (e.g., regional regulations, industry standards)). Tags provide a structured representation of verification actions 1012 and allow for easier retrieval, manipulation, and analysis of regulatory content. Tags and associated meta-data can be stored in a structured format (such as a database) where each verification action 1012 is linked to its corresponding tag and / or regulatory deployment.

[0133] Metamodel 1010 assesses the compliance of AI applications with vector constraints using verification actions 1012 (e.g., semantic search, pattern recognition, and machine learning techniques). Further assessment methods for determining the compliance of AI applications are discussed with reference to Figures 12 to 14.

[0134] Figure 11 is a block diagram illustrating one example environment 1100 for automatically performing correction actions on an AI model according to some embodiments of the present technology. Environment 1100 includes a training dataset 1102, a meta-model 1104 (which includes validation models 1106A to 1106D, validation actions 1108, and an AI application 1110), results and interpretations 1112, suggestions 1114, and correction actions 1116. Meta-model 1104 is the same as or similar to meta-model 1010, which is illustrated and described in more detail with reference to Figure 10. Meta-model 1104 and AI application 1110 are implemented using components of example device 200 and computing device 302, which are illustrated and described in more detail with reference to Figures 2 and 3, respectively. Similarly, embodiments of example environment 1100 may include different and / or additional components, or may be connected in different ways.

[0135] Training dataset 1102, containing data used to train machine learning models, is input into meta-model 1104. Meta-model 1104 is a composite model encompassing several sub-models customized to address specific aspects of AI compliance. Within meta-model 1104 are various specialized models, such as a bias model 1106A (described in further detail with reference to Figure 5), a toxicity model 1106B (described in further detail with reference to Figure 6), an IP infringement model 1106C (described in further detail with reference to Figure 7), and other verification models 1106D. Each model is responsible for detecting and evaluating specific types of non-compliant content within the AI ​​model. When processing training dataset 1102, each model generates verification actions customized to assess the presence or absence of specific types of non-compliant content. Further evaluation techniques using verification actions generated by meta-model 1104 are discussed with reference to Figures 12 through 14.

[0136] The generated verification action 1108 is provided as input to an AI application 1110 in the form of a prompt. The AI ​​application 1110 processes the verification action 1108 and generates a result and an explanation 1112 detailing how the result was determined. Subsequently, based on the result and explanation 1112 provided by the AI ​​application 1110, the system can generate suggestions 1114 for corrective actions. These suggestions are derived from the analysis of the verification action results and are intended to address any identified problems or deficiencies. For example, if certain verification actions fail to meet desired criteria due to specific attribute values ​​or patterns, the suggestions may recommend adjustments to those attributes or modifications to the basic procedures.

[0137] For a bias detection model (such as the ML model discussed in Figure 5), if certain attributes exhibit unexpected correlations or distributions, the system can retrain the tested AI model with a revised weighting scheme to better align with the desired vector constraints. In a toxicity model (such as the ML model discussed in Figure 6), corrective actions may include implementing post-processing techniques in the tested AI model to filter out responses that violate vector constraints (e.g., filtering out responses containing identified vector representations of alphanumeric characters). Similarly, in an IP infringement model (such as the ML model discussed in Figure 7), corrective actions may include implementing post-processing techniques in the tested AI model to filter out responses that infringe on IP rights (e.g., filtering out responses containing predetermined alphanumeric characters).

[0138] In some implementations, based on the results and interpretations, the system applies predefined rules or logic to determine appropriate corrective actions. These rules can be established by the user and can take into account factors such as regulatory compliance, risk assessment, and corporate objectives. For example, if an application is rejected due to insufficient revenue, the system may suggest that the applicant request additional financial documentation.

[0139] In some implementations, the system can use machine learning models to generate recommendations. The model learns from historical data and past decisions to identify patterns and trends that indicate a series of actions the AI ​​model can take to comply with vector constraints. By training on a dataset of past corrective actions and results, the machine learning model can predict the most effective recommendations for new cases. Further discussion of artificial intelligence and training methods is illustrated in Figure 7. Recommendation 1114 can be automatically implemented by the system as corrective action 1116. This automation simplifies the process of resolving identified problems and ensures rapid remediation of non-compliant content within the AI ​​model, thereby enhancing overall compliance and reliability.

[0140] Figure 11 is a block diagram illustrating one example environment 1100 for automatically performing correction actions on an AI model according to some embodiments of the present technology. Environment 1100 includes a training dataset 1102, a meta-model 1104 (which includes validation models 1106A to 1106D, validation actions 1108, and an AI application 1110), results and interpretations 1112, suggestions 1114, and correction actions 1116. Meta-model 1104 is the same as or similar to meta-model 1010, which is illustrated and described in more detail with reference to Figure 10. Meta-model 1104 and AI application 1110 are implemented using components of example device 200 and computing device 302, which are illustrated and described in more detail with reference to Figures 2 and 3, respectively. Similarly, embodiments of example environment 1100 may include different and / or additional components, or may be connected in different ways.

[0141] Training dataset 1102, containing data used to train machine learning models, is input into meta-model 1104. Meta-model 1104 is a composite model encompassing several sub-models customized to address specific aspects of AI compliance. Within meta-model 1104 are various specialized models, such as a bias model 1106A (described in further detail with reference to Figure 5), a toxicity model 1106B (described in further detail with reference to Figure 6), an IP infringement model 1106C (described in further detail with reference to Figure 7), and other verification models 1106D. Each model is responsible for detecting and evaluating specific types of non-compliant content within the AI ​​model. When processing training dataset 1102, each model generates verification actions customized to assess the presence or absence of specific types of non-compliant content. Further evaluation techniques using verification actions generated by meta-model 1104 are discussed with reference to Figures 12 through 14.

[0142] The generated verification action 1108 is provided as input to an AI application 1110 in the form of a prompt. The AI ​​application 1110 processes the verification action 1108 and generates a result and an explanation 1112 detailing how the result was determined. Subsequently, based on the result and explanation 1112 provided by the AI ​​application 1110, the system can generate suggestions 1114 for corrective actions. These suggestions are derived from the analysis of the verification action results and are intended to address any identified problems or deficiencies. For example, if certain verification actions fail to meet desired criteria due to specific attribute values ​​or patterns, the suggestions may recommend adjustments to those attributes or modifications to the basic procedures.

[0143] For a bias detection model (such as the ML model discussed in Figure 5), if certain attributes exhibit unexpected correlations or distributions, the system can retrain the tested AI model with a revised weighting scheme to better align with the desired vector constraints. In a toxicity model (such as the ML model discussed in Figure 6), corrective actions may include implementing post-processing techniques in the tested AI model to filter out responses that violate vector constraints (e.g., filtering out responses containing recognized vector representations of alphanumeric characters). Similarly, in an IP infringement model (such as the ML model discussed in Figure 7), corrective actions may include implementing post-processing techniques in the tested AI model to filter out IP infringement responses (e.g., filtering out responses containing predetermined alphanumeric characters).

[0144] In some implementations, based on the results and interpretations, the system applies predefined rules or logic to determine appropriate corrective actions. These rules can be established by the user and can take into account factors such as regulatory compliance, risk assessment, and corporate objectives. For example, if an application is rejected due to insufficient revenue, the system may suggest that the applicant request additional financial documentation.

[0145] In some implementations, the system can use machine learning models to generate recommendations. The model learns from historical data and past decisions to identify patterns and trends that indicate a series of actions the AI ​​model can take to comply with vector constraints. By training on a dataset of past corrective actions and results, the machine learning model can predict the most effective recommendations for new cases. Further discussion of artificial intelligence and training methods is illustrated in Figure 7. Recommendation 1114 can be automatically implemented by the system as corrective action 1116. This automation simplifies the process of resolving identified problems and ensures rapid remediation of non-compliant content within the AI ​​model, thereby enhancing overall compliance and reliability. The platform uses data to identify and address gaps in AI use cases.

[0146] Figure 12 is a block diagram illustrating one instance environment 1200 for identifying and correcting compliance gaps in AI use cases using a generative AI model, according to some embodiments of the present technology. Environment 1200 includes a computing device 1202, operational data 1204, an inventory module 1206, model use cases 1208, guidelines 1210, a storage module 1212, a compliance engine 1214, models 1216a, 1216b, and 1216c, a risk category 1218, compliance documents 1220, and a feedback loop 1222. The compliance engine 1214 is the same as or similar to the generative model engine 120 illustrated and described in more detail with reference to Figure 1. Similarly, embodiments of instance environment 1200 may include different and / or additional components, or may be connected in different ways.

[0147] Computing device 1202 refers to any electronic device capable of processing data and executing instructions, such as a server, desktop computer, or mobile device. Computing device 1202 and compliance engine 1214 are implemented using components of example device 200 and computing device 302, which are illustrated and described in more detail with reference to Figures 2 and 3, respectively. A computing device 1202 can process operational data 1204, which includes real-time and historical data related to AI use cases, such as performance metrics, user interactions, and system logs. Operational data 1204 covers data generated and used by AI systems within the organization. The data is collected from various sources and used to monitor the performance and behavior of AI models. For example, operational data 1204 may include user interaction logs from a chatbot AI, performance metrics from a predictive maintenance AI, and / or system logs from an image recognition AI.

[0148] In some implementations, operational data 1204 is collected in real time from various endpoints and sensors. For example, user interaction logs from a chatbot AI can be retrieved via API calls and stored in a centralized database (e.g., inventory module 1206). Similarly, system logs from an image recognition AI can be collected from server logs that record each instance of image processing and recognition events. In some implementations, operational data 1204 is accompanied by additional contextual information before being sent to inventory module 1206. For example, operational data 1204 may be accompanied by background information (such as timestamps, user IDs, and location data) to provide background content on the scope of the AI ​​application's influence. For example, user interaction logs from a chatbot AI may use user demographics or geographic coordinates of the deployment area, or system logs from an image recognition AI may use annotations of the image type being processed.

[0149] Operational data 1204 is sent from computing device 1202 to inventory module 1206. Inventory module 1206 maintains a list of AI use cases (e.g., model use case 1208) within the organization. Inventory module 1206 categorizes and stores information about each model use case 1208, including its purpose, functionality, and associated operational data 1204. For example, the inventory module may list an AI use case for fraud detection in financial transactions, detailing the algorithm used, data inputs (e.g., transaction records), and expected outputs (e.g., fraud alerts). In some implementations, operational data 1204 may include documents (e.g., model use case 1208) associated with a type, category, or degree of risk for a particular organizational system. For example, the documents may provide a description of the AI ​​model, its intended use, performance metrics, validation results, and / or risk assessment. In some implementations, the documents may further include historical performance data, such as past event reports for a particular organizational system. Alternatively or concurrently, operational data 1204 may include documents relating to one type, category, or extent of risk associated with new or unimplemented organizational systems, activities, and / or measures. For example, documents may be prior performance assessments of risks introduced by new systems within the organization (or even considerations of systems used by others, hypothetical systems, or systems not implemented by the organization). By incorporating such documents, compliance engine 1214 can assess the potential impact of new activities on the overall risk profile and / or the risk profile of other existing systems.

[0150] Work data 1204 may further include documents associated with each (or nearly all) of the computer systems used by the organization, and may include internal and / or third-party documents. Documents may contain information related to the organization's computer systems, such as the purpose of the system, data inputs and outputs, software and hardware components, and / or any known associated risks (e.g., data labels following a pre-loaded portion of work data 1204 marked with a perceived risk level, as per past organizational guidelines). Similarly, work data 1204 may include documents associated with third-party computer systems, such as third-party software and / or hardware documents, which may be internal and / or external to the third party. Work data 1204 for third-party systems may include systems used to support the organization's activities. Alternatively, work data 1204 for third-party systems may include systems not used by the organization. In some implementations, third-party system documents may further include third-party audit reports and / or compliance certifications. Including operational data associated with third-party systems 1204 enables organizations to ensure that third-party systems meet the same regulatory and compliance standards as internal systems, thereby mitigating the risk of regulatory leaks attributable to third-party components.

[0151] In some implementations, operational data 1204 may encompass both structured and unstructured data from the organization. Structured data may include organizational information such as performance metrics, user interaction logs, and system logs, which are stored in a database in a structured (e.g., tabular) format. On the other hand, unstructured data includes text, images, audio, and video files that do not follow a predefined structure. Examples include user feedback in the form of text comments, images processed by an image recognition AI, and audio recordings from a voice assistant. The method for extracting the model from operational data 1204, use case 1208, is further discussed in detail with reference to Figure 14.

[0152] Inventory module 1206 organizes operational data 1204 into one or more model use cases 1208 and sends the model use cases 1208 to compliance engine 1214. A model use case 1208 refers to a specific application or case using an AI model. Each model use case 1208 includes at least a portion of the operational data 1204 containing information about the AI ​​system's objectives, data inputs and outputs, and the algorithms employed. For example, a model use case 1208 for a suggestion system might include an objective to increase user engagement, data inputs (such as user behavior data), and outputs (such as personalized content suggestions). The information can be used to meet regulatory requirements applicable to each use case and to assess compliance using the methods discussed with reference to Figure 14. In some implementations, model use cases 1208 may include post-processing data describing the model's training data, validation methods, and performance metrics. For example, the post-processing data might include details about the sources of the training data, the preprocessing steps applied, the algorithms and hyperparameters used, and / or the methods used to validate the model's accuracy and robustness. Performance metrics such as accuracy, recall, and / or F1 score can be included in Model Use Case 1208 to provide a quantitative assessment of model effectiveness and / or to evaluate model compliance.

[0153] Inventory module 1206, referencing guideline 1210, generates an inventory of model use cases containing model use cases 1208 using the method described in reference figure 14. In some implementations, the inventory of model use cases includes model use cases 1208 defined by the organization as corresponding operational data 1204 associated with an AI system / application. Using organization-specific definitions to filter operational data 1204 enables the compliance engine 1214 to identify, for example, computer systems identified by the organization as AI. Alternatively, the inventory of model use cases may be included in a general inventory of computer systems used by the enterprise, which may further encompass internal systems and third-party systems operating internally and / or externally. The general inventory ensures that all operational data 1204 (e.g., organizational systems) can still be evaluated against regulatory definitions, even if they were not initially identified as AI operational data, because some operational data 1204 may meet the extended legal or regulatory definitions of AI, even if operational data 1204 was not initially classified as such by the enterprise. In some implementations, two sets are constructed (complete inventory and inventory filtered by organization definition), and a user can compare compliance documents 1220 from the compliance engine 1214 from the two sets of inventory.

[0154] Guidance 1210 (e.g., Guidance 602, Guidance 902) may contain regulatory standards and / or best practices that organizations must adhere to when deploying AI systems. Guidance 1210 may be derived from various regulatory agencies and industry standards. For example, Guidance 1210 may include transparency and accountability requirements under the EU AI Act / California SB-1047 and / or GDPR data protection principles. In another instance, Guidance 1210 may include a Software Development Lifecycle Control (SDLC) document. An SDLC document may cover one or more requirements / guidelines of the software development lifecycle (e.g., requirements for collection, coding, testing, debugging, deployment, and monitoring). Including an SDLC document in Guidance 1210 ensures that the development and maintenance of software (including AI systems) comply with the best practices and regulatory standards outlined in the SDLC document. In some implementations, the SDLC document may include automated test instruction files and / or a continuous integration / continuous deployment (CI / CD) pipeline.

[0155] In some implementations, guidance 1210 may be stored in storage module 1212. Storage module 1212 serves as a repository for storing all relevant data, guidance, and compliance documents. Storage module 1212 ensures that information related to the guidance is stored and easily accessible for compliance assessments and audits. For example, storage module 1212 may store compliance reports, audit logs, and / or regulatory guidance. In some implementations, storage module 1212 may utilize cloud storage solutions to provide scalable and secure data storage options.

[0156] Compliance Engine 1214 uses the criteria within Guideline 1210 to assess the compliance of Model Use Case 1208 and identify any gaps that need to be addressed using the methods discussed in Figure 14. Compliance Engine 1214 manages the assessment of Model Use Case 1208 using the criteria from Regulatory Guideline 1210. In some implementations, Compliance Engine 1214 uses one or more generative AI models (e.g., models 1216a, 1216b, 1216c) to analyze operational data, model use cases, and guidelines to identify compliance gaps. For example, Compliance Engine 1214 may analyze operational data from a facial recognition AI to ensure model use cases comply with privacy regulations, identify any non-compliant data handling practices, and generate actionable recommendations to address gaps. In some implementations, Compliance Engine 1214 may use machine learning algorithms to continuously improve its compliance assessment capabilities based on feedback from the feedback loop 1222 and new regulatory updates.

[0157] Within the compliance engine 1214, models 1216a, 1216b, and 1216c are used to determine risk category 1218. Models 1216a, 1216b, and 1216c can be the same AI model applied to different forms of compliance assessment or different AI models specifically designed for various compliance tasks. For example, model 1216a may focus on data quality and bias detection, model 1216b on algorithmic transparency and interpretability, and model 1216c on performance measurement and robustness. Models 1216a, 1216b, and 1216c can be implemented using various machine learning and deep learning architectures. For example, model 1216a can use a combination of statistical analysis and machine learning classifiers to detect bias in data. Model 1216b can use natural language processing (NLP) techniques to analyze text interpretation and ensure transparency, and model 1216c can use neural networks or holistic methods to evaluate the performance and robustness of the AI ​​model. Compliance is assessed based on the specific use case of each model use case 1208 and applicable regulatory guidance(s). For example, model use case 1208 could be an NLP model for sentiment analysis, a machine learning model for predictive analytics, or a deep learning model for image classification. Model use case 1208 may include contextual details such as user, location, etc. In some implementations, models 1216a, 1216b, and 1216c may be deployed on different hardware platforms (such as GPUs or TPUs) to improve performance for specific tasks.

[0158] Models 1216a, 1216b, and 1216c can be used to determine risk category 1218. Risk category 1218 for model use case 1208 is based on different risk levels of the guidance. Risk categories can be predefined levels, such as low, medium, and high risk, or more granular categories based on specific regulatory guidance. Each risk category can be associated with compliance requirements and mitigation strategies (e.g., the EU AI Act, California SB-1047). For example, the EU AI Act defines "AI systems" as categories such as "prohibited," "high-level," "limited," and "minimal" risk. Compliance Engine 1214 uses risk categories to classify model use case 1208 under specific regulatory categories. For example, a model use case 1208 associated with a higher level of risk (such as autonomous driving) could be classified as a "high-level" risk, which could be associated with stricter compliance checks and monitoring under a specific regulation. Similarly, California SB-1047 focuses on mitigating "critical hazards." Critical hazards are defined as serious risks that could result in mass casualties, significant property damage, or other serious threats to public safety and security. In this example, risk category 1218 could be either "critical hazard" or "non-critical hazard." In some implementations, compliance engine 1214 may use a scoring system to quantify risk category 1218, referring to model use case 1208 discussed in Figure 14.

[0159] Compliance document 1220 can be derived from the guidelines of risk category 1218. Compliance document 1220 contains reports, records, and / or evidence of compliance with model and regulatory guidelines. Compliance document 1220 can be generated by compliance engine 1214 and / or stored in storage module 1212. For example, compliance document 1220 may contain audit reports, compliance checklists, and evidence of corrective actions taken to address identified gaps. In some implementations, the method described with reference to Figure 14 can be used to automatically generate and update compliance document 1220 based on real-time data from the compliance engine.

[0160] Feedback loop 1222 is a mechanism for continuously monitoring and improving compliance. Feedback loop 1222 collects feedback from various sources (such as user interaction, system performance, and regulatory updates) and uses this information to refine compliance procedures. For example, feedback loop 1222 can incorporate user updates for a specific model use case, changes in system performance metrics, and updates to regulatory guidelines. In some implementations, feedback loop 1222 can be used to tune models 1216a, 1216b, and 1216c using the methods described with reference to Figure 14. For example, hyperparameters in a neural network (such as learning rate, regularization parameters, and / or number of layers) can be adjusted to enhance model performance. Additionally, algorithm selection can be performed by evaluating and selecting different algorithms based on feedback data. For example, if a random forest demonstration shows improved performance, compliance engine 1214 can switch from using a decision tree to using a random forest. In addition, the compliance engine 1214 can add new input features derived from user feedback or remove features that do not contribute to model performance.

[0161] Figure 13 is a block diagram illustrating one instance environment 1300 for continuously monitoring compliance in an AI use case using a generative AI model, according to some embodiments of the present technology. Environment 1300 includes model use case 1302, compliance engine 1304, model 1306, risk category 1308, criteria 1310, gaps 1312, compliance actions 1314, and monitoring loop 1316. Model use case 1302 is the same as or similar to model use case 1208, which is illustrated and described in more detail with reference to Figure 12. Model 1306 is the same as or similar to models 1206a, 1206b, and 1206c, which are illustrated and described in more detail with reference to Figure 12. Model 1306 is implemented using components of instance device 200 and computing device 302, which are illustrated and described in more detail with reference to Figures 2 and 3, respectively. Risk category 1308 is the same as or similar to risk category 1218, which is illustrated and described in more detail with reference to Figure 12. Similarly, the implementation of instance environment 1300 may include different and / or additional components, or may be connected in different ways.

[0162] Model use case 1302 is a specific application or case that utilizes or can utilize an AI model (e.g., model use case 1208), and defines the boundaries and expectations of AI model operation, including inputs, outputs, and background content of the applied (or to be / applied) model. Model 1306 within compliance engine 1304 represents various AI models employed within environment 1300. Model 1306 processes operational data associated with model use case 1302 (e.g., operational data 1204) and generates compliance indicators based on the defined model use case 1302. Model 1306 may contain various types of machine learning algorithms, such as supervised learning models, unsupervised learning models, reinforcement learning models, and deep learning models. Each model 1306 can be trained on a specific dataset to perform tasks such as classification, regression, clustering, and anomaly detection. Model 1306 can be continuously updated and retrained with new data to improve its accuracy and performance.

[0163] Model 1306 can be domain-specific or general. A domain-specific model 1306 is tailored to address specific regulations within a specific domain of a regulation. For example, a domain-specific model for cybersecurity compliance can be trained to ensure compliance with regulations such as the EU AI Act, focusing on cybersecurity requirements. A domain-specific model 1306 can be trained on a dataset containing regulatory texts, compliance checklists, and historical compliance data related to a specific regulation or a specific domain of the regulation. For example, a domain-specific model for GDPR cybersecurity compliance would focus on data encryption, access controls, and breach notification requirements. A domain-specific model 1306 can output risk categories and identify domain-specific compliance gaps using the methods discussed in Figure 14. On the other hand, a general model 1306 can output broader risk categories and identify general compliance gaps applicable across various regulatory frameworks using the methods discussed in Figure 14.

[0164] Risk category 1308 categorizes potential risks associated with model use case 1302. Criteria 1310 for each risk category in risk category 1308 (e.g., "Unacceptable Risk" 1308a, "High Risk" 1308b, "Limited Risk" 1308c, and "Minimum Risk" 1308d) provide the standards that model use case 1302 must meet, and any gap 1312 between criteria 1310 and model performance can be identified by a system (e.g., compliance engine 1304, compliance engine 1214 in Figure 12). Risk categories may include risks across various dimensions of regulations, such as data privacy risks, model bias risks, and operational risks, depending on the guidelines used (e.g., guideline 1210 in Figure 12). Each risk category 1308 may be associated with specific risk factors and metrics used by model 1306 to assess the risk level of model use case 1302. For example, data privacy risks can be assessed based on the sensitivity of the data used by the model, while model bias risks can be assessed based on the fairness and impartiality of the model's predictions.

[0165] Guideline 1310 sets forth the specific requirements or standards that model use case 1302 must meet. Guideline 1310 ensures that model use case 1302 operates within acceptable parameters and complies with regulatory, ethical, and performance standards. Guideline 1310 may include various technical and non-technical requirements, such as accuracy thresholds, data quality standards, ethical guidelines, and regulatory compliance requirements. For example, accuracy thresholds may specify the minimum acceptable accuracy of the model's predictions, while data quality standards may define the required quality and completeness of the data used by the model. Further refer to Figure 14 for a method of evaluating model use case 1302 in accordance with Guideline 1310.

[0166] Gap 1312 refers to the difference or deviation between the expected criteria and the actual performance or characteristics of model use case 1302, and may include any type of defect, missing obligation, or missing requirement. Gap 1312 can be identified using the various methods described with reference to Figure 14. Compliance action 1314 is the steps taken to address the identified gap and ensure that model use case 1302 meets the defined criteria. Compliance action 1314 may include modifications to model use case 1302, updates to data, and / or changes to operating procedures. In some implementations, compliance action 1314 may include various technical and non-technical measures, such as retraining the model with new data, updating the model's algorithm, implementing data privacy controls, and conducting periodic audits and reviews. For example, retraining one of the models in model use case 1302 with new data can help improve its accuracy and reduce bias, while implementing data privacy controls can help protect the privacy and confidentiality of data used by the model. In some implementations, compliance actions can be performed automatically using the methods described with reference to Figure 14.

[0167] Monitoring loop 1316 continuously observes model use case 1302, updating and refining it as needed to maintain compliance over time. Monitoring loop 1316 is a continuous process of observing, evaluating, and updating the AI ​​model and its use cases, and can be the same as or similar to feedback loop 1222 in Figure 12. Monitoring loop 1316 ensures the model maintains compliance over time, adapting to new data, changing conditions, and evolving standards. For example, continuous performance monitoring using monitoring loop 1316 may involve real-time tracking of performance metrics for model use case 1302 and identifying any deviations from expected criteria, or periodically evaluating the data, algorithms, and procedures of model use case 1302 at predefined time intervals to ensure compliance with defined criteria. In some implementations, the system may continuously scan for regulatory changes or receive regulatory updates from a user input. When a new regulation or regulatory update is identified, the system automatically updates the compliance criteria using the method described in Figure 14 and maps the new criteria to a specific AI use case, thereby updating the existing compliance rules and parameters.

[0168] Figure 14 is a flowchart illustrating one of the procedures 1400 for identifying and correcting compliance gaps in AI use cases using a generative AI model according to some embodiments of the present technology. In some embodiments, procedure 1400 is performed by components of instance device 200 and computing device 302, which are illustrated and described in more detail with reference to Figures 2 and 3, respectively. Specific entities, such as groups of AI models, are illustrated and described in more detail with reference to Figures 12 and 13. Similarly, embodiments may include different and / or additional steps or may perform steps in a different order.

[0169] In operation 1402, the system (e.g., compliance engine 1214 in Figure 12, compliance engine 1304 in Figure 13) receives from a computing device: (1) a set of alphanumeric characters (e.g., regulations, guidelines), which define one or more operational boundaries of a set of expected model use cases configured to comply with the constraints of that set of alphanumeric characters; and (2) a set of operational data, which contains one or more of structured or unstructured data. The set of operational data may include data from, for example, third-party tools (e.g., downstream suppliers).

[0170] For example, the anticipated model use cases could be "AI systems" as defined in regulations such as the EU AI Act, California SB-1047, and / or Guideline 1210 in Figure 12. Regulatory alphanumeric characters represent specific regulatory requirements or guidelines for the AI ​​model in the model use case. By receiving input, the compliance engine can evaluate the AI ​​model against defined constraints to ensure that the model use case complies with regulatory standards. Operational data (which may be structured (e.g., databases) or unstructured (e.g., text, images)) provides context for evaluating the model's performance and compliance. For operational structured data, the system can convert the data type and remove missing values; for unstructured data, NLP can be used for textual data or image preprocessing for visual data. The compliance engine can use the methods discussed in Figure 8 to convert alphanumeric characters representing regulatory guidelines into actionable rules and constraints, and then map the actionable rules to portions of the model use case. For example, the EU AI Act may require documentation of data control practices that can be mapped to specific data handling and storage requirements within the model's use case.

[0171] This set of anticipated model use cases may include a common set of attributes across all anticipated model use cases within this set. This set of attributes may include characteristics defined as "AI systems" in regulations, such as those outlined in the EU AI Act or the "covered model" in California SB-1047. For example, for a regulation that defines an "AI system" as "a system designed to operate with a degree of autonomy and produce outputs (such as content, predictions, recommendations, or decisions that affect real or virtual environments)," this set of attributes may include autonomy, output generation, data control, algorithmic transparency, risk management, human oversight, ethical considerations, and / or compliance documentation.

[0172] In operation 1404, using the common set of attributes among the expected model use cases, the system constructs a set of observed model use cases (e.g., an inventory) from the set of operational data using a first set of AI models. Each specific observed model use case in this set of observed model use cases may contain a set of features for that specific observed model use case. For example, the set of features for a specific observed model use case may include a text-based description of that specific observed model use case, the expected input of that specific observed model use case, the expected output of that specific observed model use case, one or more AI models configured to generate the expected output of that specific observed model use case using the expected input of that specific observed model use case, and / or data supporting one or more AI models.

[0173] To obtain the set of observed model use cases from operational data, the system can examine a set of existing (e.g., internal) explicit and implicit definitions of AI, machine learning, and / or generative AI, and compare the definitions with: (1) definitions provided by a specific regulatory framework (such as the EU AI Act and California SB-1047), and / or (2) other regulations lacking a definition, such as GDPR, NIST, and OECD. In some implementations, the system can identify definition gaps and generate recommendations (or automated actions) to close the definition gaps by using the methods discussed in Figures 5 through 8 to align organizational definitions with regulatory definitions.

[0174] In some implementations, the implementation may receive a structured AI use case input and use a rule-based system and / or one or more AI models to determine the applicable predefined regulations for the AI ​​use case. Specific regulations can be selected based on the background content and requirements of the AI ​​use case being evaluated. After analyzing and associating regulations with those stored in its database, the system labels, keywords, or categories the data and uses NLP to parse the text of the regulations, identifying terms and clauses relevant to the AI ​​application domain. Regulations can be stored in a vector space, allowing the system to calculate the similarity between vectors representing domain-specific terms and other relevant terms, applying a similarity threshold to filter out insufficiently similar terms. Additionally, the system tags relevant regulations with attributes that categorize regulations based on criteria such as regulatory subject matter or jurisdictional relevance.

[0175] When processing unstructured data (such as text documents, emails, or social media posts), the system can extract relevant operational data that matches the set of attributes (such as autonomy or output generation). For example, data related to autonomy can be identified by searching for phrases indicating independent action, while output generation data can be extracted by identifying mentions of content creation, prediction, or suggestion. Once the relevant data is extracted, the operational data can be mapped to regulatory attributes and organized into a structured model use case. Each use case may include a description of the data sources, algorithms, or other characteristics of the models(s) used.

[0176] In some implementations, the system uses a third set of AI models (which may be the same as or different from the first and / or second set of AI models) to identify a portion of the observed model use cases. For example, the third set of AI models may use these attributes to perform a capture-amplify-generate (RAG) search of job data. During training, the third set of AI models learns to capture relevant information from a large dataset based on input attributes and to generate meaningful responses or insights. The capture component of the RAG model searches the job data using techniques such as vector similarity search to find relevant model use cases that match the input attributes (e.g., model type used, location, user type), where the attributes are converted into vector representations and the model searches for similar vectors within the dataset. The generation component uses the captured information to produce a context-relevant output that can be identified as a portion of the observed model use cases.

[0177] For each specific observed model use case within the observed model use case set, in operation 1406, the system uses a second set of AI models to map the set of text characters and features of a specific observed model use case to a risk category defined within that set of text characters. The second set of AI models can select a risk category from multiple risk categories defined within that set of text characters based on a risk level associated with that set of features. For example, the second set of AI models can be trained on a dataset containing regulatory text and features of various model use cases, along with their corresponding risk categories. During training, the second set of AI models learns to identify the pattern and relationship between regulatory requirements (which can be expressed as vector representations) and the features of model use cases (which can be expressed as vector representations).

[0178] In some implementations, text embedding techniques (such as Word2Vec, GloVe, or BERT) can be used to convert the set of alphanumeric characters into vector representations to capture the semantic and contextual meaning of the regulations, while simultaneously transforming them into numerical vectors. In some implementations, similar methods are used to convert features of observed model use cases (such as data quality, algorithmic transparency, and potential biases) into numerical vectors. Once both the regulatory text and model use case features are represented as vectors, the vectors are combined into a unified feature space. A second set of AI models can evaluate the combined vectors based on the learned patterns and relationships within the training data of the second set of AI models, assigning a risk category based on the consistency between the regulatory requirements and the observed model use case features. In some implementations, the second set of AI models uses a distance between a set of words within the set of features of a specific observed model use case and the respective mapped alphanumeric characters within that set of alphanumeric characters to identify a mapping between the set of features of a specific observed model use case and the set of vector representations of those alphanumeric characters. A distance metric can be used to calculate the similarity or dissimilarity between vectors. For example, Euclidean distance measures the straight-line distance between two points in a vector space, cosine similarity measures the cosine of the angle between two vectors, and Manhattan distance measures the sum of the absolute differences between the coordinates of two points. Calculating distances helps identify the mapping between features of a model's use case and regulatory text. Features that are closer in vector space to a specific regulatory requirement indicate stronger consistency. For example, if a feature vector representing "data encryption" is close to a regulatory vector emphasizing "data security," then this indicates strong consistency.

[0179] In some implementations, a second set of AI models (which may be the same as or different from the first set of AI models) can generate a risk score for each observed model use case. This risk score can be used to classify use cases into one of a predefined risk category, such as low, medium, or high risk. For example, a use case with high data quality, transparent algorithms, minimal bias, and low decision-making impact can be classified as low risk. Conversely, a use case with poor data quality, opaque algorithms, significant bias, and high decision-making impact can be classified as high risk. To generate a risk score, the system can use a multi-agent architecture within the second set of AI models, where different agents (i.e., several models) are specifically designed to assess specific risk factors, such as data quality, algorithmic transparency, and potential bias. Each agent independently assesses the assigned factors and generates a subset of risk scores. The subset scores are aggregated to form an overall risk score for one use case. In some implementations, the system uses a holistic approach, such as averaging scores from multiple models. In some implementations, the system uses a simpler model to generate an initial risk score and uses a more complex model to refine the risk score. For example, a basic decision tree can first categorize model use cases into broad risk categories, which are then fine-tuned by a neural network that considers more granular factors (e.g., regulatory factors). In some implementations, a multi-agent architecture is similarly used by a first set of AI models to construct the set of observed model use cases.

[0180] In operation 1408, for each specific observed model use case within the set of observed model use cases, the system uses a second set of AI models (or user input) to identify a set of criteria for a specific observed model use case within the set of alphanumeric characters. The system extracts a set of keywords from the set of alphanumeric characters. The system can map the extracted keywords to the set of criteria associated with the mapped risk category within the set of alphanumeric characters. The system can search the set of alphanumeric characters for sentences or phrases containing the identified keywords and extract relevant information. For example, if the keyword "encryption" is identified, the system searches for sentences mentioning "encryption" and extracts criteria related to data encryption. In addition, using the identified criteria, the system can compare different model use cases to determine whether certain model use cases pose a lower risk than others.

[0181] After extracting the criteria, the system can associate and map them to specific risk categories. The system can use one or more models trained on labeled datasets, where each criterion is labeled with its corresponding risk category. Such models, which may include logistic regression, SVM, or neural networks, can be trained to identify the patterns and relationships between criteria and risk categories. During the training phase, the model learns to identify which criteria indicate a specific risk level. For example, the model learns that "data encryption," "multi-factor authentication," and "periodic audits" indicate a "high-risk" category, while "basic password protection" and "firewall implementation" are associated with "medium risk," and "antivirus software" is linked to "low risk." Once trained, the model can be deployed to automatically classify new criteria into appropriate risk categories.

[0182] In Operation 1410, for each specific observed model use case within the set of observed model use cases, the system uses a second set of AI models to identify a set of gaps in a specific observed model use case by comparing that set of criteria with that set of features. A gap is identified when a feature does not fully meet the criteria, thus indicating a potential compliance or performance issue. For example, if the criteria require "data encryption" and "multi-factor authentication," but the operational data within the model use case only includes "basic password protection," the second set of AI models can mark this as a gap. Refer to Figures 6 through 8 to discuss the method for identifying a gap.

[0183] In some implementations, the system determines the type of gap between the set of criteria and the set of features of a specific observed model use case for each gap in the set of gaps. The type of gap may be associated with one or more of the following: the presence of a specific set of literal characters in a specific observed model use case; or the semantic meaning of one or more sets of literal characters in a specific observed model use case. The system can classify the gaps in the set of gaps using the difference type. In some implementations, the system triggers one or more alerts in response to the classification of a specific gap in the set of gaps reaching a predetermined threshold.

[0184] In operation 1412, using the identified gaps, the system generates a set of actions to be performed related to a specific observed model use case, which are configured to satisfy the set of criteria for that specific observed model use case. In some implementations, the system uses a third set of AI models to verify that the set of actions satisfies the set of criteria for that specific observed model use case.

[0185] In some implementations, the system uses the generated actions to automatically trigger one of the automated workflows indicated by the generated actions to perform the generated actions. In some implementations, the system converts existing documents (e.g., previously compiled documents) into new documents that meet new regulatory requirements. For example, financial institutions are expected to comply with both MRM requirements and EU AI Act obligations, which differ in their risk rating methodologies. The MRM takes a holistic view of risk classification (e.g., an AI related to a loan does not necessarily guarantee a high-risk rating), while the EU AI Act uses a use-case-based classification method (e.g., all AI related to a loan is high-risk and has corresponding obligations, even if there may be exceptions). The system can combine textual analysis of existing documents with relevant regulations to classify risk categories (disregarding pre-assigned risk ratings provided by other regulations). The system then uses the existing documents to generate new documents that comply with the updated regulatory requirements.

[0186] In some implementations, when generating and / or executing the set of actions, the system enables a set of agents to operate autonomously and make decisions and / or execute the set of actions based on programming, learned behavior, and / or suggestions from other models (e.g., AI models, LLM, GenAI). For example, the agent uses LLM to query suggestions about what actions to trigger and extracts parameters for those actions. For example, a user can prompt the agent with a command that triggers actions to automatically adjust a set of operational data to comply with a set of criteria for a regulatory category (e.g., a risk category) by correcting any identified gaps. Actions can be triggered by the user and based on any operational data received from a first set of AI models and / or a second set of AI models.

[0187] To verify and control agent actions to ensure that the agent does not take any unnecessary or unauthorized actions, the system can intercept and inspect actions performed by an autonomous agent before execution and / or audit and record actions after execution. For example, the system can compare each anticipated action with a set of predefined rules and boundaries established by a human operator (e.g., through a set of prompts, pre-loaded query background content). If an action falls outside the boundaries, it can be marked as unauthorized, and the system prevents its execution. In some implementations, the system verifies that all actions taken by the agent comply with industry-specific standards, laws, regulations, and ethical guidelines. The system can cross-reference actions with a database containing permitted actions for the agent (e.g., a static database, a dynamic database). In some implementations, the system can use one or more LLMs (or a voting mechanism among multiple LLMs) to ensure that the agent's behavior is fair, impartial, and / or consistent with ethical standards. For example, the system can query multiple LLMs to evaluate proposed actions rather than relying on a single algorithm. If a majority of LLMs agree that the action is appropriate and unbiased, the action can proceed. If a discrepancy exists or a potential deviation is detected, the unit can automatically adjust its actions or flag the actions for human review. The system can generate reports summarizing the agent's compliance status, any unauthorized actions prevented, and any deviation corrections made.

[0188] In some implementations, for each specific observed model use case within the set of observed model use cases, the system maps each specific observed model use case to a timestamp of one of the set of alphanumeric characters and the time when the specific observed model use case was constructed (e.g., for auditing purposes). Using a third set of AI models, the system can aggregate a set of documents associated with a specific observed model use case. This set of documents meets the set of criteria corresponding to the mapped risk category. In some implementations, this set of documents includes mapped timestamps.

[0189] In some implementations, a system version (e.g., assigning a version) and / or a timestamp is attached to one or more of the categories of observed model use cases, the set of alphanumeric characters, operational data, etc. For example, each version of the classification (such as version A) may be dated (e.g., version A, date 09 / 2030) and include a detailed classification of one of the observed model use cases. This classification may be accompanied by an explanation of a specific subset of the set of alphanumeric characters used to classify the observed model use cases (thus enabling users to track the system's decision-making process). In some implementations, the explanation may include basic instructions (e.g., queries, prompts, assumptions, operational data, post-operational data, preloaded query context) used by the models (e.g., a first set of AI models, a second set of AI models, and a third set of AI models). In some implementations, the explanation defines the set of parameters used by at least one or more of the first or second set of AI models to generate one or more of the following: the risk category of a particular observed model use case, the set of criteria, or the set of gaps.

[0190] For example, if an observed model use case is classified under a high-risk category, the system may record one or more of the following: (i) the classification (e.g., high-risk), (ii) an interpretation of the classification (e.g., the definition of "high-risk" in regulations, specific rules and criteria within regulations that lead to the classification), or (iii) the underlying assumptions used by one or more models to classify the observed model use case (e.g., the interpretation of regulations as of a specific date). For example, the document may contain prompts used by (several) AI models to interpret regulatory text and map compliance criteria to the observed model use case. By maintaining a versioned record of the classification, the system ensures that decisions are traceable and can be audited at any time to verify that the system's classifications are consistent with the latest regulatory requirements and standards.

[0191] The system can build and store a set of versioned files for each observed model use case. These files may contain mapped timestamps and / or versioned categories. The files can be aggregated and stored in a centralized repository. For example, the set of files may contain reports on compliance criteria, risk categories, and / or specific actions taken to address any identified compliance gaps. As discussed above, each file in this set may be accompanied by a version number and / or timestamp.

[0192] For each specific observed model use case in the set of observed model use cases, the system can generate a layout that indicates one of the summary files in the set for display on a computing device. For example, the layout includes a first representation of one of the specific observed model use case model use cases and a second representation of one of the corresponding files in the set of summary files.

[0193] In some implementations, the system monitors each specific observed model use case within a set of criteria for a particular observed model use case at predetermined time intervals. In response to detecting a set of deviations, the system can automatically update the set of gaps. The system can use this updated set of gaps to generate a set of actions for observed model use cases configured to meet the set of criteria for the observed model use cases. In some implementations, using this set of generated actions, the system automatically triggers an automated workflow indicated by the generated actions to perform those actions. The system can track compliance with a specific regulation both externally and internally. For example, externally, the system monitors required statements or documentation issued by AI providers as mandated by legislation. Internally, the system can automatically reassess and identify new gaps, and generate an alert based on these gaps (e.g., an anomalous trigger for potentially "prohibited" or "high-risk" AI use cases).

[0194] In some implementations, the set of attributes is a set of expected attributes, and the set of actions is a first set of actions. The system can compare the common expected attributes across the various expected model use cases in the set of expected model use cases with a set of observed attributes defining one or more of the observed model use cases in the set of observed model use cases. In response to the absence of one or more attributes from the set of expected attributes in the set of observed attributes, the system can generate a configuration to add the missing one or more attributes from the set of expected attributes to a second set of actions in the set of observed attributes, using the set of expected attributes.

[0195] In some implementations, this set of gaps is a first set of gaps. The system can map each specific observed model use case of the set of observed model use cases to a timestamp of one of the set of alphanumeric characters and the time of construction of the specific observed model use case. Using a third set of AI models, the system can aggregate a first set of documents associated with a specific observed model use case, wherein the first set of documents meets the criteria of that set corresponding to the mapped risk category. The system can provide a second set of documents associated with a specific observed model use case. The system can identify a second set of gaps by comparing the first set of documents with the second set of documents. For example, the second set of gaps may include documents that are present in the first set of documents but missing in the second set of documents, and / or documents that are present in the second set of documents but missing in the first set of documents.

[0196] The system can automatically trigger a set of corrective actions in response to the identification of a second set of gaps. These corrective actions may include adding missing files from the first set of files to the second set of files, modifying one or more files in the second set of files, removing one or more files from the second set of files, requesting a set of additional features associated with a specific model use case via a computing device, modifying that set of features for the specific model use case, and / or adjusting at least one parameter of one or more AI models from the first and / or second sets of AI models. in conclusion

[0197] Unless otherwise expressly required by the background context, throughout the description and scope of the invention application, the terms "comprise" and similar terms shall be interpreted as inclusive rather than exclusive or exhaustive—that is, "including but not limited to." As used herein, the terms "connection," "coupling," and any variations thereof mean any direct or indirect connection or coupling between two or more elements; the coupling or connection between elements may be physical, logical, or a combination thereof. Furthermore, the terms "in this document," "above," "below," and similar terms, when used in this application, refer to the entire application and not any particular part thereof. Where the background context permits, the use of singular or plural terms in the above embodiments may also include both singular and plural forms. Referring to a list of two or more items, the term "or" encompasses all of the following interpretations: any one of the items in the list, all of the items in the list, and any combination of the items in the list.

[0198] The implementation of this technology described above is not intended to be exhaustive or to limit the technology to the precise form disclosed above. While specific implementations of the technology have been described above for illustrative purposes, various equivalent modifications are possible within the scope of this technology, as will be recognized by those skilled in the art. For example, although programs or blocks are presented in a given order, alternative implementations may execute routines with steps or employ systems with blocks in a different order, and some programs or blocks may be deleted, moved, added, subdivided, combined, and / or modified to provide alternative examples or sub-combinations. Various programs or blocks can be implemented in different ways. Furthermore, although programs or blocks are sometimes shown to execute sequentially, such programs or blocks may alternatively be executed in parallel or may be executed or implemented at different times. Moreover, any specific numbers mentioned herein are merely examples; alternative implementations may employ different values ​​or ranges.

[0199] The teachings of this technology provided herein are applicable to other systems, not necessarily those described above. Elements and actions of the various examples described above can be combined to provide further embodiments of this technology. Some alternative embodiments of this technology may include additional elements of those embodiments mentioned above, or may include fewer elements.

[0200] In view of the above embodiments, such and other changes may be made to this technology. Although the above description describes certain instances of this technology and the best mode carefully considered, this technology can be practiced in many ways, however detailed the description may be. The details of the system vary significantly in its particular embodiments, while still being covered by the technology disclosed herein. As mentioned above, specific terms used when describing certain features or manners of this technology should not be construed as implying that the terms are redefined herein as limited to any particular characteristic, feature, or manner of the technology associated with that term. Generally speaking, unless such terms are explicitly defined in the above embodiments section, the terms used in the following claims should not be construed as limiting the technology to the specific instances disclosed herein. Therefore, the actual scope of this technology covers not only the disclosed instances but also all equivalent ways of practicing or implementing the technology within the claims.

[0201] To reduce the number of claims, certain aspects of the present invention are presented hereinafter as certain claims, but the applicant has carefully considered the various aspects of the present invention in any number of claims. For example, while only one aspect of the present invention is described as a computer-readable medium claim, other aspects may similarly be embodied as a computer-readable medium claim or in other forms, such as embodied in a component plus function claim. Any claim intended to be processed under 35 USC § 112(f) will begin with the phrase “for a component of…”, but the use of the term “for” in any other context is not intended to invoke processing under 35 USC § 112(f). Therefore, the applicant reserves the right to seek such additional forms of claims in this application or a subsequent application after the filing of this application.

[0202] As should be understood from the foregoing, specific embodiments of the invention have been described herein for illustrative purposes, but various modifications may be made without departing from the scope of the invention. Therefore, the invention is not limited, except as provided in the appended patent applications.

[0203] 100: Environment 102: Data Generation Platform 104: Data Node 108a to 108n: Third-party database 112: Communication Engine 114: Access Control Engine 116: Engine leakage mitigation 118: Performance Engine 120: Generative Model Engine 150: Internet 200: Device 204: Input Component 206: Output Components 208: Processor 210: Storage 212: Application 214: Model 216: Network connectivity components 218: Permanent storage device 220: Computer-readable media disk drive 300: Computing Environment 302a to 302d: User-end computing devices 304: Internet 306: Server computing unit 308: Database 310a to 310c: Server computing devices 312a to 312c: Database 400: Artificial Intelligence (AI) Model 402: Machine Learning Model 404: Input 406: Output 500: Environment 502: User 504: Platform 506: Data Provider 508: Artificial Intelligence (AI) Model Agent 510: Large Language Models (LLM) 512: Data Cache Area 514: Notice to store 516: Execute stored log 600: Environment 602: Guidelines 604: Operating Standards 606: Gap 608: Platform 610: Mapping gap 700: Program 702: Action 704: Action 706: Action 708: Action 710: Action 712: Action 800: Environment 802: Guidelines 804: Platform 806a to 806n: Text subsets 808a to 808n: Tips 810a to 810n: Operable Projects 900: Environment 902: Guidelines 904: Jurisdiction 906: Organizational Regulations 908: Specific Regulations for Artificial Intelligence (AI) Applications 910: Vector storage 912: Verification Engine 1000: Environment 1002: Training Data 1004: Enterprise Applications 1006: Other Artificial Intelligence (AI) Applications 1008: Internal Document Search - Artificial Intelligence (AI) 1010: Metamodel 1012: Verification Action 1014: Cache Area 1016: Vector storage 1100: Environment 1102: Training Data Set 1104: Metamodel 1106A: Validation Model / Bias Model 1106B: Validation Model / Toxicity Model 1106C: Validation Model / Intellectual Property (IP) Infringement Model 1106D: Validation Model / Other Validation Models 1108: Verification Action 1110: Artificial Intelligence (AI) Applications 1112: Results and Explanation 1114: Recommendation 1116: Correction Action 1200: Environment 1202: Computing device 1204: Homework Materials 1206: Inventory Module 1208: Model Use Case 1210: Guidelines 1212: Storage Module 1214: Compliance Engine 1216a: Model 1216b: Model 1216c: Model 1218: Risk Category 1220: Compliance Documents 1222: Feedback Loop 1300: Environment 1302: Model Use Case 1304: Compliance Engine 1306: Model 1308: Risk Category 1308a: Unacceptable Risk 1308b: High risk 1308c: Limited Risk 1308d: Minimum Risk 1310: Guidelines 1312: Gap 1314: Compliance Actions 1316: Monitoring Loop 1400: Program 1402: Operation 1404: Operation 1406: Operation 1408: Operation 1410: Operation 1412: Operation

Claims

1. A non-transitory computer-readable storage medium having instructions thereon, wherein when executed by at least one data processor of a system, the system causes the system to: receive from a computing device: (1) a set of alphanumeric characters, defined and configured to adhere to one or more operational boundaries of a set of expected model use cases; and (2) a set of operational data, containing one or more of structured or unstructured data, wherein the set of expected model use cases includes a common set of attributes among the expected model use cases in the set; transmit the common set of attributes among the set of expected model use cases to one or more nodes of an input layer of a first set of AI models to receive a set of observed model use cases from one or more nodes of an output layer of the first set of AI models; Each specific observed model use case in the group of observed model use cases includes a set of features of that specific observed model use case, and the set of features of that specific observed model use case includes two or more of the following: a text-based description of that specific observed model use case, the expected input of that specific observed model use case, the expected output of that specific observed model use case, one or more AI models configured to generate the expected output of that specific observed model use case using the expected input of that specific observed model use case, or data supporting the one or more AI models;The observed model use cases are transmitted to one or more nodes of the input layer of a second set of AI models, which are trained to: map the set of literal characters and the set of features of the specific observed model use case to a risk category defined within a set of vector representations of the set of literal characters, wherein the second set of AI models is configured to select the risk category from a plurality of risk categories defined within the set of vector representations of the set of literal characters based on a risk level associated with the set of features, and to identify the specific observed model use case within the set of literal characters by: (1) extracting a set of keywords from the set of literal characters, and (2) mapping the extracted keywords to the set of criteria associated with the mapped risk category within the set of literal characters, and identifying a set of discrepancies of the specific observed model use case by comparing the set of criteria of the specific observed model use case with the set of features of the specific observed model use case. Using this set of gaps, a set of actions to be performed is generated in relation to the specific observed model use case. These actions are configured such that the set of features of the specific observed model use case satisfies the set of criteria for that specific observed model use case. A representation is presented via the computing device, including one or more graphical user interface components or a set of text, wherein the representation indicates at least one of the gaps or the actions. In response to user input received via the computing device, the actions are automatically executed to modify the set of operational data. Each specific observed model use case in the set is transmitted to one or more nodes in the input layer of the second set of AI models to verify the satisfaction of the set of criteria for each observed model use case.

2. The non-transitory computer-readable storage medium of request item 1, wherein the instructions further cause the system to: automatically trigger an automated workflow indicated by the group of generated actions, wherein the automated workflow includes performing the group of generated actions.

3. As in request 1, a non-transitory computer-readable storage medium, wherein such instructions further cause the system to, for each specific observed model use case of the set of observed model use cases, map each specific observed model use case to a timestamp of one of the set of alphanumeric characters and the time of construction of the specific observed model use case; and use a third set of AI models to aggregate a set of documents associated with the specific observed model use case, wherein the set of documents satisfies the set of criteria corresponding to the mapped risk category, and wherein the set of documents contains the mapped timestamp.

4. The non-transitory computer-readable storage medium as requested in item 3, wherein the instructions further cause the system to: generate instructions for displaying one layout of the set of aggregated files on the computing device for each specific observed model use case of the set of observed model use cases, wherein the layout includes a first representation of one model use case of the specific observed model use case and a second representation of one of the corresponding files in the set of aggregated files.

5. The non-transitory computer-readable storage medium of claim 1, wherein the second set of AI models is configured to identify a mapping between the set of features of the specific observed model use case and the set of vector representations of the set of alphanumeric characters by using a distance between a set of words in the set of features of the specific observed model use case and the respective mapped alphanumeric characters in the set of alphanumeric characters.

6. For the non-transitory computer-readable storage medium of Request 1, determine, for each gap in the set of gaps, a type of gap between the set of criteria for that particular observed model use case and the set of features for that particular observed model use case, wherein the type of gap is associated with one or more of the following: the presence of one of a particular set of alphanumeric characters in that particular observed model use case, or the semantic meaning of one or more sets of alphanumeric characters in that particular observed model use case; classify each gap in the set of gaps using a difference type; and trigger one or more alerts in response to the classification of one particular gap in the set of gaps reaching a predetermined threshold.

7. The non-transitory computer-readable storage medium of request item 1, wherein the instructions further cause the system to: monitor each specific observed model use case of the group of observed model use cases for deviations from the set of criteria of the specific observed model use case at predetermined time intervals; automatically update the set of gaps in response to the detection of a set of deviations; generate a set of actions for the observed model use case configured to meet the set of criteria of the observed model use case using the set of updated gaps; and automatically trigger an automated workflow indicated by the set of generated actions using the set of generated actions, wherein the automated workflow includes performing the set of generated actions.

8. A method for identifying and remediating gaps in artificial intelligence use cases using one or more artificial intelligence (AI) models, the method comprising: Receive from a computing device: (1) a set of alphanumeric characters, whose definitions are configured to comply with the constraints of the set of alphanumeric characters, and one or more operational boundaries of a set of expected model use cases; and (2) a set of operational data, wherein the set of expected model use cases includes a common set of attributes among the expected model use cases in the set of expected model use cases; transmit the common set of attributes among the set of expected model use cases to one or more nodes of an input layer of a first set of AI models, so as to receive from one or more nodes of the layer, and from one or more nodes of the output layer of the first set of AI models, a set of observed model use cases from the set of operational data; Each specific observed model use case in the set of observed model use cases includes a set of features for that specific observed model use case, wherein the set of features for that specific observed model use case includes at least one of the following: a text-based description of that specific observed model use case, an expected input of that specific observed model use case, an expected output of that specific observed model use case, one or more AI models configured to generate the expected output of that specific observed model use case using the expected input of that specific observed model use case, or data supporting the one or more AI models; Each specific observed model use case in the set of observed model use cases is transmitted to one or more nodes of the input layer of a second set of AI models, the second set of AI models being trained to: map the set of literal characters and the set of features of the specific observed model use case to a risk category defined within the set of literal characters, wherein the second set of AI models is configured to select the risk category from a plurality of risk categories defined within the set of literal characters based on a risk level associated with the set of features. The following criteria are used to identify a specific observed model use case within the set of alphanumeric characters: A set of keywords is extracted from the set of alphanumeric characters, and the extracted keywords are mapped to the set of criteria associated with the mapped risk category within the set of alphanumeric characters; a set of discrepancies in the specific observed model use case is identified by comparing the set of criteria of the specific observed model use case with the set of features of the specific observed model use case; and a set of actions to be performed related to the specific observed model use case is generated using the identified discrepancies, the set of actions being configured such that the set of features of the specific observed model use case satisfies the set of criteria of the specific observed model use case; and a representation indicating at least one of the discrepancies or the set of actions is transmitted via the computing device.And in response to an input received via the computing device, triggering the execution of one or more of the following actions, wherein the set of actions includes one or more of the following: (1) requesting a set of additional features associated with a particular model use case via the computing device, (2) modifying the set of features of the particular model use case, or (3) adjusting at least one parameter of at least one or more of the first set of AI models or the second set of AI models.

9. The method of request item 8, wherein the set of attributes is a set of expected attributes, and the set of actions is a first set of actions, further comprising: Compare the common expected attributes in each expected model use case in the group of expected model use cases with the observed attributes defined in one or more observed model use cases in the group of observed model use cases; and respond to the absence of one or more attributes in the group of expected attributes in the group of observed attributes, use the group of expected attributes to generate a configuration to add the missing one or more attributes in the group of expected attributes to one of the second group of actions in the group of observed attributes.

10. The method of claim 8, further comprising: A third set of AI models is used to identify a portion of the use cases of the observed models, wherein the third set of AI models is configured to perform a capture augmentation generation (RAG) search of one of the job data using the set of attributes.

11. The method of claim 8, further comprising: The group of generated actions is used to automatically trigger one of the automated workflows indicated by the group of generated actions, wherein the automated workflow includes performing the group of generated actions.

12. The method of claim 11, wherein the set of criteria includes a set of threshold metrics for power consumption and data center usage for the particular observed model use case, wherein the set of observed metrics for power consumption and data center usage for the particular observed model use case is higher than the set of threshold metrics, and wherein the automated workflow reduces power consumption and data center usage for the particular observed model use case.

13. The method of claim 11, further comprising: Each specific observed model use case of the observed model use case group is transmitted to one or more nodes of the input layer of the second set of AI models to verify that the set of criteria for each observed model use case is met.

14. The method of claim 8 further includes, for each specific observed model use case of the group of observed model use cases: mapping each specific observed model use case to a timestamp of one of the group of alphanumeric characters and the time of construction of the specific observed model use case; and using a third set of AI models to aggregate a set of documents associated with the specific observed model use case, wherein the set of documents satisfies the set of criteria corresponding to the mapped risk category, and wherein the set of documents contains the mapped timestamp.

15. A computer system comprising: At least one hardware processor; and at least one non-transitory memory storing instructions, which, when executed by the at least one hardware processor, cause the system to: receive from a computing device: (1) a set of alphanumeric characters, whose definitions are configured to comply with the constraints of the set of alphanumeric characters, one or more operational boundaries of a set of expected model use cases; and (2) a set of job data, wherein the set of expected model use cases includes a common set of attributes among the expected model use cases in the set of expected model use cases; transmit the common set of attributes among the set of expected model use cases to one or more nodes of the input layer of a first set of AI models, so as to receive a set of observed model use cases from the set of job data from one or more nodes of the output layer of the first set of AI models; Each specific observed model use case in the set of observed model use cases includes a set of features for that specific observed model use case, and the set of features for that specific observed model use case includes one or more of the following: a text-based description of that specific observed model use case, an expected input of that specific observed model use case, an expected output of that specific observed model use case, one or more AI models configured to generate the expected output of that specific observed model use case using the expected input of that specific observed model use case, or data supporting the one or more AI models; Each specific observed model use case in the set of observed model use cases is transmitted to one or more nodes of the input layer of a second set of AI models, the second set of AI models being trained to: map the set of literal characters and the set of features of the specific observed model use case to a risk category defined within the set of literal characters, wherein the second set of AI models is configured to select the risk category from a plurality of risk categories defined within the set of literal characters based on a risk level associated with the set of features. Identify a set of criteria for a specific observed model use case within the set of alphanumeric characters, and generate a set of gaps for the specific observed model use case by comparing the set of criteria with the set of features of the specific observed model use case; transmit an indication of the set of gaps via the computing device; trigger the execution of a set of actions to modify the set of operational data in response to an input received via the computing device; and transmit each specific observed model use case of the set of observed model use cases to one or more nodes of the input layer of the second set of AI models to verify the satisfaction of the set of criteria of each observed model use case.

16. The system of claim 15, wherein the set of discrepancies is a first set of discrepancies, wherein the system further causes the system to: map each specific observed model use case of the set of observed model use cases to one or more of the following: a timestamp of one of the set of alphanumeric characters, or the time of construction of the specific observed model use case, and use a third set of AI models to summarize a first set of documents associated with the specific observed model use case, wherein the first set of documents meets the set of criteria corresponding to the mapped risk category; provide a second set of documents associated with the specific observed model use case; and identify a second set of discrepancies by comparing the first set of documents with the second set of documents, wherein the second set of discrepancies includes one or more of the following: documents present in the first set of documents and missing in the second set of documents, or documents present in the second set of documents and missing in the first set of documents.

17. The system of request item 16, wherein the set of actions includes one or more of the following: requesting a set of additional features associated with a particular model use case via the computing device, modifying the set of features of the particular model use case, or adjusting at least one parameter of one or more of the first set of AI models or the second set of AI models.

18. The system of request item 15, wherein the system further causes: to automatically trigger an automated workflow indicated by the group of generated actions, wherein the automated workflow includes performing the group of generated actions.

19. The system of claim 15, further causing the system to: determine, for each gap in the set of gaps, a type of gap between the set of criteria for that particular observed model use case and the set of features for that particular observed model use case, wherein the type of gap is associated with one or more of the following: the presence of one of a particular set of alphanumeric characters in that particular observed model use case, or the semantic meaning of one or more sets of alphanumeric characters in that particular observed model use case; classify each gap in the set of gaps using a difference type; and trigger one or more alerts in response to the classification of one particular gap in the set of gaps reaching a predetermined threshold.

20. The system of claim 15, further causing the system to: assign, for each specific observed model use case of the group of observed model use cases, a version of one or more of the following instructions for each specific observed model use case: (i) the group of features of the specific observed model use case, (ii) the risk category of the group of criteria of the specific observed model use case, (iii) the group of criteria of the specific observed model use case, (iv) the group of gaps of the group of criteria of the specific observed model use case, (v) a set of corresponding alphanumeric characters of the group of criteria of the specific observed model use case, or (vi) an interpretation of one or more of the following sets of parameters used by at least one or more of the first group of AI models or the second group of AI models to produce the risk category, the group of criteria, or the group of gaps of the specific observed model use case.