Generative artificial intelligence output verification engine in artificial intelligence system

By using a multi-category analysis model and human feedback mechanism in the generative AI output verification engine, the efficiency and accuracy issues of generative AI output evaluation in existing technologies are solved, enabling efficient and targeted improvement of the quality assessment of generative AI output and supporting quality assurance in a wide range of application scenarios.

CN122003679APending Publication Date: 2026-05-08MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MICROSOFT TECHNOLOGY LICENSING LLC
Filing Date
2024-10-17
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies lack efficient information-based verification metrics to evaluate the quality of generative AI outputs, especially in providing targeted improvement opportunities when understanding high-quality or low-quality outputs. Furthermore, the evaluation process is time-consuming and difficult to scale to different types of scenarios, and suffers from subjectivity and a lack of benchmark truth values.

Method used

A generative AI output verification engine is adopted, including a lexical analysis model, a semantic analysis model, and a human-oriented clarity analysis model. Through multi-category analysis, the quality of generative AI output is quantified. Combined with a human-participatory feedback mechanism, the accuracy and appropriateness of generative AI output are automatically evaluated and verified.

Benefits of technology

It improves the evaluation efficiency and accuracy of generative AI outputs, provides opportunities for targeted improvements, reduces scoring time, enhances the reliability and robustness of generative AI models, and supports quality assurance in a wide range of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122003679A_ABST
    Figure CN122003679A_ABST
Patent Text Reader

Abstract

Methods, systems, and computer storage media for providing generative artificial intelligence (AI) output verification using a generative AI output verification engine in an artificial intelligence system. A generative AI output validation engine evaluates and determines the quality of a generative AI output (e.g., LLM output) (e.g., quantified to an output validation score). In operation, a generative AI output including summary data is accessed. Raw data from which summary data is generated is accessed. A plurality of output verification operations associated with the generative AI output verification engine are performed. The generative AI output verification engine includes a multi-class analysis model that provides corresponding output verification operations for quantifying the quality of the generative AI output. An output verification score is generated for the summary data using a generative AI output verification engine. The output verification score is transmitted. A feedback loop is established to incorporate human feedback for trimming the generative AI output verification engine model.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 596,290, filed November 11, 2023, the entire contents of which are incorporated herein by reference. Background Technology

[0002] Users rely on computing environments with applications and services to complete computational tasks. Users can interact with different types of applications and services supported by artificial intelligence (AI) systems. Specifically, generative AI systems can support text generation, image generation, music and audio generation, video generation, and data synthesis. Generative AI can refer to a class of AI systems and algorithms designed to generate new data or content that is similar to, or in some cases quite different from, the data on which they are trained. Generative AI can encompass a wide range of models and algorithms designed to generate new data or content. For example, Large Language Models (LLMs) are a specific category of generative AI models that primarily focus on generating human-like text. LLMs and other generative AI models leverage computational architectures, extensive pre-training on datasets, and task-specific fine-tuning to support natural language processing applications ranging from chatbots and virtual assistants to content generation and language translation. Summary of the Invention

[0003] The various aspects of the technology described herein generally relate to systems, methods, and computer storage media, particularly for providing generative artificial intelligence (AI) output verification engines that utilize artificial intelligence systems. Generative AI output verification engines evaluate and verify outputs from generative AI models (e.g., large language models). These engines incorporate multi-category analysis models (e.g., lexical analysis models, semantic analysis models, and human-oriented clarity analysis models) that provide corresponding operations for quantifying the quality of the generative AI output. The engines assess and determine quality (e.g., quality quantified as an output verification score) based on verifying that the generated content meets criteria and standards consistent with the intended use or context of the generated AI output. In one example, generative AI output verification may specifically involve verifying that summary data is complete, accurate, and clear, where the summary data is generalized from source data.

[0004] Generative AI output validation can be based on a generative AI output evaluation framework. This framework supports the identification of informational metrics indicating the quality of generative AI outputs. It helps determine when and why a generative AI output is of low quality. The framework operates to provide diagnostics to identify areas where the generative AI model associated with the output can be improved. Furthermore, it can operate to reduce the scoring time for generative AI outputs (e.g., event summaries) by automatically evaluating the quality of each instance of the output (e.g., event summary) based on raw data (e.g., event data). The framework supports multiple categories and uses respective techniques to evaluate generative AI outputs based on three analytics engines: lexical analysis, semantic analysis, and human-oriented clarity analysis. Each analytics engine generates a corresponding score, which can be parsed to a final score (e.g., output validation score) for each instance of the generative AI output being evaluated. This final score can be used as a rating metric to represent the overall quality of the generative AI output. Imagine enabling small-sample human-participatory evaluations of generative AI outputs to assess trust in a generative AI output evaluation framework. This trust could be replaced by customer feedback on the generative AI outputs, thereby automating the evaluation process.

[0005] Typically, AI systems lack the comprehensive computational logic and infrastructure necessary to efficiently provide information-based verification metrics for generative AI outputs. Existing evaluation metrics often produce general assessments, lacking specificity to identify opportunities for improvement in suboptimal AI-generated outputs. Generative AI evaluation metrics lack measures for human clarity (e.g., quantitative measures of how easily and unambiguously the generative AI output is understood by a human audience). Generative AI outputs can be evaluated manually; for example, security researchers can score event summaries based on several parameters, or other types of evaluators can examine generative AI outputs to determine their quality. Manual evaluation is time-consuming and cannot scale to support large volumes of generative AI outputs across different types of scenarios. Generative AI algorithms can include evaluation metrics, but these algorithms are limited in their ability to assist in understanding whether outputs are high-quality or low-quality. Manual techniques must be employed to understand the factors leading to low-quality outputs and then identify opportunities for improvement.

[0006] Furthermore, evaluating generative AI outputs can be challenging due to their subjective and context-dependent nature. For example, evaluating LLM-based event summaries can be difficult given the complexity of LLMs and the creativity of their outputs. A single technique for evaluating generative AI outputs across multiple dimensions does not exist in conventional generative AI output validation systems. Moreover, generative AI outputs may lack a benchmark truth and include model biases and illusions that make evaluating their quality challenging.

[0007] Technical solutions addressing the limitations of conventional AI systems can resolve the challenges of implementing a generative AI output evaluation framework within an AI system that supports multiple categories and uses their respective techniques to evaluate generative AI outputs; and the challenges of providing generative AI output verification operations and interfaces via a generative AI output verification engine within an AI system. The generative AI output evaluation framework can provide quality assurance that can be implemented across a wide range of applications to ensure the accuracy and appropriateness of the outputs generated by generative AI models. Therefore, AI systems can be improved based on generative AI output verification operations, which are used to efficiently perform the evaluation and verification of generative AI model outputs.

[0008] The operation involves accessing the generative artificial intelligence (AI) output, which includes summary data. It also accesses the raw data used to generate the summary data. Several output validation operations associated with the generative AI output validation engine are performed. The generative AI output validation engine includes a multi-category analysis model that provides corresponding operations for quantifying the quality of the generative AI output. Using the generative AI output validation engine, output validation scores are generated for the summary data. These output validation scores are then transmitted to indicate the quality of the generative AI output, based on verification that the content of the generated AI output meets criteria and standards consistent with the intended use or context of the generated AI output.

[0009] This summary is provided to present, in a simplified form, the selection of concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to help determine the scope of the claimed subject matter. Attached Figure Description

[0010] The technology described herein is described in detail below with reference to the accompanying drawings, wherein: Figure 1A and Figure 1B This is a block diagram of an exemplary artificial intelligence system including a generative AI output verification engine, based on various aspects of the technology described herein; Figures 1C to 1G This is a schematic diagram relating to an exemplary artificial intelligence system, including a generative AI output verification engine, based on various aspects of the techniques described herein; Figure 2A This is a block diagram of an exemplary artificial intelligence system including a generative AI output verification engine, based on various aspects of the technology described herein; Figure 2B This is a block diagram of an exemplary artificial intelligence system including a generative AI output verification engine, based on various aspects of the technology described herein; Figure 3 This paper provides a first exemplary method for providing generative AI output verification using a generative AI output verification engine based on various aspects of the techniques described herein; Figure 4 A second exemplary method is provided for providing generative AI output verification using a generative AI output verification engine based on various aspects of the techniques described herein; Figure 5 A third exemplary method for providing generative AI output verification using a generative AI output verification engine based on various aspects of the techniques described herein is provided. Figure 6 Block diagrams are provided of exemplary distributed computing environments suitable for implementing various aspects of the techniques described herein; and Figure 7 This is a block diagram of an exemplary computing environment applicable to implementing various aspects of the techniques described herein. Detailed Implementation

[0011] Overview An artificial intelligence (AI) system refers to an AI computing environment or architecture that includes the infrastructure and components supporting the development, training, and deployment of AI models. It provides developers with the necessary hardware, software, and frameworks to create and run AI applications. AI systems can be cloud-based AI solutions that utilize cloud computing infrastructure to develop, train, deploy, and manage AI models and applications. AI models can specifically refer to generative AI models, which are designed to generate new data or content that is similar to or, in some cases, entirely different from the data on which they were trained. Applications can be associated with different types of domains, from cloud computing to security management.

[0012] Artificial intelligence systems can support different types of AI models. Generative AI models can be used in various ways, including: content generation, product image generation, personalized product recommendations, natural language chatbots, and content creation. Overview Yes. Traditional AI models encompass a wide range of algorithms and techniques and can be used in various ways, including: recommender systems, predictive analytics, search algorithms, fraud detection, user segmentation, image classification, natural language processing (NLP), and A / B testing and optimization.

[0013] Artificial intelligence systems can include transformer models capable of performing complex neural language processing tasks. Transformer models (including, but not limited to, large language models "LLMs") have applications across a wide range of industries. LLMs are trained deep learning models that can use very large datasets to identify, summarize, translate, predict, and generate content. LLMs and other types of generative AI models are associated with training and inference phases. In the training phase, the model is taught to learn patterns, relationships, and knowledge from the training dataset; the inference phase includes predicting, classifying, or generating outputs for real-world tasks or queries. Artificial intelligence systems can also include convolutional neural networks, which are typically used for image tasks and rely primarily on convolutional operations.

[0014] Generative AI outputs from different types of generative AI models can be associated with different types of applications, services, and systems. For example, security analysts can use an AI assistant (e.g., Microsoft Copilot) that leverages generative AI models (e.g., large language models) to generate summaries of security incidents (“incident summaries”). Incident summaries help to concisely contextualize security incidents, focusing on key points. Improving the quality of security summaries can help improve the functionality of security systems by enabling faster mitigation actions against potentially harmful incidents.

[0015] Typically, AI systems lack the comprehensive computational logic and infrastructure necessary to efficiently provide information verification metrics. Evaluating generative AI outputs can be challenging due to their subjective and context-dependent nature. Existing evaluation metrics often provide general results and lack specificity to identify opportunities for improvement in suboptimal AI-generated outputs, thus limiting their ability to aid in understanding situations involving high-quality or low-quality outputs. Manual techniques must be employed to understand the factors leading to low-quality outputs and then identify opportunities for improvement. For example, security researchers might score event summaries based on several parameters, or other types of evaluators might examine generative AI outputs to determine their quality. Manual evaluation is time-consuming and cannot be scaled to support large volumes of generative AI outputs across different types of scenarios.

[0016] Furthermore, manual evaluation of generative AI outputs can be challenging due to the subjectivity associated with human perspectives, and quality scores may depend on the nature of how humans perceive the definition of a metric. For example, when evaluating an LLM-based event summary that includes misunderstandings, one evaluator might identify it as a hallucination, while another might identify it as a misunderstanding. A single technique for evaluating generative AI outputs across multiple aspects does not exist in conventional generative AI output verification systems. Moreover, generative AI outputs may lack a benchmark truth in real-world scenarios; for example, in the case of LLM-generated event summaries, there is no "perfect" event summary that can be used as a "benchmark truth" for comparison, which also makes evaluating the quality of generative AI outputs challenging. Therefore, AI systems can be improved based on generative AI output verification operations to efficiently perform the evaluation and verification of generative AI model outputs.

[0017] Embodiments of this technical solution relate to systems, methods, and computer storage media for providing generative artificial intelligence (AI) output verification engines, particularly those using artificial intelligence systems. The generative AI output verification engine evaluates and verifies outputs from generative AI models (e.g., large language models). The generative AI output verification engine includes multi-class analysis models (e.g., lexical analysis models, semantic analysis models, and human-oriented clarity analysis models). Multi-class analysis models refer to models that quantify the quality of generative AI outputs based on operations and techniques associated with each corresponding model. Multiple models in output verification lead to more robust, reliable, and accurate results, thereby enabling a deeper understanding of the verification.

[0018] Generative AI output verification engines assess and determine quality (e.g., quantified as output verification scores) based on criteria and standards that verify the generated content meets the intended use or context of the generated AI output. In one example, generative AI output verification might specifically involve verifying summary data (e.g., a security "event summary") against raw data (e.g., security "event data"). Generative AI output verification is provided using a generative AI output verification engine that is operatively integrated into the artificial intelligence system. The artificial intelligence system supports a generative AI output evaluation framework for the computational components associated with providing generative AI output verification. The generative AI output evaluation framework operates to provide diagnostics to identify areas where the generative AI model associated with the generative AI output can be improved.

[0019] refer to Figures 1C-1GThe generative AI output validation engine 110 C can be based on a generative AI output evaluation framework (i.e., evaluation framework 111 C). For example, evaluation framework 111 C can operate to reduce the scoring time of generative AI outputs (e.g., event summaries) by automatically evaluating the quality of each instance of the output (e.g., event summaries 104 C) against the raw data (e.g., raw event data 102 C). Evaluation framework 111 C supports multiple analysis categories and uses respective techniques to evaluate generative AI outputs based on three analysis engines: lexical analysis 120 C, semantic analysis 130 C, and human-oriented clarity analysis 140 C. The analysis engines are associated with corresponding validation criteria (e.g., omissions and illusions 122 C for lexical analysis; contextual understanding, information accuracy, and completeness 132 C for semantic analysis; and conciseness, relevance, fluency, coherence, tone, and logical flow 142 C for human-oriented clarity analysis). Validation criteria can refer to predefined standards or guidelines associated with the corresponding analysis engine, which are used to evaluate and assess the quality of the event summaries.

[0020] The evaluation process can be divided into three categories (e.g., models): lexical analysis model 120C, semantic analysis model 130C, and human-oriented clarity analysis model 140C (collectively referred to as the "multi-category analysis model"). Each model can be used to generate a corresponding final analysis score (e.g., a quality assessment score 150C_1 including lexical analysis score 152, semantic analysis score 154C, and clarity analysis score 156C). The final analysis score provides targeting for improvement opportunities for low-quality outputs because each final analysis score has its own informational validation metric and provides specific insights into understanding high-quality versus low-quality outputs. Imagine that each individual score can be parsed into an aggregated final score calculation module to assign the aggregated final score to each event based on criteria. The aggregated final score can be used as a quality assessment metric representing a summary of the source data. Each model in the evaluation framework 111C can be generalized and reused in a wide range of applications, from security to chatbot interfaces. In this way, the output validation engine 110 C can generate an output validation score, which can refer to any one of the quality assessment scores of the analytical model or an aggregate score based on two or more quality assessment scores.

[0021] It is also envisioned that human-participatory quality assessments could be implemented for small samples to evaluate trust in the framework. Human-participatory assessments could include presenting event summaries and their scores to human verifiers for review. Verifiers would assess the accuracy, relevance, and completeness of the summaries and provide feedback on discrepancies or errors. This feedback would then be used to refine and improve the assessment framework 111 C, enhancing its ability to verify event summaries against raw event data. For example, a data sample, including logs and scores output from the generative AI output engine, could be obtained from the generative AI output engine storage device 160 C. This data sample, logs, and scores would be transmitted to a dashboard 170 C to report quality metrics. The generative AI output engine scores could be retrieved (e.g., by the generative AI output model developer 180 C_1) from both the generative AI output engine storage devices 160 C and 170 C to transmit quality metrics and enhance the generative AI model to generate more accurate summaries.

[0022] Go to Figure 1D Lexical analysis 120 C (or lexical analysis model 120 C) includes preprocessing 120 C_1 of event data 102 C and preprocessing of event summaries 104 C as the initial stage of data analysis. This includes cleaning, transforming, and preparing the event data 102 C and event summaries 104 C for further analysis. The event summaries undergo additional processing, where lexical analysis model 102 C performs lexical reduction 120 C_3 to reduce the content to its basic or root form, while ensuring that the simplified form is linguistic and meaningful. In this way, lexical analysis model 120 C operates by comparing the lexical form of the source data (i.e., the raw data or raw event data) with its summary (i.e., the summary data, the event summary data). Validation criteria 122 C are associated with lexical analysis 120 C, including omissions and illusions. For example, the source data could be tabular event data of security events from which entities with field names are extracted. Tabular event data refers to structured data organized by rows and columns, typically stored in a tabular format such as a spreadsheet or database table. In the context of a security incident, tabular event data may include fields such as timestamps, event IDs, severity levels, event descriptions, affected entities, and any actions taken in response to the event. Entities extracted from this data may include users, devices, IP addresses, application names, and other relevant entities involved in the security incident. Field names may vary but are generally designed to capture key information about the incident for analysis, investigation, and response purposes.

[0023] Part-of-speech taggers (i.e., POS taggers) 102 C_4 are used to extract key information (called entities) from generative AI output. POS tagger 102 C_4 is a natural language processing algorithm that assigns grammatical tags to words in a sentence based on their syntactic roles and relations. These tags typically indicate the part of speech of each word, such as noun, verb, adjective, adverb, etc., as well as additional information such as tense, number, and gender. For example, generative AI output (i.e., summary data) generated from a security incident (i.e., raw data) contains key information as entities (e.g., IP addresses, domain names, and threat indicators (“TI”) entities), and the POS tagging algorithm is used to extract these entities from the generative AI output. Specifically, threat indicators (also known as intrusion indicators (IOCs)) are observable patterns that indicate the presence of cybersecurity threats within a system or network. These indicators can range from anomalous network traffic to suspicious user account activity and anomalous system behavior. The goal of this analysis is to verify whether all key information from the raw data is available in the summary data; and whether any unidentifiable information (e.g., hallucinatory information or entities) exists in the summary data.

[0024] In this way, lexical analysis model 120 C can support the extraction of evidential entities from the event summary. Lexical analysis model 120 C uses a POS tagging algorithm, which assigns a label (e.g., verb, noun, and adjective) to each word in the text input. Lexical analysis model 120 C can then filter out only specific labels such as nouns or pronouns from these labels (i.e., filter entity data 120 C_5), thus filtering out data. Lexical analysis model 102 C may also include heuristic rules 120 C_6, which are applied to compare the event data and the event summary. Heuristic rules 120 C_6 are guidelines or strategies for solving problems or making decisions when exhaustive search or formal algorithmic methods are impractical or impossible. Heuristic rules are provided to compare the summary data with the original data, aiming to identify any instances where the summary data lacks significant entities from the original data (omissions) or contains additional information (illusions) not present in the original data.

[0025] Using heuristic rule 120C_6, preprocessed event data 120C_1 and filtered entities from event summaries 120C_5 are compared to identify matching entities, missing entities from the event summaries, and phantom entities in the event summaries. These are then used to calculate the number of missing entities and the number of phantom entities. Once the mechanism is running, a score table with fields such as event id, org id, summary id, matching entities, missing entities, phantom entities, number of entities missing from evidence, number of entities missing from TI, and number of phantom entities can be generated. Therefore, lexical analysis model 120C supports the identification of missing information and phantom information from summaries with the help of a lexical analysis model. Phantom information refers to data or details created or fabricated by the generative AI model in the event summary, rather than based on real or observed patterns in the original event data.

[0026] Referring to Semantic Analysis 130 C (or Semantic Analysis Model 130 C), Semantic Analysis Model 130 C is responsible for performing contextual analysis and evaluating the completeness of the generative AI output (e.g., a completeness determination algorithm). For example, using a completeness determination algorithm, an event summary generated by an LLM can be compared with event data to determine the accuracy and completeness of the information conveyed in the output. The goal of this analysis is to assess the conceptual similarity between the summary and the source data. Semantic Analysis Model 130 C determines whether the summary captures the core meaning of the original content and does not introduce unexpected or incorrect interpretations. Semantic analysis can be configured to first create contextual embeddings (i.e., contextual embedding 130 C_1) from the original data (e.g., event data); and create contextual embeddings (i.e., contextual embedding 130 C_2) from the generative AI output (e.g., event summary). Semantic Analysis Model 130 C then operates to compute the similarity of those embeddings (i.e., cosine similarity 130 C_3). Contextual embeddings can be created using a transformer-based bidirectional encoder representation (BERT), a robustly optimized BERT method "ROBERTA", or a paraphrase-MiniLM-L6-v2 model. Similarity scores can be computed using several techniques, including ("BERTScore") or cosine similarity. Semantic analysis models can be specifically implemented to address the limitations of lexical analysis models. For example, if within an event, there are two alerts, A1 and A2, where A1 occurred on July 12, 2023, and A2 occurred on July 13, 2023, and an AI-generated summary indicates that A2 occurred on July 12, 2023, and A1 occurred on July 13, 2023, then high lexical overlap exists even with contextual errors because the original entity from the event data is itself included in the event summary.

[0027] When executing the semantic analysis model, a context similarity score (i.e., semantic similarity 130 C_4) is generated. The context similarity score is used to calculate the usefulness score. In one example implementation, the context similarity score can be used to calculate a usefulness score library by subtracting the number of TI omissions calculated from the lexical analysis process, to place more emphasis on the availability of security risk information in the profile; for example, a lower score is assigned if threat actor data is not provided in the profile (if such data exists in the source data). Usefulness score = Context similarity score - Number of significant omissions. More generally, a usefulness score can be generated based on the context similarity score and the number of key omissions.

[0028] A combination of lexical and semantic analysis scores can be used to determine whether a generative AI output summary of the source data is an illusion or an explanatory response, or a good fit to the context of the source data. A detailed explanation of the scores follows.

[0029]

[0030] Referring to human-oriented clarity analysis (or clarity model 140 C), clarity model 140 C can be implemented to assess user trust and satisfaction (e.g., user trust and evaluation algorithms) based on identified metrics (e.g., relevance, coherence, fluency, and conciseness). Clarity model 140 C can be specifically implemented to address the limitations of lexical and semantic analysis models. As an illustration, while lexical and semantic analysis provide coverage from the perspectives of key information coverage and contextual understanding of information, respectively, these models lack a human-centric perspective. For example, a summary might contain all the key information, but it is piled up in the same sentence, making it time-consuming and difficult for users to understand.

[0031] Clarity Model 140 C employs Cueing Engineering 140 C_1, which includes designed and refined cues to guide the LLM Model 140 C_2 in natural language processing tasks that support the assessment of the clarity of event summaries. Clarity Model 140 C can also consider the tone and logical flow of information when evaluating criteria such as relevance, coherence, fluency, and conciseness. LLM 140 C_2 (e.g., generating the pre-trained transformer GPT-4) can be implemented as an evaluation model based on user trust and satisfaction assessment algorithms.

[0032]

[0033] Once executed, each prompt can measure the event summary in terms of conciseness, coherence, fluency, and relevance. Each prompt can be used to generate an intermediate score (e.g., intermediate score 140 C_3), which can be further processed to generate a clarity score 140 C_4 and a validation score 140 C_5. Scores related to relevance, coherence, fluency, and conciseness can be reduced if the tone is not neutral and the information flow in the summary is illogical. In addition, each of the prompts defined above also validates the tone and logical flow of information in the Copilot response. When executing the model, four individual scores (e.g., conciseness score, coherence score, fluency score, and relevance score) can be generated, which are then passed to another mechanism to calculate the final clarity score as follows: Clarity Score = (Conciseness Score + Coherence Score + Fluency Score + Relevance Score) / 4. More generally, a clarity score can be generated based on the individual scores. A high-resolution score calculated for any event can indicate a summary that is highly understandable to humans, while a lower score will indicate a summary that is least understandable to humans.

[0034] The generative AI output validation engine may include an output validation scoring engine 150 C that supports processing final analysis scores (e.g., lexical analysis scores, semantic analysis scores, and clarity scores). The output validation scoring engine 150 C may also process validation results, including validation results corresponding to each multi-class analysis model. For example, validation results (e.g., the number of phantom entities, the number of missing entities, usefulness scores, clarity scores, entity matching, etc.) may each be fed to the output validation scoring engine 150 C. The final analysis scores or validation result scores associated with validation may be fed individually and in combination (e.g., the output validation scores) as evaluation metrics that provide an understanding of the quality of instances of the generative AI output. For example, each score (e.g., lexical analysis score, semantic analysis score, or clarity score) may indicate a specific information metric used to evaluate instances of the generative AI output. It is envisioned that the output validation scoring engine 150 C may be configured to generate different types of overall quality score calculations (e.g., assigning weights to individual scores) based on the corresponding context of the analyzed generative AI output. For example, the overall quality score (e.g., the final score) for a security application may be calculated differently from the overall quality score for applications in different domains.

[0035] refer to Figure 1E , Figure 1EA customer feedback tool 180C supporting human participation in the evaluation is shown. Customers can provide feedback on the feedback UI 184C, which can be stored in the generative AI output verification engine's storage along with a quality score generated by the generative AI output verification engine (evaluation framework) 110C. The stored data can be used to create dashboards to continuously monitor any quality inconsistencies 182C between the customer and the generative AI output verification engine. Feedback loops 186C can be created to trigger fine-tuning of the generative AI output verification engine model to adjust the quality evaluation mechanism in case of any inconsistency. Therefore, human participation in the evaluation can be implemented to determine the effectiveness of the evaluation framework. Specifically, the practice can involve sending a sample of event IDs and their generated summaries to a human evaluator to assess all the aforementioned criteria and verify whether the model score and the human score are comparable. An "accuracy" evaluation metric can be used to determine the ratio of cases where the model scores are comparable to cases where no match exists. Furthermore, the scores from this evaluation can be used to fine-tune the model by feeding back mismatches (e.g., FP (false positive) or FN (false negative) cases) to the underlying model. Fine-tuning using human scores can be a manual process or automated in a feedback loop.

[0036] refer to Figure 1F , Figure 1F A sample customer feedback interface 112F associated with a Security Operations Center Analyst 112F is provided. The customer feedback interface 112F includes an event title 120F, event details 122F, an event summary 130F, and feedback options 132F, 134F, and 136F. The human-involved evaluation within the evaluation framework 110C can be automated by using customer feedback tools on the generative AI output. In this human-involved feedback mechanism, the event summary generated by the LLM undergoes a verification process performed by human verifiers. These verifiers, possessing expertise and experience in cybersecurity, carefully review the summary to assess its accuracy, relevance, and completeness. They meticulously evaluate the content to identify any discrepancies, errors, or missing information. In identifying such issues, verifiers provide detailed feedback, highlighting specific areas that need improvement or clarification. This feedback loop 186C serves as valuable input for refining and enhancing the output verification engine algorithms and processes. By analyzing feedback from human verifiers, the output verification engine learns to identify patterns, correct errors, and adjust its approach to produce more accurate and reliable quality scores in the future. This iterative process ensures continuous improvement and optimization of the output verification engine and evaluation framework 110 C, resulting in a higher quality generative AI output quality assessment process and enhancing the overall effectiveness of the security incident response system.

[0037] As an example, in the context of a security management system, the event summaries generated by the LLM are reflected on a security management portal for user access. Users have the flexibility to provide feedback on the portal, including whether the event summary is useful to them, whether any key information is missing, and the additional ability to add their comments. Imagine that user feedback responses can be compared with quality scores generated by a generative AI output evaluation framework to assess the "accuracy" of the framework. Furthermore, the scores from this evaluation can be used to fine-tune the underlying model. Fine-tuning involves further training the model on the feedback to improve its performance. For example, providing feedback on mismatches (e.g., FP (false positive) or FN (false negative) cases). Fine-tuning using human scores can be a manual process or automated within the feedback loop.

[0038] Figure 1G A first SOC interface 110 G without AI assistance and a second SOC interface 120 G with AI assistance are provided. The event summary in the second SOC interface 120 G includes an indication 122 G that the event summary is AI-generated and its accuracy has been verified (e.g., using the techniques described herein). Interface elements (e.g., various visual cues and interactive features) can be employed to display the verified summary event data. These include verification status indicators, such as checkmarks or green ticks, located next to each verified event summary, providing the user with quick visual confirmation of the verification status. Verification scores, presented as numerical or categorical values, can also be attached to each summary, providing a quantitative measure of confidence in its accuracy after verification.

[0039] The second SOC interface 120 G offers clarity and simplicity, presenting a concise overview of key details of a security incident without overwhelming stakeholders with the massive amounts of data required by the first SOC interface 110 G. The second SOC interface 120 G enhances understanding by presenting information in a more accessible format, especially for stakeholders less familiar with technical terminology. Furthermore, the second SOC interface 120 G supports decision-making by helping stakeholders quickly assess the severity and impact of an incident, enabling timely response and mitigation efforts. As the summary simplifies information sharing between different teams and stakeholders, it facilitates efficient communication, thereby promoting collaboration and coordination in incident response. By focusing on actionable insights and reducing cognitive load, the summary directs attention to key aspects of the incident, enabling stakeholders to prioritize actions and allocate resources effectively. Overall, the security incident summary with validated accuracy indicators improves incident communication, decision-making, and response efforts in computing environments.

[0040] Advantageously, embodiments of this technical solution include several inventive features (e.g., operations, systems, engines, and components) associated with an artificial intelligence system having a generative AI output verification engine. The generative AI output verification engine supports generative AI output verification operations for evaluating and verifying outputs from generative AI models, and provides AI system operations and interfaces via the generative AI output verification engine within the AI ​​system. The generative AI output verification operations are solutions to specific problems in AI systems (e.g., the lack of information evaluation metrics for assessing generative AI outputs). The generative AI output verification engine provides an ordered combination of operations for multi-category analysis models (e.g., lexical analysis models, semantic analysis models, and human-oriented clarity analysis models), providing corresponding operations for quantifying the quality of generative AI outputs, thus improving computational operations within the AI ​​system.

[0041] Example systems and operations You can use examples and references Figure 1A-Figure 1B To describe various aspects of the technical solution. Figure 1A A cloud computing system (environment) 100 is shown, which includes an artificial intelligence system 100 A; a generative AI output verification system 100 B; a network 100C; a generative AI output verification engine 110 with generative AI output verification operation 112, a multi-class analysis model engine 114; an output verification scoring engine 120; an artificial intelligence client 130, an application client 132 and application interface data 134; a machine learning engine 140 including a machine learning model 142 (“LLM 142”); an application 150 and an artificial intelligence assistant 160.

[0042] Cloud computing environment 100 provides computing system resources for different types of managed computing environments. For example, cloud computing environment 100 supports the delivery of computing services (including servers, storage, databases, networks, software-integrated applications and services, collectively referred to as "(multiple) services") and artificial intelligence systems (e.g., artificial intelligence system 100 A). Multiple artificial intelligence clients (e.g., artificial intelligence client 130) include hardware or software that accesses resources within cloud computing environment 100. Artificial intelligence client 130 may include applications or services that support client-side functionality associated with cloud computing environment 100. Multiple artificial intelligence clients can access the computing components of cloud computing environment 100 via a network (e.g., network 100 C) to perform computing operations.

[0043] Artificial intelligence system 100 A is responsible for providing an artificial intelligence computing environment or architecture, including infrastructure and components that support the development, training, and deployment of artificial intelligence models. Artificial intelligence system 100 A is responsible for providing generative AI output verification associated with generative AI output verification engine 110. Artificial intelligence system 100 A operates to support the generation of inference for machine learning model 142 (“LLM” 142). Artificial intelligence system 100 A may integrate with components that support providing generative AI output verification for generative AI outputs from different types of generative AI models (e.g., LLM 142).

[0044] Artificial intelligence system 100 A provides an integrated operating environment for applications 150 (e.g., security applications) operating with LLM 142, based on a generative AI output evaluation framework that is related to the verification of generative AI outputs (e.g., these components may participate in accessing or generating event summary "summary data" from event data "raw data" of security events). Artificial intelligence system 100 A integrates a generative AI output verification operation 112 (which supports providing a model with corresponding operations for quantifying the quality of the generative AI outputs to a multi-category analysis model engine), and artificial intelligence system 100 A operates with artificial intelligence system operations and interfaces to effectively provide generative AI outputs for application-related generative AI models.

[0045] The generative AI output verification engine 110 is responsible for providing generative AI output verification operations 112 that support the functions associated with it. Generative AI output verification operations 112 are performed to support multi-category analysis models (e.g., lexical analysis models, semantic analysis models, and human-oriented clarity analysis models) that provide corresponding operations for quantifying the quality of the generative AI output. The generative AI output verification engine 110 includes a multi-category analysis model engine 114 and an output verification scoring engine 120, which operate together to support the functionality of the output verification engine 110. The multi-category analysis model engine 114 is a computational engine that analyzes the generative AI output data using three different techniques (e.g., lexical analysis, semantic analysis, and clarity analysis). The output verification scoring engine 120 is a computational component that scores the generative AI output data based on criteria and standards that verify the generated content meets the intended use or context of the generated AI output, indicating the quality of the generative AI output data.

[0046] Machine Learning Engine 140 is a machine learning framework or library that operates as a tool to provide the infrastructure, algorithms, and capabilities for designing, training, and deploying machine learning models. Machine Learning Engine 140 may include pre-built functions and APIs capable of building and applying machine learning techniques. Machine Learning Engine 140 can provide a machine learning workflow from data processing and feature extraction to model training, evaluation, and deployment. Machine Learning Engine 140 may include LLM 142.

[0047] LLM 142 can refer to a type of machine learning model (e.g., a transformer model). Specifically, LLM 142 can use lexical units as the basic units for processing and understanding text. Lexical units can be as short as a single character or as long as a word, and the model's understanding of text is based on these lexical units. LLM 142 can support natural language understanding, including text generation and machine translation. LLM 142 can support contextual responses, question answering, content generation, language translation, text summarization, task automation, learning and assistance, and accessibility tools. For example, chat interfaces for search engines and other chat interfaces associated with LLM can be generated based on operations performed on the inference phase of the LLM. For illustration, AI client 130 can refer to a user's device. Application client 132 can be a web browser, mobile application, or any software connected to AI system 100 A. Application 150 is hosted in AI system 100 A. AI system 100 A processes requests from AI client 130. Specifically, the user interacts with application client 132 and provides input (e.g., text prompts). Input can be a text request, question, or instruction that the user wants LLM 142 to process. The user submits input through application client 132. Input, along with any additional parameters or context, is typically transmitted to LLM 142 in a structured format over a secure HTTPS connection.

[0048] Upon receiving input, the AI ​​system 100A processes the request, including passing the input through layers of the LLM 142, leveraging its pre-trained knowledge, and applying the components and neural network architecture of the LLM 142 to generate a response. The LLM 142 performs inference on the input, which involves making predictions based on patterns and information learned during pre-training and fine-tuning. The LLM 142 generates a response based on the input. This response can be in the form of text. The response can include answering questions, providing recommendations, completing text-based or any other language-related task, depending on the nature of the prompts associated with the input and the fine-tuning of the model.

[0049] LLM 142 sends the generated response back to AI client 130 (e.g., a structured data object containing the output of LLM 142). AI client 130 can receive the response and cause it to be displayed. The response can be part of application interface data 134, including responses integrated into text, integrated into a chat interface, or otherwise used based on application design. A request-response interaction occurs where the client sends a prompt, LLM 142 processes the prompt, generates a response, and sends it back to the client for display or further action.

[0050] Application 150 refers to a generative AI-enabled application that can be associated with a broad domain of natural language understanding and generative capabilities based on generative AI models (e.g., LLM 142). Application 150 can support use cases ranging from text generation and autocomplete to chatbots and virtual assistants. Application 150 can integrate with LLM 142 (e.g., via AI Assistant 160) and provide access to LLM 142 via a client (e.g., AI Client 130) that operates based on user interaction, sending queries or prompts, request processing, response processing, and result display. LLM 142 can enhance the capabilities of Application 150 by providing integrated LLM services.

[0051] Application 150 provides a security management system that supports the management of security aspects of data, resources, and workloads in a computing environment. This system helps protect against threats, mitigate risks across different types of computing environments, and strengthen the security posture of the computing environment (i.e., the security status and remediation recommendations for computing resources, including networks and devices). For example, the security management system can provide real-time security alerts, centralized insights into different resources, and offer preventative protection, post-intrusion detection and automated investigation and response. It can also support security management operations (e.g., security investigation queries) that identify potential and actual threats to security posture management.

[0052] The security management system can operate in conjunction with an AI assistant 160 that integrates an AI security engine. AI attack monitoring data can be correlated with the interface of an application that connects the AI ​​model to the computing environment, where the interface between the AI ​​model (e.g., a large language model) and the application supports AI-assisted features tailored to the application (e.g., Microsoft CO-PILOT). AI attack monitoring data can be based on model inputs, model outputs, model behavior, model training and updates, user behavior, contextual validation, and anomaly detection. The security management system can be the security management system described in U.S. Patent Application No. 18 / 451,405, filed August 17, 2023, entitled "AI Engine in a Security Management System," which is incorporated herein by reference in its entirety.

[0053] Therefore, the AI ​​system 100A can provide a generative AI output verification engine 110 that supports verification of generative AI output for LLM 142. The generative AI output verification engine 110 can support multi-category analysis models (e.g., lexical analysis models, semantic analysis models, and human-oriented clarity analysis models) that provide corresponding operations for quantifying the quality of the generative AI output. The generative AI output verification engine supports LLM 142 integrated with application 150, enabling LLM 142 to generate inference using memory 114 and processor 120, and the inference is transmitted to the application's intelligent client 130. The inference for application 150 can be associated with a wide range of domains supported by natural language understanding and generative capabilities.

[0054] refer to Figure 1B , Figure 1B The diagram shows an artificial intelligence system 100A, an output verification system 100B, raw data 122, summary data 124, a generative AI output verification engine 110 with a generative AI output verification operation 112, a multi-class analysis model engine 114 with a lexical analysis engine 114A, a semantic analysis engine 114B, and a clarity analysis engine 114C, an output verification scoring engine 120, an artificial intelligence client 130, a machine learning engine 140 including a machine learning model 142 (“LLM 142”), an application 150, and an artificial intelligence assistant 160.

[0055] Artificial intelligence system 100 A provides output verification system 100 B, which includes a generative AI output verification operation 112 of output verification engine 110. The generative AI output verification operation 112 includes retrieving raw data 122 and summary data 124 (e.g., from the output verification engine storage) to support operations that verify the summary data 124 against the raw data 122. Output verification engine 110 is designed to evaluate and verify outputs generated by generative AI models. Generative AI models (such as generative adversarial networks (GANs) or transformer-based language models such as GPT (generative pre-trained transformers)) can generate various types of output, such as text, images, audio, or video. Output verification engine 110 evaluates these outputs based on predefined criteria or quality metrics to ensure their accuracy, relevance, coherence, and compliance with specified guidelines or constraints. Output verification engine 110 employs generative AI output verification operation 112 to analyze and interpret the generated content. Generative AI output verification operation 112 is a series of tasks designed to evaluate and verify outputs generated by generative AI models. Generative AI Output Validation Operation 112 covers a range of tasks focused on evaluating and validating the accuracy, coherence, and relevance of outputs generated by generative AI models.

[0056] The Multi-Class Analysis Model Engine 114 is a computational framework designed to analyze and process data across multiple categories or dimensions to support the validation of generative AI outputs. The Multi-Class Analysis Model Engine 114 utilizes advanced analytics techniques, including machine learning algorithms and statistical methods, to derive insights and patterns from complex datasets. The Multi-Class Analysis Model Engine 114 includes a lexical analysis engine 114A, a semantic analysis engine 114B, and a clarity analysis engine 114C.

[0057] The output validation scoring engine 120 is designed to evaluate and assess the quality and accuracy of outputs generated by artificial intelligence models (particularly generative AI models). The output validation scoring engine 120 analyzes the generated outputs against predefined criteria associated with statistical analysis, natural language processing, and machine learning algorithms linked to lexical analysis engine 114A, semantic analysis engine 114B, and clarity analysis engine 114C. The output validation scoring engine 120 assigns a score or rating to the output based on its alignment with the desired goal or expected result, as well as its reliability and usefulness for downstream applications or decision-making processes.

[0058] Raw data 122 can refer to unprocessed, unorganized, and unstructured information collected from various sources; raw data 122 may not have undergone any transformation or manipulation. Raw data 122 can represent the original form of the data. Raw data 122 can be secure data from different secure data sources. Summary data 124 refers to a summary of the raw data, corresponding to an overview or compressed representation of the raw data. The purpose of summarizing raw data is to extract key insights, trends, and patterns that can inform decision-making or further analysis. Summary data can be output from LLM 142, which is used to summarize the raw data 122 of summary data 124.

[0059] The generative AI output validation engine 110 refers to a dedicated computing system designed to evaluate and validate the output generated by a generative AI model (e.g., machine learning model 142). The generative AI output validation engine evaluates the generative AI output based on a lexical analysis engine 114A, a semantic analysis engine 114B, and a clarity analysis engine 114C. Each analysis engine generates a corresponding score, which can be parsed into a final score for each instance of the generative AI output being evaluated. The final score can be generated using the output validation scoring engine 120. The final score can be used as a metric to represent the quality of the generative AI output.

[0060] Evaluation of the generative AI output can be performed via: a lexical analysis engine 114A, which performs surface analysis on the generative AI output by decomposing it into basic lexical forms; a semantic analysis engine 114B, which provides an interpretation of the meaning of the generative AI output; and a clarity analysis engine 114C, which provides human-like clarity within the context of the generative AI output. Each model can be used to generate a final score, which can be processed to generate an aggregated final score for each generative AI output. For illustration, the generative AI output data corresponds to a summary (e.g., summary data 124) of security events (e.g., raw data 122) in the computing environment, and the score (i.e., the final score or the aggregated final score) can be used as an evaluation metric to represent the quality of the summary of the security events.

[0061] The generative AI output verification engine 110 includes an output verification scoring engine 120 that supports processing final scores (e.g., lexical analysis scores, semantic analysis scores, and clarity scores). The output verification scoring engine 120 can also process verification results, including verification results corresponding to each multi-class analysis model. For example, verification results (e.g., the number of phantom entities, the number of missing entities, usefulness scores, clarity scores, entity matching, etc.) can each be individually transmitted to the output verification scoring engine 120. The final analysis score or verification result score associated with the verification can be transmitted individually and in combination (e.g., transmitting output verification scores) as an evaluation metric providing an understanding of the quality of instances of the generative AI output. For example, each score (e.g., lexical analysis score, semantic analysis score, or clarity score) can indicate a specific information metric used to evaluate instances of the generative AI output. It is envisioned that the output verification scoring engine 120 can be configured to generate different types of overall quality score calculations (e.g., assigning weights to individual scores) based on the corresponding context of the analyzed generative AI output. For example, the overall quality score (e.g., the final score) for a security application can be calculated differently from the overall quality score for applications in different domains.

[0062] In this manner, output verification engine 110 accesses generative AI output including summary data 124 associated with raw data 122. Output verification engine 110 performs multiple output verification operations associated with the generative AI output verification engine, employing a multi-class analysis model 114 with corresponding operations for quantifying the quality of the generative AI output. Based on the execution of the multiple output verification operations, output verification engine 110 generates an output verification score associated with the summary data. Output verification engine 110 transmits the output verification score 110.

[0063] The output validation engine 110 may include human-involved evaluation, which can be implemented to determine the effectiveness of the evaluation framework. Specifically, a practice can be implemented where samples of event IDs and their generated summaries are sent to human evaluators to assess all the aforementioned criteria and verify whether the model scores and human scores are comparable. An "accuracy" evaluation metric can be used to determine the ratio of cases where model scores are comparable to cases where no match exists. Furthermore, the scores from this evaluation can be used to fine-tune the model by feeding back mismatches (e.g., FP (false positive) or FN (false negative) cases) to the underlying model.

[0064] Various aspects of the technical solution can be illustrated through examples and references. Figure 2A and Figure 2B To describe. Figure 2A Based on reference Figure 6 and Figure 7A block diagram of an exemplary technical solution environment is described, illustrating an example environment for implementing embodiments of the technical solution. Typically, the technical solution environment includes a technical solution system suitable for providing an example artificial intelligence system 100A in which the methods of this disclosure can be employed. Specifically, Figure 2A A high-level architecture of an artificial intelligence system 100 A according to an embodiment of this disclosure is illustrated. Among other engines, managers, generators, selectors, or components (collectively referred to herein as "components") not shown, the technical solution environment of the artificial intelligence system 100 A corresponds to... Figure 1A and Figure 1B .

[0065] refer to Figure 2A , Figure 2A A cloud computing system (environment) 100 is shown, which includes an artificial intelligence system 100 A; an output verification system 100 B, raw data 122, summary data 124, a generative AI output verification engine 110 with a generative AI output verification operation 112, a multi-class analysis model engine 114 with a lexical analysis engine 114 A, a semantic analysis engine 114 B and a clarity analysis engine 114 C; an output verification scoring engine 120; an artificial intelligence client 130, a machine learning engine 140 including a machine learning model 142 (“LLM 142”); an application 150 and an artificial intelligence assistant 160.

[0066] In some embodiments, a system such as the computerized system described in any of the embodiments above includes at least one computer processor and a computer storage medium storing computer-usable instructions that, when used by the at least one computer processor, cause the system to perform operations. The operations include: accessing generative artificial (AI) output including summary data 124; accessing raw data 122 associated with the summary data 124; performing a plurality of output verification operations associated with a generative AI output verification engine 110, wherein the generative AI output verification engine 110 includes a multi-class analysis model 114 having corresponding operations for quantifying the quality of the generative AI output; generating an output verification score associated with the summary data based on the execution of the plurality of output verification operations; and transmitting the output verification score.

[0067] In any combination of the above embodiments of the system, the summary data 124 includes a summary of the raw data 122 and multiple quality assessment scores, including lexical analysis scores, semantic analysis scores, and clarity analysis scores.

[0068] In any combination of the above embodiments of the system, the plurality of multi-category analysis models 114 include lexical analysis model 114 A, semantic analysis model 114 B and clarity analysis model 114 C.

[0069] In any combination of the above embodiments of the system, performing multiple output verification operations includes: performing a first plurality of output verification operations associated with a lexical analysis model that supports comparing the lexical form of the raw data with the lexical form of the summary data; performing a second plurality of output verification operations associated with a semantic analysis model that supports comparing the contextual analysis of the raw data with the summary data; and performing a third plurality of output verification operations associated with a clarity analysis model that supports evaluating user trust and satisfaction based on multiple identified metrics.

[0070] In any combination of the above embodiments of the system, the output verification score is generated based on: accessing a first final score associated with the lexical analysis model; accessing a second final score associated with the semantic analysis model; accessing a third final score associated with the clarity analysis model; and generating the output verification score based on the first output verification score, the second output verification score, and the third output verification.

[0071] In any combination of the above embodiments of the system, the operation further includes: receiving a request for a security posture of the computing environment; generating a security posture visualization associated with an output verification score based on the request for a security posture of the computing environment; and transmitting the security posture visualization to cause the display of the security posture visualization.

[0072] In any combination of the above embodiments of the system, the customer feedback mechanism supports presenting summary data to human verifiers for review and providing feedback on any discrepancies or errors, where the feedback is used to refine the generative AI output verification engine.

[0073] In some embodiments, one or more computer storage media have computer-executable instructions contained thereon that, when executed by a computing system having a processor and memory, cause a processor to perform operations. These operations include: transmitting a request for a security posture of a computing environment; based on the request for the security posture of the computing environment, accessing a security posture visualization associated with an output verification score, wherein the output verification score is generated using a generative AI output verification engine, the generative AI output verification engine including a multi-class analysis model having corresponding operations for quantifying the quality of the generative AI output; and causing a display of the security posture visualization.

[0074] In any combination of the above embodiments of the medium, the plurality of multi-category analysis models 114 include lexical analysis model 114 A, semantic analysis model 114 B and clarity analysis model 114 C.

[0075] In any combination of the above embodiments of the medium, the lexical analysis model 114 A employs a part-of-speech tagging algorithm to support the comparison of the lexical form of the raw data with the lexical form of the summary data.

[0076] In any combination of the above embodiments of the medium, semantic analysis model 114 B employs an integrity determination algorithm to support the comparison of contextual analysis of the raw data with summary data.

[0077] In any combination of the above embodiments of the medium, the clarity analysis model 114 C employs user trust and satisfaction assessment algorithms to support the assessment of user trust and satisfaction based on multiple identifier metrics.

[0078] In any combination of the above embodiments of the medium, the operation further includes: receiving a request for a security posture of the computing environment; generating a security posture visualization associated with an output verification score based on the request for a security posture of the computing environment; and transmitting the security posture visualization to cause the display of the security posture visualization.

[0079] In any combination of the above embodiments of the medium, the operation further includes: transmitting a request for a security posture of the computing environment; based on the request for a security posture of the computing environment, accessing a security posture visualization associated with the output verification score; and causing a display of the security posture visualization.

[0080] In some embodiments, a computer-implemented method is provided. The method includes: accessing generative AI output associated with a generative artificial intelligence (AI) model; using a generative AI output verification engine, performing a plurality of output verification operations associated with a multi-class analysis model having corresponding operations for quantifying the quality of the generative AI output, wherein performing the plurality of output verification operations includes: performing a first plurality of output verification operations associated with a lexical analysis model 114A; performing a second plurality of output verification operations associated with a semantic analysis model 114B; and performing a third plurality of output verification operations associated with a clarity analysis model 114C; generating an output verification score based on the performance of the plurality of output verification operations; and transmitting the output verification score.

[0081] In any combination of the above embodiments of the method, the lexical analysis model 114 A of the multi-class analysis model 114 employs a part-of-speech tagging algorithm to support the comparison of the lexical form of the raw data with the lexical form of the summary data.

[0082] In any combination of the above embodiments of the method, the semantic analysis model 114 B of the multi-class analysis model 114 employs an integrity determination algorithm to support the comparison of contextual analysis of the raw data with summary data.

[0083] In any combination of the above embodiments of the method, the clarity analysis model 114C of the multi-category analysis model 114 employs user trust and satisfaction assessment algorithms to support the assessment of user trust and satisfaction based on multiple identifiers.

[0084] In any combination of the above embodiments of the method, the output verification score is generated based on: accessing a first final score associated with a lexical analysis model; accessing a second final score associated with a semantic analysis model; accessing a third final score associated with a clarity analysis model; and generating the output verification score based on the first final score, the second final score, and the third final score.

[0085] In any combination of the above embodiments of the method, the method further includes receiving a request for a security posture of the computing environment; generating a security posture visualization associated with an output verification score based on the request for a security posture of the computing environment; and transmitting the security posture visualization to cause the display of the security posture visualization.

[0086] refer to Figure 2B , Figure 2B A schematic diagram of an exemplary cloud computing system 100, including application 110, output verification engine 114, and application client 130, is shown. For illustration, application 110 and application client 130 are associated with a security application; however, other applications can be implemented using the functionality of the components described herein. At box 10, application client 130 transmits a request for a security posture of the computing environment. At box 12, application 12 accesses the request for a security posture of the computing environment; and at box 14, it transmits a request for an output verification score for application data. At box 16, output verification engine 114 accesses the request for an output verification score for application data; at box 18, it accesses a generative AI output including summary data associated with security events; accesses raw data associated with the summary data of security events; at box 22, it performs multiple output verification operations associated with a multi-class analysis model; at box 24, it generates an output verification score associated with the summary data; and at box 26, it transmits the output verification score to the application.

[0087] In box 28, application 110 accesses the output verification score for application data; in box 30, a security posture visualization is generated based on the verification score; and in box 32, the security posture verification is transmitted to the application client. In box 34, the application client accesses the security posture visualization associated with the computing environment based on this request; and in box 36, the security posture visualization is displayed based on the output verification score.

[0088] Example Method refer to Figure 3 , Figure 4 and Figure 5 A flowchart illustrating a method for providing generative AI output verification using a generative AI output verification engine in an artificial intelligence system is provided. This method can be performed using the artificial intelligence system described herein. In embodiments, one or more computer storage media having computer-executable or computer-usable instructions thereon, when executed by one or more processors, can cause one or more processors to perform the method (e.g., a computer-implemented method) within an artificial intelligence system (e.g., a computerized system or computing system).

[0089] Go to Figure 3 A flowchart illustrating a method 300 for providing generative AI output verification using a generative AI output verification engine in an artificial intelligence system is provided. At box 302, the generative AI output, including summary data, is accessed. At box 304, the raw data associated with the summary data is accessed. At box 306, multiple output verification operations associated with the generative AI output verification engine are performed. At box 308, an output verification score is generated for each summary data point. At box 310, the output verification score is transmitted.

[0090] Go to Figure 4 A flowchart illustrating a method 400 for providing generative AI output verification using a generative AI output verification engine in an artificial intelligence system is provided. At box 402, the generative AI output is accessed. At box 404, a first plurality of output verification operations associated with a lexical analysis model are performed. At box 406, a second plurality of output verification operations associated with a semantic analysis model are performed. At box 408, a third plurality of output verification operations associated with a clarity analysis model are performed. At box 410, an output verification score is generated based on the execution of the first, second, and third plurality of output verification operations. At box 412, the output verification score is transmitted.

[0091] Go to Figure 5 A flowchart illustrating a method 500 for providing generative AI output verification using a generative AI output verification engine in an artificial intelligence system is provided. At box 502, a generative AI output including summary data associated with a security event is accessed. At box 504, raw data associated with the summary data of the security event is accessed. At box 506, multiple output verification operations associated with the generative AI output verification engine are performed to compare the summary data and the raw data. At box 508, an output verification score associated with the summary data is generated. At box 510, the output verification score is displayed via a graphical user interface of a security application.

[0092] Technological improvements Embodiments of this technical solution have been described with reference to several inventive features (e.g., operations, systems, engines, and components) associated with artificial intelligence systems. The described inventive features include: operations, interfaces, data structures, and arrangements of computational resources associated with the functionality described herein in the generative AI output verification engine. The functionality of embodiments of this technical solution has been further described through implementations and exemplary examples to demonstrate operations (e.g., using multi-category analysis models (lexical analysis models, semantic analysis models, and human-oriented clarity analysis models) that provide corresponding operations for quantifying the quality of generative AI output. The generative AI output verification engine is a solution to a specific problem (e.g., the lack of information evaluation metrics for assessing generative AI output). The generative AI output verification engine improves the computational operations associated with providing generative AI output verification using artificial intelligence systems.

[0093] Additional support for detailed description Example Distributed Computing System Environment Now for reference Figure 6 , Figure 6 An example distributed computing environment 600 in which an implementation of the present disclosure may be adopted is shown. Specifically, Figure 6 illustrates a high-level architecture of an exemplary cloud computing platform 610 that may host a technology solution environment or a portion thereof (e.g., a data hosting environment). It should be understood that the arrangements described herein and others are presented as examples only. For example, as mentioned above, the various elements described herein may be implemented as independent or distributed components, or combined with other components, and implemented in any suitable combination and location. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, and functional groupings) may be used in addition to or in lieu of those shown.

[0094] The data center can support a distributed computing environment 600, which includes a cloud computing platform 610, racks 620, and nodes 630 (e.g., computing devices, processing units, or blade servers) within the racks 620. The technical solution environment can be implemented using the cloud computing platform 610, which runs cloud services across different data centers and geographical regions. The cloud computing platform 610 can implement a structure controller 640 component for provisioning and managing the allocation, deployment, upgrades, and management of cloud services. Typically, the cloud computing platform 610 is used to store data or run service applications in a distributed manner. The cloud computing platform 610 in the data center can be configured to host and support the operation of endpoints for specific service applications. The cloud computing platform 610 can be a public cloud, a private cloud, or a dedicated cloud.

[0095] Node 630 may be equipped with a host 650 (e.g., an operating system or runtime environment) on which a defined software stack runs. Node 630 may also be configured to perform specialized functions (e.g., compute nodes or storage nodes) within the cloud computing platform 610. Node 630 is assigned to run one or more portions of a tenant's service application. A tenant may refer to a customer that utilizes the resources of the cloud computing platform 610. The service application components of the cloud computing platform 610 that support a particular tenant may be referred to as multi-tenant infrastructure or leases. The terms service application, application, or service are used interchangeably herein and refer generally to any software or portion of software that runs on top of or accesses storage and compute equipment locations within a data center.

[0096] When node 630 supports more than one individual service application, node 630 can be partitioned into virtual machines (e.g., virtual machine 652 and virtual machine 654). Physical machines can also run their respective service applications simultaneously. Virtual machines or physical machines can be configured as individualized computing environments supported by resources 660 (e.g., hardware and software resources) in the cloud computing platform 610. It is envisioned that resources can be configured for specific service applications. Furthermore, each service application can be partitioned into functional parts, allowing each functional part to run on a separate virtual machine. In the cloud computing platform 610, multiple servers can be used to run service applications and perform data storage operations in a cluster. In particular, servers can perform data operations independently but appear as a single device referred to as a cluster. Each server in the cluster can be implemented as a node.

[0097] Client device 680 can connect to service applications in cloud computing platform 610. Client device 680 can be any type of computing device, which can correspond to the reference... Figure 7 The described computing device 700, for example, client device 680, can be configured to issue commands to cloud computing platform 610. In embodiments, client device 680 can communicate with service applications via Virtual Internet Protocol (IP) and load balancers or other means of directing communication requests to designated endpoints in cloud computing platform 610. Components of cloud computing platform 610 can communicate with each other via a network (not shown), which may include, but is not limited to, one or more local area networks (LANs) and / or wide area networks (WANs).

[0098] Example computing environment Having briefly described an overview of embodiments of this technical solution, the following describes an example operating environment in which embodiments of this technical solution may be implemented, in order to provide a general context for various aspects of this technical solution. First, refer to... Figure 6The illustration shows an example operating environment for implementing an embodiment of the present technical solution, and is generally designated as computing device 600. Computing device 600 is merely an example of a suitable computing environment and is not intended to impose any limitation on the scope or functionality of the technical solution. Computing device 600 should also not be construed as having any dependency or requirement on any one or combination of the components shown.

[0099] The technical solution can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions (such as program modules) that are executed by a computer or other machine (such as a personal data assistant or other handheld device). Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This technical solution can be implemented in various system configurations, including handheld devices, consumer electronics, general-purpose computers, and more specialized computing devices. It can also be implemented in distributed computing environments, where tasks are performed by remote processing devices linked through a communication network.

[0100] refer to Figure 7 The computing device 700 includes a bus 710 that directly or indirectly couples to the following devices: a memory 712, one or more processors 714, one or more presentation components 716, an input / output port 718, an input / output component 720, and an illustrative power supply 722. The bus 710 can represent one or more buses (such as an address bus, a data bus, or a combination thereof). For clarity of concept, Figure 7 The various boxes are shown with lines, and other arrangements of the described components and / or component functions are also envisioned. For example, a presentation component such as a display device can be considered an I / O component. Furthermore, the processor has memory. We recognize this as an inherent characteristic of the art and reiterate... Figure 7 The figures are merely illustrations of example computing devices that can be used in conjunction with one or more embodiments of this technical solution. No distinction is made between categories such as "workstation," "server," "laptop," and "handheld device," as all of these are... Figure 7 Within the scope and referring to "computing devices" Computing device 700 typically includes a variety of computer-readable media. Computer-readable media can be any available medium accessible by computing device 700, and includes volatile and non-volatile media, removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.

[0101] Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store the required information and is accessible by the computing device 700. The signal itself is excluded from the definition of computer storage media.

[0102] Communication media typically embody computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and include any information transmission medium. The term "modulated data signal" refers to a signal whose one or more characteristics are set or altered in a manner that encodes information in the signal. By way of example and not limitation, communication media include wired media such as wired networks or direct wired connections, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.

[0103] Memory 712 includes computer storage media in the form of volatile and / or non-volatile memory. The memory can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. Computing device 700 includes one or more processors that read data from various entities such as memory 712 or I / O components 720. Presentation component 716 presents data indications to a user or other device. Exemplary presentation components include display devices, speakers, printing components, vibration components, etc.

[0104] I / O port 718 allows computing device 700 to be logically coupled to other devices including I / O components 720, some of which may be built-in. Illustrative components include microphones, joysticks, game controllers, disc-shaped satellite antennas, scanners, printers, wireless devices, etc.

[0105] Additional structural and functional features of embodiments of the technical solution Various components used herein have been identified, and it should be understood that any number of components and arrangements can be employed to achieve the desired functionality within the scope of this disclosure. For example, components in the embodiments depicted in the figures are shown with lines for clarity of concept. Other arrangements of these and other components can also be implemented. For example, although some components are depicted as single components, many elements described herein can be implemented as discrete or distributed components or combined with other components, and implemented in any suitable combination and location. Some elements may be omitted entirely. Furthermore, the various functions described herein as being performed by one or more entities can be performed by hardware, firmware, and / or software, as described below. For example, various functions can be performed by a processor executing instructions stored in memory. Therefore, other arrangements and elements (e.g., machines, interfaces, functions, sequences, and functional groupings) may be used in addition to or in lieu of those shown.

[0106] The embodiments described in the following paragraphs may be combined with one or more of the specifically described alternatives. In particular, in the alternatives, the claimed embodiments may include references to more than one other embodiment. The claimed embodiments may specify further limitations on the claimed subject matter.

[0107] This document specifically describes embodiments of technical solutions to meet legal requirements. However, the specification itself is not intended to limit the scope of this patent. Rather, the inventors have envisioned that the claimed subject matter may also be embodied in other ways in combination with other current or future technologies to include different steps or combinations of steps similar to those described in this document. Furthermore, although the terms “step” and / or “box” may be used herein to refer to different elements of the method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.

[0108] For the purposes of this disclosure, the word "comprising" has the same broad meaning as the word "including," and the word "access" includes "receiving," "referencing," or "retrieval." Furthermore, the word "communication" has the same broad meaning as the words "receiving" or "transmitting," which is facilitated by using a software- or hardware-based bus, receiver, or transmitter of the communication medium described herein. Additionally, unless otherwise indicated, words such as "a" and "an" include both plural and singular forms. Thus, for example, in the presence of one or more features, the constraint of "feature" applies. Furthermore, the term "or" includes both, alternatives, and both (a or b therefore includes a or b as well as a and b).

[0109] For the purposes of the detailed discussion above, embodiments of the present technical solution are described with reference to a distributed computing environment; however, the distributed computing environment depicted herein is merely exemplary. Components can be configured to perform novel aspects of the embodiments, wherein the term "configured for" can mean "programmed to" perform a specific task or implement a specific abstract data type using code. Furthermore, while embodiments of the present technical solution can generally be referenced to the technical solution environment and schematic diagrams described herein, it should be understood that the described technology can be extended to other implementation contexts.

[0110] For the purposes of this disclosure, the term "support" refers to the provision of functionality, services, or assistance by computing components or through computing operations within the broader computing system. When a computing component or set of operations supports a specific function, it means that it plays a role in enabling or performing that specific aspect of the computing system. Such support can manifest in various ways, including data processing, operation execution, resource management, and ensuring compatibility or interoperability with other components. Additionally, support can involve providing interfaces, APIs (Application Programming Interfaces), or protocols that allow seamless interaction and integration with other elements of the computing system. The concept of support extends beyond simply providing functionality to encompass the maintenance, troubleshooting, and overall optimization of computing resources, thereby ensuring the robust and efficient operation of the computing system.

[0111] Embodiments of the present invention have been described with respect to specific examples, which are intended in all respects to be illustrative rather than limiting. Alternative embodiments will become apparent to those skilled in the art without departing from the scope of the present invention.

[0112] As can be seen from the foregoing, this technical solution is well-suited to achieving all the purposes and objectives described above, as well as other obvious and inherent structural advantages.

[0113] It should be understood that certain features and sub-combinations are useful and can be used without reference to other features or sub-combinations. This is contemplated by the claims and is within the scope of the claims.

Claims

1. A computerized system, comprising: One or more computer processors; as well as A computer memory storing computer-usable instructions that, when used by the one or more computer processors, cause the one or more computer processors to perform operations, the operations including: Access (302) includes generative human (AI) output of summary data; Access (304) the raw data associated with the summary data; Perform (306) multiple output verification operations associated with a generative AI output verification engine, wherein the generative AI output verification engine includes a multi-class analysis model having corresponding operations for quantifying the quality of the generative AI output; Based on the execution of the plurality of output verification operations, an output verification score associated with the summary data is generated (308); and The output verification score is transmitted (310).

2. The system according to claim 1, wherein the summary data includes a summary of the original data and multiple quality assessment scores, the multiple quality assessment scores including lexical analysis scores, semantic analysis scores, and clarity analysis scores.

3. The system according to claim 1, wherein the plurality of multi-category analysis models include a lexical analysis model, a semantic analysis model, and a clarity analysis model.

4. The system of claim 1, wherein performing the plurality of output verification operations includes: Perform a first plurality of output verification operations associated with a lexical analysis model that supports comparing the lexical form of the raw data with the lexical form of the summary data; Perform a second set of output verification operations associated with a semantic analysis model that supports comparing contextual analysis of the raw data with summary data. as well as Perform a third set of multiple output verification operations associated with the clarity analysis model, which supports the evaluation of user trust and satisfaction based on multiple identified metrics.

5. The system of claim 1, wherein the output verification score is generated based on the following: Access the first final score associated with the lexical analysis model; Access the second final score associated with the semantic analysis model; Access to the third final score associated with the clarity analysis model; and The output verification score is generated based on the first output verification score, the second output verification score, and the third output verification.

6. The system according to claim 1, wherein the operation further includes: Receive requests regarding the security posture of the computing environment; Based on the request regarding the security posture of the computing environment, a security posture visualization associated with the output verification score is generated; as well as The security posture visualization is transmitted to trigger its display.

7. The system of claim 1 further includes a customer feedback mechanism that supports presenting summary data to human verifiers for review and providing feedback on any discrepancies or errors, wherein the feedback is used to refine the generative AI output verification engine.

8. One or more computer storage media having computer-executable instructions contained thereon, which, when executed by a computing system having a processor and a memory, cause the processor to perform operations, the operations including: Transmit (10) a request for the security posture of the computing environment; Based on the request for the security posture of the computing environment, access (34) a security posture visualization associated with the output verification score, wherein the output verification score is generated using a generative AI output verification engine, the generative AI output verification engine including a multi-class analysis model having corresponding operations for quantifying the quality of the generative AI output; as well as This leads to the visualization of the security situation described in (36).

9. The medium according to claim 8, wherein the plurality of multi-category analysis models include a lexical analysis model, a semantic analysis model, and a clarity analysis model.

10. The medium according to claim 9, wherein the lexical analysis model employs a part-of-speech tagging algorithm to support the comparison of the lexical form of the raw data with the lexical form of the summary data.

11. The medium of claim 9, wherein the semantic analysis model employs an integrity determination algorithm to support the comparison of contextual analysis of the raw data with summary data.

12. The medium of claim 9, wherein the clarity analysis model employs user trust and satisfaction assessment algorithms to support the assessment of user trust and satisfaction based on multiple identified metrics.

13. The medium according to claim 11, further comprising: Receive requests regarding the security posture of the computing environment; Based on the request regarding the security posture of the computing environment, a security posture visualization associated with the output verification score is generated; as well as The security posture visualization is transmitted to trigger its display.

14. The medium according to claim 8, wherein the operation further comprises: Transmit a request for the security posture of the computing environment; Based on the request regarding the security posture of the computing environment, access the security posture visualization associated with the output verification score; and This leads to the visualization of the aforementioned security situation.

15. A computer-implemented method, the method comprising: Access (402) the generative AI output associated with the generative artificial intelligence (AI) model; Using a generative AI output verification engine, multiple output verification operations are performed associated with a multi-class analysis model, which has corresponding operations for quantifying the quality of the generative AI output. The execution of these multiple output verification operations includes: Perform the (404) first plurality of output verification operations associated with the lexical analysis model; Perform (406) a second plurality of output verification operations associated with the semantic analysis model; and Perform (408) the third multiple output verification operation associated with the sharpness analysis model; Based on the execution of multiple output validation operations, an output validation score of (410) is generated; and The output verification score is transmitted (412).

16. The method of claim 15, wherein the lexical analysis model in the multi-category analysis model employs a part-of-speech tagging algorithm to support the comparison of the lexical form of the raw data with the lexical form of the summary data.

17. The method of claim 15, wherein the semantic analysis model in the multi-category analysis model employs an integrity determination algorithm to support the comparison of contextual analysis of the raw data with summary data.

18. The method of claim 15, wherein the clarity analysis model in the multi-category analysis model employs user trust and satisfaction assessment algorithms to support the assessment of user trust and satisfaction based on multiple identified metrics.

19. The method of claim 15, wherein the output verification score is generated based on: Access the first final score associated with the lexical analysis model; Access the second final score associated with the semantic analysis model; Access to the third final score associated with the clarity analysis model; as well as The output verification score is generated based on the first final score, the second final score, and the third final score.

20. The method of claim 15, further comprising: Receive requests regarding the security posture of the computing environment; Based on the request regarding the security posture of the computing environment, a security posture visualization associated with the output verification score is generated; as well as The security posture visualization is transmitted to trigger its display.

Citation Information

Patent Citations

  • Artificial intelligence security engine in a security management system

    US20250061195A1