Enterprise-generated artificial intelligence anti-hallusion and attribution architecture

By introducing anti-hallucination and attribution architectures into generative AI systems, the problem of hallucinations in traditional systems is solved, the accuracy and reliability of generated content are achieved, the traceability of responses is ensured, and the system can be adapted to different models and deployed without refactoring existing systems.

CN121548824APending Publication Date: 2026-02-17SIRUI ARTIFICIAL INTELLIGENCE CO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480044238.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-04-30
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Traditional generative AI systems cannot effectively detect, prevent, or mitigate hallucinations, nor can they verify the accuracy of the generated content, leading to the spread of inaccurate and contradictory information in the enterprise environment.

Method used

It employs an anti-hallucination and attribution architecture, which detects, prevents, and mitigates hallucinations through anti-hallucination and attribution modules. It utilizes a combination of agents and tools to process inputs from different data sources and performs validation and attribution before generating responses to verify the accuracy of the responses.

Benefits of technology

It improves the accuracy and reliability of generative AI content, ensures the traceability and reliability of responses, reduces the possibility of illusions, adapts to different models, and can be deployed without refactoring existing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121548824A_ABST
    Figure CN121548824A_ABST
Patent Text Reader

Abstract

Disclosed herein is an anti-hallucination and attribution architecture for enterprise generative AI systems that increases the accuracy and reliability of generative artificial intelligence content (e.g., responses or answers) by detecting, preventing, and mitigating hallucination. The anti-hallusion and attribution architecture may be added to the deployed generative artificial intelligence system as a separate tool or module, which allows the architecture to work with the deployed system without having to reconstruct or redesign those systems. The anti-hallusion and attribution architecture may also be deployed with minimal impact on the field production system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to generative artificial intelligence and machine learning. More specifically, the present disclosure relates to an anti-delusion and attribution architecture for enterprise generative artificial intelligence. BACKGROUND

[0002] Generative artificial intelligence (generative AI) refers to a subfield of machine learning that concerns algorithms that can generate new instances of data. These algorithms are typically deep learning models that are trained on large datasets to learn the underlying statistical properties of the data. Generative AI methods employ artificial neural networks to model the statistical properties of large training data. Unlike traditional AI methods that can focus on analyzing existing data based on predefined rules or completing specific tasks, generative AI can use its learned understanding of the data to create entirely new outputs. Traditional generative AI methods have various shortcomings, such as delusions and users being unable to verify or corroborate generative artificial intelligence responses, among others. BRIEF DESCRIPTION OF DRAWINGS

[0003] Figure 1 A schematic diagram depicting an example logical flow of an enterprise generative artificial intelligence system with an anti-delusion and attribution architecture, in accordance with some embodiments.

[0004] Figure 2 A schematic diagram depicting an example resolution and information graph generation process of an enterprise generative artificial intelligence system using an anti-delusion and attribution architecture, in accordance with some embodiments.

[0005] Figure 3 A flow diagram depicting an example anti-delusion and attribution process for an enterprise generative artificial intelligence system, in accordance with some embodiments.

[0006] Figure 4 A schematic diagram depicting an example structure of a response segment generated by an anti-delusion and attribution method for an enterprise generative artificial intelligence system, in accordance with some embodiments.

[0007] Figure 5 A schematic diagram depicting an example enterprise generative artificial intelligence system architecture and environment, in accordance with some embodiments.

[0008] Figure 6A A schematic diagram depicting an example enterprise search graphical user interface and underlying architecture, in accordance with some embodiments.

[0009] Figure 6B A schematic diagram depicting an example enterprise generative artificial intelligence response graphical user interface, in accordance with some embodiments.

[0010] Figure 7 A schematic diagram depicting an example hierarchical architecture and environment of an enterprise generative artificial intelligence system, in accordance with some embodiments.

[0011] Figure 8 A schematic diagram of an example network system for enterprise generative artificial intelligence is depicted according to some embodiments.

[0012] Figure 9 A flowchart is depicted illustrating an example iterative generative artificial intelligence process using unstructured data according to some embodiments.

[0013] Figure 10A A flowchart is depicted illustrating an example iterative generative artificial intelligence process using unstructured data according to some embodiments.

[0014] Figure 10B A flowchart is depicted illustrating an example non-iterative generative artificial intelligence process using unstructured data according to some embodiments.

[0015] Figure 10C A flowchart is depicted illustrating an example non-iterative generative artificial intelligence process using unstructured data according to some embodiments.

[0016] Figure 11 A flowchart is depicted illustrating an example iterative generative artificial intelligence process using unstructured data according to some embodiments.

[0017] Figure 12 A flowchart is depicted for an example enterprise generative artificial intelligence method according to some embodiments.

[0018] Figure 13 A flowchart is depicted illustrating an example enterprise generative artificial intelligence approach using an agent orchestrator architecture, according to some embodiments.

[0019] Figure 14 A flowchart is depicted for an example anti-hallucination and attribution method for an enterprise generative artificial intelligence system according to some embodiments.

[0020] Figure 15 A flowchart is depicted for an example information retrieval method for enterprise generative artificial intelligence methods according to some embodiments.

[0021] Figure 16 A flowchart is depicted illustrating an example method, according to some embodiments, for routing requests to different agent programs and verifying responses to enterprise generative artificial intelligence.

[0022] Figure 17 A schematic diagram of an example computer system for implementing the features disclosed herein, according to some embodiments, is depicted. Detailed Implementation

[0023] Generative AI is an artificial intelligence technology that uses machine learning algorithms to mimic human cognitive intelligence and generate content. This content can take the form of text, audio, video, images, etc. However, traditional generative AI processes often present biased or erroneous information due to machine learning model illusions (or simply hallucinations). In generative AI, artificial hallucinations, fictions, or delusions refer to specific types of output generated by the model that deviate from factual accuracy, ground truth, or expected context. Traditional generative AI may create content that does not correspond to reality or verifiable information; this content is often caused by statistical inaccuracy, model overfitting, and data bias. The generated content may appear reasonable and internally coherent, but lacks accuracy and coherence with the provided cues or surrounding information. Enterprise environments may also include contradictory information, further increasing the likelihood of hallucinations. Hallucinations in enterprise environments can be complex due to the often incomplete and disparate information propagating across various and incompatible enterprise systems. Currently, generative AI systems cannot detect, prevent, or mitigate hallucinations. Traditional generative AI processes also fail to provide users with any mechanism to verify, validate, or confirm generative AI content. Therefore, users cannot know whether generative AI content is accurate or the result of illusion.

[0024] This paper discloses an anti-illusion and attribution architecture for enterprise generative AI systems, which increases the accuracy and reliability of generative AI content (e.g., responses or answers) by detecting, preventing, and mitigating illusions. Furthermore, the anti-illusion and attribution architecture can be added as a standalone tool or module to deployed generative AI systems, allowing the architecture to work seamlessly with the deployed systems without requiring restructuring or redesign. The anti-illusion and attribution architecture can also be deployed with minimal impact on on-site production systems.

[0025] The enterprise generative AI system described in this paper can transform the interaction with enterprise information by fundamentally changing the human-computer interaction (HCI) model used in enterprise software. Enterprises running sensitive workloads in cloud-native, on-premises, or air-gap environments can implement an enterprise generative AI architecture to generate enterprise-wide insights in response to simple, intuitive inputs, leveraging agents that develop and coordinate complex operations, and using tools for rapid information location and retrieval. This enterprise generative AI architecture enables enterprise users to ask open-ended, multi-level, context-specific questions, which are processed using generative AI with machine learning capabilities to understand the request, identify relevant information, and generate new context-specific insights using predictive analytics. The enterprise generative AI architecture supports simplified human-computer interaction with intuitive natural language interfaces and advanced accessibility features for adaptable forms of input, including but not limited to text, audio, video, images, etc.

[0026] Anti-illusion and attribution architectures include anti-illusion and attribution modules that can work with deployed enterprise generative AI systems and components (e.g., retrieval tools, large language models, or other generative AI models) to detect, prevent, and mitigate illusions caused by deployed large language models or other models (e.g., generative AI models, multimodal models). For example, a user might submit a prompt (e.g., a question, query, etc.) to an enterprise generative AI system such as “How many different engineers has John Doe worked with in his engineering department?” This might require the enterprise generative AI system to identify John Doe, identify John Doe’s department, identify the engineers in that department in a third iteration, determine which of those engineers John Doe has interacted with, and then finally combine these results to generate an answer to the query. Such complex queries can introduce illusions at any step in the process of generating the answer. For example, there might be several John Does within the organization, and the system might be prone to illusions in order to determine which John Doe the query involves. Anti-hallucination and attribution architectures can prevent or mitigate this hallucination by providing answers to the anti-hallucination and attribution modules before generating the final answer.

[0027] The anti-illusion and attribution module can parse the response generated by the enterprise generative AI system into several chunks (e.g., several sentences). This module processes these chunks along with the original paragraphs retrieved by the enterprise generative AI system to find relevant paragraphs for each chunk. The module then combines the chunks and the retrieved relevant paragraphs to generate an attributed response. For example, the attributed response may include source identifiers indicating the documents and paragraphs used to generate the response. Thus, the system provides traceable attribution to confirm that the response is reliable and that the model is not hallucinating. If the anti-illusion and attribution module cannot locate any relevant paragraphs, it can determine that the model is hallucinating and rerun the query to find alternative results or notify the user that a reliable result cannot be determined, rather than simply providing the user with a response.

[0028] The enterprise generative AI system described herein can also utilize a combination of agents and tools to efficiently process various inputs received from different data sources (e.g., with different data formats) and return results in a common data format (e.g., natural language). The enterprise generative AI architecture includes an orchestrator agent (or simply orchestrator) for supervising, controlling, and / or otherwise managing many different agents, tools, and / or modules (e.g., anti-hallucination and attribution modules). The orchestrator may include one or more machine learning models and may perform supervisory functions such as routing inputs (e.g., queries, instruction sets, natural language input, or other human-readable or machine-readable input) to specific agents to perform a prescribed set of tasks (e.g., a retrieval request prescribed by the orchestrator to answer a query). The machine learning model may include some or all of the different types or modalities of models described herein (e.g., multimodal machine learning models, large language models, data models, statistical models, audio models, visual models, audiovisual models, etc.). The agent may include one or more multimodal models (e.g., large language models) to use a variety of different tools to perform the prescribed tasks. Different intelligent agents can use a variety of tools to execute and process unstructured data retrieval requests, structured data retrieval requests, and API calls (e.g., for accessing insights from artificial intelligence applications). Tools may include one or more specific functions and / or machine learning models to accomplish a given task (or set of tasks).

[0029] Agents can be adapted to perform differently based on context. Context can relate to a specific domain (e.g., an industry), and agents can employ specific models (e.g., large language models, other machine learning models, and / or data models) already trained on industry-specific datasets (such as healthcare datasets). A particular agent can use a healthcare model when receiving inputs associated with a healthcare environment, and can also be easily and efficiently adapted to use different models based on different inputs or contexts. In fact, some or all of the models described in this paper can be trained for a specific domain, in addition to or instead of a more general purpose. Enterprise generative AI architectures leverage domain-specific models to generate accurate context-specific retrieval and insights.

[0030] The orchestrator manages agents to efficiently process diverse inputs or different parts of those inputs. For example, inputs might require the system to access and retrieve data records from different data sources (e.g., unstructured data stores, structured data stores, and time-series data stores), database tables from different types of databases, and machine learning insights from different machine learning applications. Different agents can handle these requests individually and in parallel, significantly improving computational efficiency.

[0031] An agent can process different data returned by different agents and / or tools. For example, a large language model typically receives input in natural language format. An agent can receive information in non-natural language format (e.g., database tables, images, audio) from a tool and transform it into natural language describing the tool's output in a format understood by the large language model. The model (e.g., a large language model, a multimodal model) can then process this input to generate an initial response, which anti-illusion and attribution modules can validate and / or attribute before providing the final output.

[0032] Figure 1 A schematic diagram 100 depicts an example logic flow of an enterprise generative AI system with an anti-illusion and attribution architecture according to some embodiments. Typically, generative AI models (e.g., large language models, multimodal models) have different capabilities and follow instructions with different performance levels. In one example, an enterprise generative AI system can generate attributions (e.g., source citations) for responses generated by the enterprise generative AI system (e.g., by a large language model) by prompting the generative AI model to cite its source. However, this approach has several drawbacks and limitations.

[0033] More specifically, different models may or may not be able to follow this instruction, and this can depend on the specific model used by the enterprise generative AI system and the task or query received. The formatting of the source hash definition used by the enterprise generative AI system to obtain citations may vary from client to client, which in turn affects citation performance. This approach may also limit the system's ability to cite only from the initially retrieved documents. This could potentially cause the system to ignore important or relevant documents (e.g., because the system does not select a sufficiently high value for the number of retrieved documents that may vary from query to query) and hinder the system from obtaining the potential benefits of leveraging fine-tuned models on the client corpus associated with the enterprise generative AI system.

[0034] Furthermore, while this attribution approach can reduce illusions to some extent, it still does not completely or optimally prevent them. Since generative attribution and illusion prevention are crucial for enterprise generative AI systems, it is unacceptable to rely on luck to make the model follow instructions (e.g., instructions used to cite the source). In fact, the enterprise generative AI systems described in this paper can be model-agnostic, and some models may implement this approach very poorly. That is, this attribution approach is compatible with different types of models, including proprietary or open-source large language models (LLMs), small language models, image generation models, audio generation models, video generation models, multimodal models, etc. Therefore, a method that not only prevents illusions but is also model-agnostic is needed.

[0035] For example, enterprise generative AI systems perform attribution as a separate step after the model has generated its response, rather than prompting the model to cite its source. This addresses the limitations discussed above and also allows the system to rely on information encoded in its corpus while using information from the retrieved paragraphs. The system can combine local and global searches within the corpus to provide better and more informative responses. This enables the system to use new components (e.g., anti-illusion and attribution module 110) as a standalone tool for fact-checking statements and providing confirmatory statements.

[0036] Figure 1 This section provides an overview of current architectures used for anti-hallucination and attribution. Figure 1 In the example, query 102 is received (e.g., a question received from a user via a graphical user interface). Query 102 is passed to retrieval unit 104 to retrieve paragraphs relevant to the query (e.g., using a similarity assessment that generates a corresponding similarity score) (e.g., from a vector store). Retrieval unit 104 passes the retrieved paragraphs through a context processor 106 for forming a prompt, which is then provided to a generative artificial intelligence model 108 for generating a response. In the new design, instead of simply providing the response to the user, the system applies an anti-hallucination and attribution module 110 to this response. The response may initially pass through a response parser 112, which resolves the response into chunks (or segments) to be attributed. These chunks and the originally retrieved paragraphs are processed by the anti-hallucination and attribution (AHA) retrieval unit 114 to locate and retrieve relevant paragraphs for each chunk (e.g., from a vector store). The chunks are then provided together with their corresponding retrieved paragraphs to a response renderer 116, which combines the retrieved paragraphs and chunks to generate an attributed response 118.

[0037] The anti-hallucination and attribution module 110 can parse and chunk the response. To this end, the anti-hallucination and attribution module 110 can decompose or segment the response into sentences and continuous fragments, or can selectively use subsets of the response as sentences and continuous fragments. These fragments or portions are then combined by the anti-hallucination and attribution module 110 into larger chunks until it reaches the maximum number of tokens in each chunk. To compute tokens, the anti-hallucination and attribution module 110 can utilize the same tokenizer selected as used by the embedding model in the anti-hallucination and attribution retrieval module 114. The individual chunks are then fed through the anti-hallucination and attribution retrieval module 114 to generate relevant paragraphs for each chunk. Furthermore, in this stage, the anti-hallucination and attribution module 110 can optionally (e.g., by setting tag values) restrict or filter the paragraphs used to attribute the original paragraphs retrieved at the beginning of the pipeline. These paragraphs, along with the chunks, are then passed to the response renderer 116 to generate an attributed response 118. In this stage, the anti-illusion and attribution module 110 can optionally include several paragraphs to support each chunk. This can be controlled by parameter values. For example, if set to a value greater than 1, the anti-illusion and attribution module 110 will only reference more paragraphs for each chunk if the paragraph's score is greater than a specific threshold of the anti-illusion and attribution retriever 114. The anti-illusion and attribution module 110 can always include at least one reference for each chunk, but it will only include more if the score associated with the paragraph is greater than the threshold. The anti-illusion and attribution module 110 can also provide visualizations such as color coding based on the maximum score of the referenced paragraphs for the chunk.

[0038] In one example, the raw response generated by generative artificial intelligence model 108 may include the following:

[0039] Plot B1 comprises approximately 40-50 acres of property and adjacent marina infrastructure (dredged to include berthing areas and turning waterways) specifically constructed for assembly (“Plot B1”). Development of Plot A is targeted for completion in the second quarter of 2024, with the remaining portion of Phase 1 expected to be completed in 2024 and 2025. Phase 2 is targeted to commence in 2024 and be completed by 2028. The development costs for Phase 1 and Phase 2 are each estimated at approximately US$550 million, excluding hydraulic dredging undertaken by the Department of Transportation. The official agency is issuing bonds to cover a portion of these costs.

[0040] The corresponding attributed response generated by the anti-hallucination and attribution module 110 may include the following:

[0041] Parcel A is targeted for completion in the second quarter of 2024, with the remaining development of Phase 1 anticipated for completion in 2024 and 2025. Phase 2 is targeted for commencement in 2024 and completion by 2028 (from

[143] , score 0.7580196261405945). The development costs for both Phase 1 and Phase 2 are estimated at approximately US$550 million each, excluding hydraulic dredging by the Department of Transportation. The official agency is issuing bonds to cover a portion of these costs (from

[143] , score 0.8340072631835938). Parcel B1 is approximately 40–50 acres of property and adjacent marina infrastructure built specifically for assembly (with dredged berthing areas and turning waterways) (“Parcel B1”) (from

[143] , score 0.7503519654273987). The development target for Plot A is to be completed in the second quarter of 2024, with the remaining portion of Phase 1 expected to be completed in 2024 and 2025. The development target for Phase 2 is to begin in 2024 and be completed by 2028 (from

[143] , score 0.7610742449760437).

[0042] In some embodiments, the anti-hallucination and attribution module 110 can visually label the attributed responses (e.g., highlight, color-code). For example, yellow can be used to highlight (e.g., relative to a threshold) chunks (e.g., sentences) that have a relatively low similarity score. Red can be used to highlight chunks that were not located as relevant paragraphs by the anti-hallucination and attribution retrieval unit 114. Green can be used to highlight (e.g., relative to a threshold) chunks that have a relatively high similarity score, indicating that the chunk was not generated by any hallucination of the generative artificial intelligence model 108.

[0043] In another example, the raw response generated by generative artificial intelligence model 108 may include:

[0044] No, you do not need to pay federal and state taxes on the interest earned on your 2023 Series A bonds, as it is not excluded from gross income for federal income tax purposes and is not included in gross income under current gross income tax laws. However, it is recommended that you consult a tax professional for further guidance.

[0045] The corresponding attributed response generated by the anti-hallucination and attribution module 110 may include the following:

[0046] No, you do not need to pay federal and state taxes on interest earned on 2023 Series A bonds because they are not excluded from gross income for federal income tax purposes and are not included in gross income under current gross income tax laws (from

[161] , score 0.7880104780197144). However, it is recommended to consult a tax professional for further guidance (from

[161] , score 0.4502568244934082).

[0047] Figure 2 A schematic diagram 200 depicts an example parsing and infographic generation process using an anti-illusion and attribution architecture in an enterprise generative artificial intelligence system according to some embodiments. Example data record preprocessing and information retrieval processing can be performed by one or more of the systems and / or subsystems described herein (e.g., enterprise generative artificial intelligence system 802). Infographics can be used to facilitate information retrieval in order to accurately attribute source paragraphs to chunks of the response and prevent illusions.

[0048] Typically, data records can include information in different modalities, such as plain text, tables, images, code, video, and audio. To efficiently and reliably retrieve information for queries, preprocessing and information retrieval processes provide multimodal methods for extracting information from data records.

[0049] Preprocessing and information retrieval can include three stages. The first stage can include parsing and extracting different modalities from these documents. This processing can be done in parallel (or substantially in parallel) for all the different modalities.

[0050] In one example, text and code can be parsed (step 202). Extracting text information as plain text and code can present different challenges depending on the file format and may require utilizing different libraries. Regardless of the data record type, the processing can involve various steps to prepare for other downstream stages. Depending on the file format, the complexity of some of these steps may increase.

[0051] One of these steps may include extracting text information from data record 201 (step 204) so ​​that it can be further parsed. For example, extracting all content that is not an image or table. The output of this step may include (for example, may need to include) all text information from the data record (for example, information that is not descriptive text for images and tables), which can then be used to further separate text and code snippets. The parsing results can be high-fidelity (e.g., not introducing random spaces or odd characters that would ruin the meaning of sentences) and robust to font, size, color, and position relative to other elements on the page.

[0052] Another step may include separating text and code (step 205). The goal of this step is to identify and separate code and plain text. This then allows the system to process these modalities further separately. Additional steps may include chunking and parsing the text and code (step 206). After the text and code modalities have been separated, the goal of this step is to identify, locate, and extract consecutive code segments (step 207), and to chunk the text content in a reasonable and as continuous manner as possible (step 206) (e.g., without cuts in sentences or paragraphs, especially due to breaks between pages, and where possible, chunks with a continuous theme).

[0053] The system (e.g., an enterprise generative AI system 802) can use an object-oriented architecture where classes can exist for both plain text and code. These classes can at least have fields for tracking content, its position within a document, and the number of tokens in the content (e.g., which implies the system should know the tokenizer used for that purpose). For text classes, the system may already be able to track which code snippets have been removed from the chunked content (or the code snippets associated with them). This may already be done as part of the system.

[0054] In another example, tables can be parsed. A key modality in data records is the table. To enable effective information retrieval, the system can first fully locate and identify the table (e.g., because a table may span multiple pages / data records / segments, or may appear in different structures within a single page). See step 206. The system can identify libraries that enable these features and measure their performance in fully identifying tables. Once the table is identified, it can be extracted as an image or as a data frame, etc. (step 209). The system can also extract the table's descriptive text or title, column headers, and potentially (one or more) row indexes, etc., as associated metadata (step 209). Similar to text and code classes, the system can also have table classes that track the extracted table content, its location, title / descriptive text, etc.

[0055] In another example, images can be parsed. Similar to how the system processes tables, the system can also begin by fully identifying and locating images (step 212). The system can include image classes for tracking images in a document, i.e., the content of the image, its location in the document, and its descriptive text / title, along with its associated metadata. The system can also extract images as associated metadata (step 213).

[0056] At the end of the first phase, the system can have several instances of text, code, table, and image classes, thus outlining different modalities across the various data records. After doing this, the system can proceed to the second phase, which involves building infographics for each data record.

[0057] In the second phase, to facilitate efficient information retrieval, the system can represent the information in each data record 201 by generating and / or using an information graph 220. Nodes 230, 241, 243, and 245 of this graph correspond to different instances of the four modalities from Phase One. This depiction is a (directed) bipartite graph, where edges originate from text nodes and extend to all other modal nodes. Establishing these edges is the primary objective of this phase.

[0058] exist Figure 2 In the example, if text node 230 has a reference (e.g., a correlation) with other modal nodes 241, 243, and / or 245, then an edge exists between them. This can be based on direct references to them in text chunks, on proximity, or even on contextual similarity with their descriptive text or headings. After the edges have been identified, the system can track which edges exist between each text node 230 and other modal nodes 241, 243, and / or 245 as part of a text class. Once the system has fully specified the graph, it can use the information graph 220 to design information retrieval processes. For example, agents (e.g., agent 906) and tools (e.g., tool 908) can use the graph to retrieve information. In another example, the anti-illusion and attribution module 110 can use the information graph 220 to attribute sources to responses to generate attributed responses. For this purpose, the information graph 220 may also include access control protocol information (e.g., as defined by the enterprise access control layer 515 and / or the enterprise access control module 918).

[0059] In the third phase, given information graph 220, the system can outline the process for information retrieval. One approach to this begins by embedding the content of text nodes (and / or other text metadata associated with other modalities) (step 229) and storing the embeddings in vector storage 226. Given a query 222, the system can take other embeddings (step 228) and find the most relevant text chunks or text nodes 230 associated with them. This will be the input to the graph. At this point, the system can follow outgoing edges to other modal nodes. The classes associated with these modalities (e.g., code, images, and tables) can have methods that enable the generation of relevant insights given a query. This method can be supported by different approaches, including multimodal models or other tools for understanding and querying specific modalities. These insights can then be combined with text chunks and queries in aggregator 250 to form hints or text bodies for use with query models (e.g., multimodal models, large language models, etc.).

[0060] In one implementation, Figure 2The functionality shown and described herein can be performed by a chunking module (e.g., chunking module 910) and / or an embedding generator module (e.g., embedding generator module 912). For example, steps 204-213 can be performed by the chunking module, and steps 228-229 can be performed by the embedding generator module. In some embodiments, the aggregator 250 includes a portion of an orchestrator module (e.g., orchestrator module 904) and / or an understanding module (e.g., understanding module 916).

[0061] Figure 3 A flowchart 300 depicts an example anti-illusion and attribution process for an enterprise generative artificial intelligence system. Typically, the anti-illusion and attribution process (e.g., performed by anti-illusion and attribution module 110) is a completely independent process (and module) that can be used to verify and validate any response (e.g., a response from a generative artificial intelligence model or a human-generated response), regardless of whether it is accompanied by inline references.

[0062] In step 302, a response is received (e.g., from an enterprise generative AI system, a generative AI model, a human user, etc.). The response may be received by an anti-illusion and attribution module (e.g., anti-illusion and attribution module 110). In step 304, the anti-illusion and attribution module determines whether the response already includes references (e.g., provided by the generative AI model of the enterprise generative AI system during the response generation process). If the anti-illusion and attribution module determines that the response is associated with references (e.g., inline references), a segmentation of the response can be extracted by segmentation based on available references (step 306). If the anti-illusion and attribution module determines that no references exist in the response, the anti-illusion and attribution module performs a general segmentation (step 308). More specifically, the general segmentation is typically based on the number of lexical units and on sentences in the response. More specifically, for each segment, the anti-illusion and attribution module may arrange as many sentences as possible in the segment until it reaches the segment's maximum lexical unit limit.

[0063] In step 310, the anti-hallucination and attribution module assigns a set of sources based on a relevance or similarity search. The anti-hallucination and attribution module may filter source segments (or sources) based on the obtained scores. Filtering may compare the obtained scores to one or more thresholds. For example, if the score is equal to or higher than the threshold, the associated source segment may be attributed to the segment. In another example, if the score is lower than the threshold, the associated source segment is not attributed to the segment. The similarity search may be an extension of the similarity score (e.g., "classical" cosine) and / or a function that can be overridden by the end user.

[0064] The anti-hallucination and attribution module uses a mapping between response segments and sources with relevance scores. The anti-hallucination and attribution module can also quantify / extract credibility scores associated with each source. The anti-hallucination and attribution module can then segment the sources (step 312). More specifically, the anti-hallucination and attribution module can use the same segmentation logic used for segmenting responses. The anti-hallucination and attribution module calculates pairwise relevance / similarity scores between response and source segments (step 314).

[0065] In step 316, the anti-illusion and attribution module further enhances the mapping using confirmation scores / labels associated with each pair within the four categories of contradictory, supportive, neutral, and suggestive. These can then be used to further prune the keys and values ​​in the mapping. Supportive labels can be added to potentially cover the shortcomings associated with general segmentation.

[0066] At this point, the anti-hallucination and attribution modules have generated and quantified the relationship between the response and source segments. This is in Figure 4 Example in Figure 4 This illustrates all the different scores that define the relationship between response and source segments. These can then be used to quantify a single score associated with each response segment (e.g., quantifying its assertiveness within the source) and a single score associated with each source segment (e.g., within the context of the corresponding response segment). See step 318. In some embodiments, the process can be adapted to user-defined heuristics for calculating some or all of these scores and for calculating a source credibility score based on available metadata of the source during inference or when ingesting the source.

[0067] Figure 4 A schematic diagram 400 depicts an example structure of response segments generated by anti-illusion and attribution methods for enterprise generative artificial intelligence systems, according to some embodiments. More specifically, Figure 4 This includes response segment 402, source 404, relevance score 406 for source 404, credibility score 408, source segment 410, relevance score 412 for source segment, and confirmation score 414.

[0068] Figure 5 A schematic diagram of an example enterprise generative artificial intelligence system architecture and environment 500 according to some embodiments is depicted. Figure 5In the example, system architecture and environment 500 include an enterprise generative artificial intelligence system 502, an enterprise system 504, an external system 506, a domain model 508, a vector data store 526, an embedded model data store 524, and an enterprise access control layer 515. In some embodiments, queries include natural language queries received through a graphical user interface. In some embodiments, one or more enterprise datasets include any of documents, document segments, and insights generated by one or more artificial intelligence applications. In some embodiments, relevance scores are each associated with a corresponding portion of one or more enterprise datasets, and wherein each relevance score is determined relative to other corresponding portions of one or more enterprise datasets.

[0069] Each data model in a plurality of data models may correspond to a different data domain in a plurality of different data domains. In some embodiments, each data model represents the corresponding relationships and attributes of the corresponding different data domains in a plurality of different data domains. The corresponding relationships and attributes include any of data types, data formats, and industry-specific information. In some embodiments, the natural language output includes a summary of at least one of the various portions of one or more enterprise datasets associated with relevance scores.

[0070] Typically, an enterprise generative AI system 502 can function to securely query and process enterprise data and applications across different domains of the enterprise information environment. This can be referred to as generative enterprise search (or simply enterprise search). As shown, the enterprise generative AI system 502 can receive questions 516 (e.g., input, prompts, natural language queries, instructions, etc.). Typically, the enterprise generative AI system 502 can process queries using a large language model 520 and a retrieval model 522. More specifically, the enterprise generative AI system 502 can use the large language model 520 to interpret, understand, and / or parse queries. The retrieval model 522 can interact with the large language model 520 and the domain model 508 to retrieve data records (e.g., documents, images, application outputs, AI insights, and objects, etc.) across different domains using domain-specific data models 512. Therefore, the enterprise generative AI system 502 can use the large language model 520, the retrieval model 522, and the domain model 508 to generate accurate, reliable, and secure enterprise search results.

[0071] Enterprise generative AI system 502 can use connectors 510, data models 512, and various persistent storage mechanisms and technologies 514 to facilitate the ingestion and persistence of enterprise system data from enterprise system 504 and / or external system data from external system 506 (e.g., systems outside the enterprise information environment). Enterprise system 504 may include CRM systems, EAM systems, and / or ERP systems, etc., and connectors can facilitate data ingestion from different data sources (e.g., Oracle systems and / or SAP systems, etc.). In some embodiments, data model 512 provides attributes, relationships, and / or functions associated with a specific domain. For example, a domain may include aerospace domain 512-1, energy domain 512-2, and / or defense domain 512-3, etc. Domain model 508 enables enterprise generative AI system 502 to provide domain-specific results without compromising the security or integrity of the underlying enterprise data, systems, and applications.

[0072] Furthermore, the enterprise generative AI system 502 utilizes or manages an enterprise access control layer 515, which can provide numerous technical benefits. In some embodiments, the enterprise access control layer 515 facilitates the separation of underlying enterprise information (e.g., enterprise data, applications, systems) from the large language model 520 and / or other machine learning models of the enterprise generative AI system 502. Therefore, the enterprise generative AI system 502 can provide domain-specific deterministic results without having to train the large language model 520 and / or other machine learning models on such enterprise information, where such training could lead to the numerous problems mentioned above (e.g., information leakage, hallucinations).

[0073] In some embodiments, the enterprise generative AI system 502 may use an enterprise access control layer 515 to implement additional enterprise controls. For example, the enterprise information environment may include users and systems with different enterprise permission levels. The enterprise access control layer 515 may ensure that responses or outputs comply with access and security protocols. The enterprise generative AI system 502 protects information to allow user access based on permissions, profiles, and controls. In one example, the enterprise access control layer 515 may filter information defined by the retrieval model 522 before it is processed by the large language model 520 or before a response or other output is presented. More specifically, the enterprise access control layer 515 may filter data sources, data records, and / or other elements of the enterprise information environment so that query responses (or those supporting traceability references) do not include information that the user is not authorized to access.

[0074] Enterprise generative AI systems can also perform similar functionality based on the context of the user submitting the query and / or the system. For example, managers and engineers can submit the same query (e.g., “Which projects are overdue?”), and enterprise generative AI system 502 can use contextual information (e.g., user roles, permissions, and domains associated with the user) to provide a response that is context-based in both substance (e.g., providing information about the smoke project tailored to a specific requester) and / or presentation of the response (e.g., engineers may receive more detailed technical information, while managers may receive less technical detail).

[0075] In some embodiments, the enterprise generative AI system 502 may use contextual information (e.g., contextual metadata) along with data record embeddings to crawl, index, and / or map a corpus of data records (e.g., data records from one or more enterprise systems or environments) to provide access control (e.g., role-based access), improved data record identification and retrieval, and mapping relationships between data records. In one example, contextual information may prevent certain users from accessing (e.g., viewing, retrieving) certain data records and improve similarity assessment used in retrieval operations (e.g., in generative AI processes).

[0076] In some implementations, the enterprise generative AI system 502 can generate embeddings based on the embedding model of the embedding model data store 524 and the content of the data records. In some implementations, the embeddings can be represented by one or more vectors that can be stored in the vector data store 526. In some implementations, the retrieval model 522 can use the embeddings to retrieve relevant data records and perform similarity or relevance assessments or other aspects of retrieval operations. As used herein, data records can include unstructured data records (e.g., documents and text data stored on a file system in formats such as PDF, DOCX, MD, HTML, TXT, PPTX, image files, audio files, video files, and application outputs), structured data records (e.g., database tables or other data records stored according to a data model or type system), time-series data records (e.g., sensor data, AI application insights), and / or other types of data records (e.g., access control lists).

[0077] Figure 6A An example enterprise search graphical user interface 600 and underlying architecture 602-608 according to some embodiments are depicted. Typically, Figure 6A and Figure 6B The graphical user interface described includes a human-machine interface (HMI) for receiving natural language queries and responding to those queries by presenting relevant information from the enterprise information environment. AlthoughFigure 6A or Figure 6B Not shown, but the relevant information may also include visualizations, predictive analytics, control commands and / or other information obtained or generated by the enterprise's generative artificial intelligence system.

[0078] exist Figure 6A In the example, the diagram includes an enterprise query input interface 600 and frameworks 602-608, which unify information access methods and application operations across traditional and new enterprise applications and a growing number of data sources in various enterprise environments. The enterprise generative artificial intelligence system described in this paper can harmonize access to information and increase the usability of complex application operations while adhering to enterprise security and privacy controls. The framework uses machine learning techniques (e.g., generative AI algorithms and models) to navigate enterprise information and applications, understand organization-specific context queues (e.g., abbreviations, nicknames, and jargon), and locate the information most relevant to the request (e.g., queries, questions, prompts, etc.). For example, this can lower the learning curve and reduce the steps users must take to access information, thereby popularizing information access currently hindered by the complexity and domain expertise required by traditional enterprise information systems.

[0079] exist Figure 6A In the example, the enterprise query input interface 600 includes graphical user interface elements configured to receive various inputs, such as natural language queries. For example, a user could ask the system, "Why did my brakes fail?", and the enterprise generative AI system could utilize a natural language processing component 602 and one or more generative pre-trained transformers 604 (e.g., large language models and / or other machine learning models) to process the query, securely, accurately, and reliably generating answers based on various enterprise data storage and applications 608. More specifically, the natural language processing component 602 and the generative pre-trained transformers 604 can generate one or more new queries 606 to process the initial user query. For example, the first new query could include an SQL query for retrieving data records from a relational database management system, and the second query could include other types of queries (e.g., instruction sets) for executing one or more applications, such that the results of the application execution can be returned and used by the system to generate an enterprise search response. An example enterprise search response interface is shown in... Figure 6A and Figure 6B It is shown in the figure and described below.

[0080] In some embodiments, Figure 6A and Figure 6B The enterprise search interface shown and described herein can be generated by the enterprise generative artificial intelligence system described herein, and frames 602-608 can represent the architecture and environment of the enterprise generative artificial intelligence system described herein.

[0081] Figure 6B An example enterprise generative AI responsive graphical user interface 650 according to some embodiments is depicted. In some embodiments, the enterprise generative AI responsive graphical user interface 650 may be generated at least in part by the enterprise generative AI system described herein. Figure 6B In the example, the enterprise generative AI response graphical user interface 650 includes an enterprise search query input section 652, a generative enterprise search results section 654, and an interactive query section 656.

[0082] The enterprise search query input section 652 presents the enterprise search query 658. In some implementations, the query can be entered through input section 652, but it can also be entered through other interfaces (e.g., Figure 6A The input is shown in interface 600 and presented in input section 652 as part of the generative enterprise search response.

[0083] The generative enterprise search results section 654 includes the type of generative AI response 660, the status of the generative AI response 665, the generative AI enterprise search results 666, the source data section used to generate the response 668, source identification 669, and the generative AI response feedback element 670. Figure 6B In the example, response type 660 is a summary, but other response types that an enterprise generative AI system can generate may exist. Response status indicates the state of the response. Response status can include, for example, processing a query, processing a query, evaluating a metric, searching a document, generating an answer (e.g., results), generating a visualization (e.g., a time-series visualization for presentation in the response's graphical user interface), and / or ending generation (e.g., as...). Figure 6B (As shown).

[0084] Source data section 668 includes at least a portion of the information from the source data used to generate the response. This can, for example, enable the verification of the response without requiring independent verification. Source identification 669 identifies the source data records used to generate the response. For example, source identification may indicate entity name, domain type, description or name of the data record (e.g., service manual, user manual, and / or technical manual, etc.) and / or type of data record (e.g., document or more specifically, PDF document), etc. This can also provide traceability and enable users to trust the response.

[0085] The response feedback section 670 enables users to provide feedback on the response (e.g., positive or negative feedback). Enterprise generative AI systems can, for example, use the received feedback to improve themselves (e.g., through reinforcement learning).

[0086] The interactive query section 656 allows the user to input additional related queries (e.g., "follow-up" questions) via the interactive input section 657. Figure 6B In the example, the interactive query section 656 includes a chat interface, but other interfaces may use different interactive query sections. The interactive query section 656 also includes a system-generated message 674 prompting the user to ask follow-up questions, and the user can provide additional related queries 676. The interactive query section 656 may also include a status section 678 indicating the processing status of the additional related queries 676. The status may include, for example, a query being processed (e.g., as...). Figure 6B (as shown), it processes queries, evaluates metrics, searches documents, generates answers (e.g., results), generates visualizations (e.g., time-series visualizations for presentation in a responsive graphical user interface) and / or terminates generation.

[0087] Time series data is a list of data points ordered chronologically, representing how values ​​of data related to a specific issue (such as inventory levels, equipment temperature, financial value, or customer transactions) change over time. Time series data provides historical information that can be analyzed by generative machine learning algorithms to generate and test predictive models. Example implementations apply cleaning, normalization, aggregation, and combination to time series data representing the state of a process over time to identify patterns and associations that can be used to create and evaluate predictions applicable to future behavior.

[0088] Figure 7 A schematic diagram 700 depicts an example layered architecture and environment of an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802) according to some embodiments. Figure 7 In the example, the enterprise generative AI system architecture and environment includes a hierarchical structure of layers. More specifically, the hierarchical structure of layers includes an input layer 702, a supervision layer 710, an agent layer 720, an agent and tool layer 730, a tool and data model layer 750, and an outer layer 780. It will be understood that these layers are shown by way of example, and other examples may include any number of such layers (e.g., any number of layers 720 and 730).

[0089] Input layer 702 represents the layer in an enterprise generative AI system architecture that receives input (e.g., queries, complex inputs, and / or instruction sets) from users or the system. For example, the interface module of an enterprise generative AI system can receive input.

[0090] Supervised layer 710 represents a subsequent layer in the enterprise generative artificial intelligence system architecture, comprising one or more large language models (e.g., an orchestrator module) that can be developed to respond to inputs received in input layer 702. The plans may include a defined set of tasks (e.g., retrieval tasks and API call tasks). In one example, supervised layer 710 may provide the preprocessing and post-processing functionality described herein, as well as the orchestrator and understanding modules described herein. Supervised layer 710 may coordinate with one or more of subsequent layers 720-780 to perform the defined set of tasks.

[0091] Agent layer 720 represents a layer in an enterprise generative artificial intelligence system architecture that includes agents capable of performing a defined set of tasks. Figure 7 In the example, agent layer 720 includes machine learning insight agent 722, information retrieval agent 724, dashboard agent 726, and optimizer agent 728. Each of agents 724-728 may include a large language model that provides reasoning functionality for completing its assigned portion of a prescribed set of tasks. More specifically, agents 724-728 may instruct agents and tools in any number of subsequent layers (e.g., layer 730) to perform tasks. For example, machine learning insight agent 722 may instruct text processing tool 732 to perform text processing tasks (e.g., transforming AI application output into natural language), instruct image processing tool 734 to perform image processing tasks (e.g., generating a natural language summary of images from AI application output), instruct time series tool 736 to obtain summarized time series data (e.g., time series data from AI application output), and instruct API tool 738 to perform API call tasks (e.g., performing an API call to trigger or access the AI ​​application).

[0092] Information retrieval agent 724 can collaborate and / or coordinate with several different agents to perform retrieval tasks. For example, information retrieval agent 724 can instruct unstructured data retrieval agent 740 to receive unstructured data records, instruct structured data retrieval agent 742 to retrieve structured data records, and instruct type system retrieval agent 744 to obtain one or more data models (or subsets of data models) and / or types from a type system. The type system provides compatibility across different data formats, protocols, operating languages, different systems, etc. Types can encapsulate some or all of the data formats from the different types or modalities described herein (e.g., multimodal, text, encoded, language, statistical, audio, visual, audiovisual, etc.). For example, data models can include various different types (e.g., in tree or graph structures), and each type can describe data fields, operations, and functions, etc. Types can represent different objects (e.g., real-world objects, such as machines or sensors in a factory) or systems (e.g., computing clusters, enterprise data storage, file systems), and each type can include a large language model context that provides context for designing or updating plans for a large language model. For example, the context can include a natural language summary or description of the type (e.g., a description of the represented object, its relationship to other types or objects, and associated methods and functions, etc.). Types can be defined in natural language format for efficient processing by the large language model. The type system retrieval agent 744 can traverse the data model 754 to retrieve subsets of the data model 754 and / or types of the data model 754. The structured data retrieval agent 742 can then use this retrieved information to efficiently retrieve structured data from structured data sources (e.g., structured data sources structured or modeled according to the data model 754).

[0093] The dashboard agent 726 can be configured to generate one or more visual and / or graphical user interfaces, such as dashboards. For example, the dashboard agent 726 can execute tools 752-5 and 752-6 to generate a dashboard based on information retrieved by other agents and / or information output by other agents (e.g., a natural language summary of associated tool outputs).

[0094] The optimizer agent 728 can be configured to execute various prescriptive analytic functions and mathematical optimizations 752-7 to assist in the computation of answers to various problems. For example, a large language model 706 can use the optimizer agent 728 to generate plans, determine prescriptive task sets, and determine whether more information is needed to generate the final result.

[0095] The tools and data model layer 750 is designed to represent the layer of an enterprise generative AI system architecture, including tools 752 and a data model 754. Agents 740-742 can execute tools 752 to retrieve information from various applications and data stores 782 in an external layer 780 (e.g., external to the enterprise generative AI system). Tools 752 may include connectors that can connect to systems and data stores external to the enterprise generative AI system.

[0096] Figure 8 A schematic diagram 800 depicts an example network system for enterprise generative artificial intelligence according to some embodiments. Figure 8 In the example, the network system includes an enterprise generative artificial intelligence system 802, enterprise systems 804-1 to 804-N (each a separate enterprise system 804, collectively referred to as enterprise system 804), external systems 806-1 to 806-N (each a separate external system 806, collectively referred to as external system 806), and a communication network 808.

[0097] Enterprise generative AI system 802 can function to iteratively and non-iteratively generate machine learning model inputs and outputs to determine a final output (e.g., an "answer" or "result") in response to initial input (e.g., provided by a user or another system). In some embodiments, the functionality of enterprise generative AI system 802 may be provided by one or more servers (e.g., cloud-based servers) and / or other computing devices. Enterprise generative AI system 802 may be implemented using type systems and / or model-driven architectures. Enterprise generative AI system 802 may also include (e.g., added to and / or connected to enterprise generative AI system 802 after it has been deployed) an anti-illusion and attribution module 110.

[0098] In various implementations, an enterprise generative AI system 802 can provide a wide range of technical features, such as effectively processing and generating complex natural language inputs and outputs, generating synthetic data (e.g., supplementing customer data obtained during the onboarding process or otherwise filling data gaps), generating source code (e.g., application development), generating applications (e.g., AI applications), providing cross-domain functionality, and numerous other technical features not offered by traditional systems. As used herein, synthetic data can refer to content generated on the fly as part of the processes described herein (e.g., by a large language model). Synthetic data can also include transient content that is not retrieved (e.g., temporary data not existing in a database), as well as combinations of retrieved information, queried information, and / or model outputs.

[0099] In some embodiments, the enterprise generative AI system 802 may provide and / or enable intuitive, uncomplicated interfaces to rapidly execute complex user requests with improved access, privacy, and security enforcement. The enterprise generative AI system 802 may include a human-machine interface for receiving natural language queries and, in response to these queries, presenting relevant information with predictive analytics from the enterprise information environment. For example, the enterprise generative AI system 802 may understand the language, intent, and / or context of a user's natural language query. The enterprise generative AI system 802 may execute the user's natural language query to identify relevant information from the enterprise information environment to present to the human-machine interface (e.g., in the form of an "answer").

[0100] In some embodiments, the generative AI model of the enterprise generative AI system 802 (e.g., a large language model of an orchestrator) can interact with agents (e.g., a retrieval agent, a retrieval agent) to retrieve and process information from various data sources. For example, the data sources may store data records and / or segments of data records that can be identified by the enterprise generative AI system 802 based on embedded values ​​(e.g., vector values ​​associated with data records and / or segments). Data records may include tables, text, images, audio, video, code, and / or application outputs (e.g., predictive analytics and / or other insights generated by AI applications), etc.

[0101] In some embodiments, the enterprise generative AI system 802 can generate context-based synthetic outputs based on information retrieved from one or more retrieval models. For example, a retrieval model (e.g., a retrieval agent or retrieval model) can provide additional retrieved information to a large language model to generate additional context-based synthetic outputs until context validation criteria are met. Once the validation criteria are met, the enterprise generative AI system 802 can output the additional context-based synthetic outputs as a result or set of instructions (collectively, a “response”). Context validation criteria may include thresholds for identifying source material from enterprise data systems used to validate the response.

[0102] In various embodiments, the enterprise generative AI system 802 provides transformed, context-based intelligent generation results. For example, the enterprise generative AI system 802 can use a natural language interface to process input from enterprise users to quickly locate, retrieve, and present relevant data across the entire corpus of the enterprise's information systems.

[0103] As discussed elsewhere in this document, the enterprise generative AI system 802 can handle both machine-readable input (e.g., compiled code, structured data, and / or other types of formats that can be processed by a computer) and human-readable input. Input can also include complex input, such as input including "and," "or," and / or input including different types of information to satisfy the input (e.g., data records, text documents, database tables, and AI insights). In one example, complex input could be "How many different engineers has John Doe worked with within his engineering department?" This might require the enterprise generative AI system 802 to identify John Doe in the first iteration, John Doe's department in the second iteration, the engineers within that department in the third iteration, and then, in the fourth iteration, which of these engineers John Doe has interacted with, and finally combine these results or portions thereof to generate a final answer to the query. More specifically, the enterprise generative AI system 802 can use portions of the results from each iteration to generate contextual information (or simply context), which can then inform subsequent iterations.

[0104] Enterprise system 804 may include enterprise applications (e.g., artificial intelligence applications), enterprise data storage, client systems, and / or other systems within the enterprise information environment. As used herein, the enterprise information environment may include one or more networks (e.g., cloud, on-premises, air-gap, or others) of enterprise systems (e.g., enterprise applications, enterprise data storage) and client systems (e.g., computing systems for accessing enterprise systems). Enterprise system 804 may include different computing systems, applications, and / or data storage along with enterprise-specific requirements and / or characteristics. For example, enterprise system 804 may include access and privacy controls. For example, an organization's private network may include an enterprise information environment that includes various enterprise systems 804. Enterprise system 804 may include, for example, CRM systems, EAM systems, ERP systems, FP&A systems, HRM systems, and SCADA systems. Enterprise system 804 may include or utilize artificial intelligence applications, and artificial intelligence applications may utilize enterprise systems and data. Enterprise system 804 may include (e.g., data flows and management of one or more organizations) with different processing capabilities and may provide access to the enterprise's systems and users while preventing access from other systems and / or users. It should be understood that in some embodiments, references to an enterprise information environment may also include an enterprise system, and references to an enterprise system may also include an enterprise information environment. In various embodiments, the functionality of the enterprise system 804 may be provided by one or more servers (e.g., cloud-based servers) and / or other computing devices.

[0105] External system 806 may include applications, data storage, and systems outside the enterprise information environment. In one example, enterprise system 804 may be part of an organization's enterprise information environment, which cannot be accessed by users or systems outside that enterprise information environment and / or organization. Therefore, example external system 806 may include Internet-based systems outside the enterprise information environment, such as news media systems and / or social media systems. In various embodiments, the functionality of external system 806 may be provided by one or more servers (e.g., cloud-based servers) and / or other computing devices.

[0106] Communication network 808 may represent one or more computer networks (e.g., LAN, WAN, air-gap network, and / or cloud-based network, etc.) or other transmission media. In some embodiments, communication network 808 may provide communication between systems, modules, engines, generators, layers, agents, tools, orchestrators, data storage, and / or other components described herein. In some embodiments, communication network 808 includes one or more computing devices, routers, cables, buses, and / or other network topologies (e.g., mesh, etc.). In some embodiments, communication network 808 may be wired and / or wireless. In various embodiments, communication network 808 may include a local area network (LAN), a wide area network (WAN), the Internet, and / or may be one or more networks that are public, private, IP-based, non-IP-based, air-gap, etc.

[0107] Figure 9 A schematic diagram 900 depicts an example enterprise generative artificial intelligence system 802 according to some embodiments. Figure 9In the example, the enterprise generative artificial intelligence system 802 includes a management module 902, an orchestrator module 904, a retrieval agent module 906-1, an unstructured data retrieval agent module 906-2, a structured data retrieval agent module 906-3, a type system retrieval agent module 906-4, a machine learning insight module 906-5, a time series processing agent 906-6, an API agent module 906-7, a mathematical agent module 906-8, a visualization agent module 906-9, a code generation agent module 906-10, an unstructured data retrieval tool 908-1, a structured data retrieval tool 908-2, a text processing tool module 908-3, an image processing tool module 908-4, a time series processing tool module 908-5, an API tool module 908-6, a visualization tool module 908-7, and an optimizer tool module. 908-8, Filter Tool Module; 908-9, Projection Tool Module; 908-10, Grouping Tool Module; 908-11, Sorting Tool Module; 908-12, Restriction Tool Module; 908-13, Code Generation Tool Module; 908-14, Chunking Module; 910, Embedding Generator Module; 912, Crawling Module; 914, Understanding Module; 916, Enterprise Access Control Module; 918, Artificial Intelligence Traceability Module; 920, Parallelization Module; 922, Model Generation Module; 924, Model Deployment Module; 926, Model Optimization Module; 928, Interface Module; 930, Communication Module; 932, Anti-illusion and Attribution Module; 934, (one or more) Vector Data Storage; 940, (one or more) Model Registry Data Storage; 950, (one or more) Feature Data Storage; 960, and (one or more) Enterprise Generative Artificial Intelligence System Data Storage; 970.

[0108] In some embodiments, the chunking module 910, the embedding generator module 912, the crawling module 914, the vector data storage 940 (e.g., its storage embedding), and a portion of the enterprise generative artificial intelligence system data storage 970 (e.g., segmented data storage) may include an intelligent crawling and chunking subsystem (e.g., intelligent crawling and chunking subsystem 120).

[0109] Management module 902 can function to (e.g., create, read, update, delete, or otherwise access) data associated with enterprise generative AI system 802. Management module 902 can store or otherwise manage or reside in any data store 940-970 and / or in one or more other local and / or remote data stores. It will be understood that a data store can be a single data store local to enterprise generative AI system 802 and / or multiple data stores remote to enterprise generative AI system 802. In some embodiments, the data stores described herein include one or more local and / or remote data stores. Management module 902 can operate manually (e.g., through user interaction with a GUI) and / or automatically (e.g., through one or more triggers of modules 904-930). As with the other modules described herein, some or all of the functionality of management module 902 can be included in and / or collaborate with one or more other modules, systems, and / or data stores.

[0110] Orchestrator module 904 may function to generate and / or execute one or more orchestrator agents (or simply orchestrators). The orchestrator may orchestrate, supervise, and / or otherwise control agents 906. In some implementations, the orchestrator includes one or more large language models. The orchestrator may interpret input, select appropriate agents for handling queries and other inputs, and route interpreted inputs to the selected agents. The orchestrator may also perform various supervisory functions. For example, the orchestrator may implement stopping conditions to prevent the understanding module from getting stuck in an infinite loop during an iterative, context-based generative artificial intelligence process. The orchestrator may also include one or more other types of models to process (e.g., transform) non-textual inputs. In addition to or in place of some or all of the large language models used in the agents and / or modules described herein, other models (e.g., other machine learning models, translation models) may also be included.

[0111] In some embodiments, the orchestrator can process data received from various data sources in different formats (which may be processed using vectorized data with the help of natural language processing (NLP) (e.g., using word segmentation, stemming, lemmatization, and normalization, etc.)) and can generate pre-trained transformers that are fine-tuned or retrained on specific data tailored for the relevant data domain or data application (e.g., SaaS applications, legacy enterprise applications, artificial intelligence applications). Further processing may include data modeling feature checks and / or machine learning model simulations to select one or more appropriate analytics channels. Example data objects may include accounts, products, employees, suppliers, opportunities, contracts, locations, digital portals, geolocated manufacturers, Supervisory Control and Data Acquisition (SCADA) information, Open Manufacturing System (OMS) information, inventory, supply chain, bill of materials, transportation services, maintenance logs, and service logs.

[0112] In some embodiments, orchestrator module 904 can utilize various components as needed to inventory or generate objects (e.g., components, functionalities, and / or data, etc.) using rich descriptive metadata to dynamically generate embeddings for developing knowledge across broad data domains (e.g., documents, tabular data, insights derived from AI applications, web content, or other data sources). In example implementations, orchestrator module 904 may leverage some or all of the components described herein. Thus, for example, orchestrator module 904 can facilitate storage, transformation, and communication to facilitate data processing and embedding. In some implementations, the orchestrator module can create embeddings for multiple data types across multiple vertical industries and knowledge domains, and even enterprise-specific knowledge. Knowledge can be explicitly modeled and / or learned by orchestrator module 904, agent 906, and / or tool 908. In the example, orchestrator module 904 (and / or chunking module 910) generates embeddings that are translated or transformed to be compatible with understanding module 916.

[0113] In some embodiments, orchestrator 904 can be configured to enable different data domains to operate or engage with components of the enterprise generative artificial intelligence system 802. In one example, orchestrator module 904 can embed objects from a specific data domain as well as across data domains, applications, data models, parsing byproducts, AI predictions, and knowledge stores to provide robust search functionality without requiring specialized programming to each different data domain or data source. For example, orchestrator module 904 can create multiple embeddings for a single object (e.g., the object can be embedded in a domain-specific or application-specific context). In some embodiments, chunking module 910, together with orchestrator module 904, can manage data domains to embed objects from those domains within the enterprise information system and / or environment. In some embodiments, orchestrator 904 can cooperate with chunking module 910 to provide the embedding functionality described herein.

[0114] In some embodiments, orchestrator module 904 enables agent 906 to perform data modeling to translate raw source data formats into target embeddings (e.g., objects and / or types). Data formats may include some or all of the different types or modalities described herein (e.g., multimodal, text, encoded, language, statistical, audio, visual, audiovisual, etc.). In example implementations, orchestrator module 904 and / or enterprise generative AI system 802 typically employ a model-driven architecture type system for data modeling to translate raw source data formats into target types. The knowledge base of enterprise generative AI system 802 and generative AI models can create the ability to integrate or combine insights from different AI applications.

[0115] As discussed elsewhere in this document, in addition to human-readable input, the enterprise generative AI system 802 can also handle machine-readable input (e.g., compiled code, structured data, and / or other types of formats that can be processed by a computer). Input can also include complex input, such as input including "and", "or", or input including different types of information to satisfy the input (e.g., text documents, database tables, and AI insights). The orchestrator 904 can (e.g., by using a large language model) decompose these complex inputs for processing by multiple agents 906 (e.g., in parallel).

[0116] As discussed above, orchestrator module 904 can function to perform and / or otherwise handle various supervisory functions. In some implementations, orchestrator module 904 can enforce conditions (e.g., stopping conditions, resource allocation, and / or priority ordering). For example, a stopping condition can indicate the maximum number of iterations (or hops) that can be performed before the iterative processing terminates. Stopping conditions and / or other features managed by orchestrator module 904 can be included in large language model hints and / or in the large language model of the orchestrator and / or the understanding module 916 discussed below. In some embodiments, stopping conditions can ensure that the enterprise generative AI system 802 will not get stuck in an infinite loop. This feature can also allow the enterprise generative AI system 802 to have the flexibility to have different numbers of iterations for different inputs (e.g., as opposed to having a fixed number of hops). In another example, orchestrator module 904 can allocate resources based on computational conditions, such as virtualization or load balancing. In some implementations, the orchestrator module 904 and / or agent 906 include a model that can convert (or transform) images, database tables, and / or other non-text inputs into text format (e.g., natural language).

[0117] In some embodiments, the orchestrator module 904 may function to collaborate with agents 906 (e.g., retrieval agent module 906-1, unstructured data retrieval agent module 906-2, and structured data retrieval agent module 906-3) to iteratively and non-iteratively process input to determine output results or answers, determine the context and principles used to inform subsequent iterations, and determine whether a large language model (e.g., that of orchestrator 904 and / or understanding module 916) requires additional information to determine the answer. For example, orchestrator module 904 may receive a query and instruct agent 906-1 to retrieve relevant information. Retrieval agent module 906-1 may then select unstructured data retrieval agent module 906-2 and / or structured data retrieval agent module 906-3 based on whether orchestrator module 904 wants to retrieve structured or unstructured data records. The appropriate agent 906 may select a suitable tool and provide the tool output to orchestrator module 904 and / or understanding module 916 to determine the final result.

[0118] Orchestrator 904 can also select and swap models as needed. For example, in addition to before or after runtime, orchestrator 904 can also replace the models (e.g., data models, large language models, machine learning models) of the enterprise generative AI system 802 during runtime or during runtime. For example, orchestrator 904, agent 906, and understanding module 916 can use a specific set of machine learning models for one domain and other models for different domains. Orchestrator 904 can select and use appropriate models for a given domain and / or input.

[0119] In some embodiments, orchestrator 904 can combine (e.g., splice) outputs / results from various agents to create a unified output. For example, one or more agent modules 906 may acquire / output a document (or its (one or more) segments) or related information (e.g., text summaries or translations), and another agent module 906 may acquire / output a database table, etc. Orchestrator 904 can then use one or more machine learning models (e.g., a large language model and / or another machine learning model) to combine the outputs / results into a unified output (e.g., with a common data format, such as natural language, etc.).

[0120] In some implementations, orchestrator 904 preprocesses the input (e.g., initial input) before it is sent to one or more agents 906 for processing. For example, orchestrator 904 may transform a first part of the input into an SQL query and send it to unstructured data retrieval agent module 906-2, and a second part of the input into an API call and send it to API agent module 906-7, etc. In another example, such transformation functionality may be performed by agent 906, either in place of orchestrator 904 or in addition to orchestrator 904.

[0121] The orchestrator module 904 can function to process, extract, and / or transform different types of data (e.g., text, database tables, images, videos, and / or code). For example, the orchestrator module 904 can accept a database table as input and transform it into natural language describing the database table. This natural language can then be provided to the understanding module 916, which can then process the transformed input as an "answer" or otherwise satisfy the query. In some embodiments, a large language model can be used to process text, while another model can be used to convert (or transform) images, database tables, and / or other non-text inputs into text formats (e.g., natural language).

[0122] It will be understood that in some embodiments, the orchestrator module 904 may include some or all of the functionality of the understanding module 916. For example, the understanding module 916 may be a component of the orchestrator module 904. Similarly, in some embodiments, the understanding module 916 may include some or all of the functionality of the orchestrator module 904.

[0123] exist Figure 9In the examples, agent module 906 includes various different example agent modules 906-1 to 906-N. It will be understood that these are shown as examples, and various embodiments may include different agents in place of agents 906-1 to 906-N or in addition to agents 906-1 to 906-N. In some embodiments, each agent 906 includes hardware and / or software, and includes one or more large language models, one or more other machine learning models and / or functions to provide inference functionality to accomplish a prescribed set of tasks. It will be understood that references to agent modules can refer to the agent itself and / or the components that generate and / or execute the agent. In some embodiments, an orchestrator is a type of agent and may be referred to as an orchestrator agent. Therefore, references to an orchestrator can refer to the orchestrator itself and / or the components that generate and / or execute the orchestrator.

[0124] In various embodiments, some or all of the agents 906 can process data with different data types and / or data formats. For example, agent module 906 can receive a database table or image as input (e.g., received from tool 908) and translate the table or image into natural language describing the table or image, which can then be output for processing by other modules, models, and / or systems (e.g., orchestrator module 904 and / or understanding module 916). In one example, a large language model can be used to process text, while another model can be used to convert (or transform) images, database tables, and / or other non-text inputs into text formats (e.g., natural language).

[0125] The retrieval agent module 906-1 can function to retrieve structured and unstructured data records. In some embodiments, the retrieval agent module 906-1 can coordinate / instruct the unstructured data retrieval agent module 906-2 to retrieve unstructured data records, and coordinate / instruct the structured data retrieval agent module 906-3 and the type system retrieval agent module 906-4 to retrieve structured data records. For example, the retrieval agent module 906-1 can collaborate with other agents 906 and tools 908 to generate SQL queries to query an SQL database.

[0126] The unstructured data retrieval agent module 906-2 can function to retrieve (e.g., from unstructured data storage) unstructured data records and / or paragraphs (or segments) of these data records. Unstructured data records may, for example, include text data stored on a file system in formats such as PDF, DOCX, MD, HTML, TXT, and PPTX.

[0127] In some embodiments, agent 906-2 may use embeddings (e.g., vectors stored in vector storage 940) when retrieving information. For example, agent 906-2 may use similarity evaluation or search on vector data storage 940 to find relevant data records based on k nearest neighbors, where embeddings that are closer to each other are more likely to be relevant.

[0128] In some embodiments, the unstructured data retrieval agent module 906-2 implements a read-extract-answer (REA) data retrieval process and / or a read-answer (RA) data retrieval process. More specifically, REA and RA may be appropriate when system 802 needs to process large amounts of data. For example, a query may identify many different data records and / or paragraphs (e.g., hundreds or thousands of data records and paragraphs). For simplicity, references to data records may include data records and / or paragraphs.

[0129] More specifically, the unstructured data retrieval agent module 906-2 can determine whether each data record is relevant to answering the query and filter out irrelevant data records. For example, agent 906-2 can calculate and assign a relevance score to each retrieved data record (e.g., using a machine learning relevance model). The relevance score can be relative to other retrieved data records. For example, the least relevant data record can be assigned a minimum value (e.g., 0), and the most relevant data record can be assigned a maximum value (e.g., 100). The unstructured data retrieval agent module 906-2 can filter out relevant documents (or irrelevant documents). For example, the unstructured data retrieval agent module 906-2 can filter out data records with a relevance score below a configurable threshold (e.g., 90). In some embodiments, the number of data records that the unstructured data retrieval agent module 906-2 can retrieve for a particular input or query can be user-defined or system-defined, and can also be configurable. For example, the system can define that it can return up to 50 data records.

[0130] In some embodiments, a large language model (e.g., of the unstructured data retrieval agent module 906-2) can identify key points in relevant documents and paragraphs, and then provide these key points to the large language model (e.g., the large language model of the editor 904). The large language model can provide a summary that can be used to generate a query answer (e.g., the summary can be the query answer). This, for example, can allow system 802 to examine a wide variety of concepts and documents (e.g., in contrast to an iterative process). In some embodiments, if the number of documents or paragraphs is below a threshold, the unstructured data retrieval agent module 906-2 can skip the “extraction” step (e.g., summarizing key points) and provide the paragraphs directly to the large language model. This can be referred to as the RA process.

[0131] The structured data retrieval agent module 906-3 can function to retrieve structured data records and / or their segments (or sections) from various structured data stores. For example, structured data records may include tabular data persisted in relational databases, key-value stores, or external databases and modeled or accessed using entity types (or simply types). Structured data records may include data records structured according to one or more data models (e.g., complex data models) and / or data records that can be retrieved based on one or more data models. Structured data records may include data records stored in structured data stores (e.g., data stores structured according to one or more data models).

[0132] In a particular implementation, the data model may include a graph structure of objects or types, and the agent 906 and / or tool 908 may traverse the graph along different paths to identify relevant types of the data model (e.g., depending on the query and the plan for answering the query provided by the orchestrator module 904), and may utilize complex joins to combine multiple tables (e.g., contrary to simply passing a single piece of data from that single table and operating on that single table). Paths may be stored in a data store (e.g., a vector data store 940) for efficient retrieval.

[0133] In some embodiments, the structured data retrieval agent module 906-3 may use various tools to retrieve structured data (e.g., structured data retrieval tool 908-2, filter tool 908-9, projection tool 908-10, grouping tool 908-11, sorting tool 908-12, and restriction tool 908-13, etc.). In some embodiments, once the structured data retrieval agent module 906-3 has traversed the data model and retrieved (one or more) relevant types and / or subsets of the data model, the structured data retrieval agent module 906-3 can then use this information, along with the agent and / or tool outputs, to construct a structured query specification. The structured data retrieval agent module 906-3 can then execute this structured query specification against one or more structured data stores to retrieve structured data records.

[0134] The Type System Retrieval Agent Module 906-4 can function to retrieve types, data models, and / or subsets of data models. For example, a data model can include various different types, and each type can describe data fields, operations, and functions. Types can represent different objects (e.g., real-world objects such as machines or sensors in a factory), and each type can include a large language model context that provides context for a large language model. Types can be defined in natural language format for efficient processing by the large language model.

[0135] In some embodiments, the type system is designed to be used by different computing systems, application developers, data scientists, operators, and / or other users to build applications, develop and execute machine learning algorithms, and manage and monitor the status of jobs running on the type system (e.g., in some embodiments, an enterprise generative AI system). The type system is a framework that enables systems, application developers, data scientists, and other users to communicate effectively with each other using the same language. Therefore, application developers can interact with the enterprise generative AI system 802 in the same way as data scientists. For example, they can use the same types, the same methods, and the same features.

[0136] In some embodiments, the type system can abstract the complex infrastructure within the enterprise generative AI system 802. In one example, developers may never need to write SQL, CQL, or some other query processing language to access the data. When a user reads the data, the enterprise generative AI system 802 can generate the correct query for the underlying data storage, submit the query to the database, and present the results back to the user as a collection of objects or results.

[0137] In some embodiments, a type may resemble a programming language class (e.g., a Java class) and describe data fields, operations, and functions (e.g., static functions) that can be invoked on or by one or more applications, but the type is not bound to any particular programming language. A type can be a definition of one or more complex objects that system 802 can understand. For example, a type can represent a broad range of objects, such as a water pump. Besides objects, types can also be used to model systems (e.g., computing clusters, key-value data stores, file systems, file storage, and enterprise data stores). In some embodiments, complex relationships such as "when which light bulbs are in which light fixtures" can be modeled as types.

[0138] The Machine Learning Insight Agent Module 906-5 can function to obtain and / or process outputs from artificial intelligence applications (e.g., AI application insights). For example, the Machine Learning Insight Module 906-5 can instruct the text processing tool 908-3 to perform text processing tasks (e.g., transforming AI applications into natural language), instruct the image processing tool 908-4 to perform image processing tasks (e.g., generating natural language summaries of images output from AI applications), instruct the time series tool 908-3 to summarize time series data (e.g., time series data output from AI applications), and instruct the API tool 908-6 to perform API call tasks (e.g., executing API calls to trigger or access AI applications).

[0139] The time series processing agent 906-6 can function to acquire and / or process time series data, such as time series data output from various applications (e.g., artificial intelligence applications), machines, and sensors. The time series processing agent 906-6 can instruct the time series processing tool module 908-3 and / or collaborate with the time series processing tool module 508-3 to acquire time series data from one or more artificial intelligence applications and / or other data sources.

[0140] API agent module 906-7 can function to coordinate and manage communication with other applications. For example, API agent module 906-7 can instruct API tool module 908-6 to perform various API calls and then process the output of these tools (e.g., transform them into natural language summaries).

[0141] The mathematical agent module 906-8 can function to determine whether agent 906 or a large language model requires additional information to generate an answer or result. In some embodiments, the mathematical agent 906-8 can instruct the optimizer tool module 908-8 to perform various canonical analytic functions and mathematical optimizations to aid in the computation of answers to various questions. For example, the orchestrator module 904 can use the mathematical agent module 906-8 to generate plans and determine whether the orchestrator module 904 needs more information to generate the final result, etc.

[0142] The visualization agent module 906-9 can function to generate one or more visualizations and / or graphical user interfaces, such as dashboards and charts. For example, the visualization agent module 906-9 can execute the visualization tool module 908-7 to generate dashboards based on information retrieved by other agents and / or information output by other agents (e.g., a natural language summary of associated tool outputs). The visualization agent module 906-9 can also function to generate summaries of visual elements (e.g., natural language summaries), such as charts, tables, and images.

[0143] The code generation agent module 906-10 can function to instruct the code generation tool module 908-14 to generate source code, machine code, and / or other computer code. For example, the code generation agent module 906-10 can be configured to determine what code is needed (e.g., to satisfy queries and create applications) and instruct the tool 908-14 to generate that code in a specific language or format.

[0144] In some embodiments, tool 908 is a specific function that an agent (e.g., agent 906, orchestrator module 904) can access or perform when attempting to complete (e.g., one or more) a prescribed task from a set of prescribed tasks determined by orchestrator module 904. Tool 908 may include software and / or hardware. Tool 908 may also include one or more machine learning models, but it may also include functionality without any machine learning model. In some embodiments, tool 908 does not include a large language model, but in other embodiments, the tool may include a large language model. In some embodiments, some or all of agent 906 and / or tool 908 may be manually configured (e.g., by a user). Agent 906 and tool 908 may also normalize data (e.g., normalize to a common data format) before outputting the data.

[0145] Unstructured data retrieval tool 908-1 can function to retrieve unstructured data records from an unstructured data store. In some embodiments, agent 906-2 can use embeddings (e.g., vectors stored in vector storage 940) when retrieving information. For example, agent 906-2 can use similarity evaluation or search to find relevant data records based on k-nearest neighbors, where embeddings that are closer to each other are more likely to be relevant. Structured data retrieval tool 908-2 can function to access and retrieve structured data records from a structured data store (e.g., structured or modeled according to a data model). Structured data retrieval tool 908-2 can be executed by a structured data retrieval agent module 906-3.

[0146] Text processing tool module 908-3 can function to retrieve and / or transform text (e.g., from unstructured data records) and perform other text processing tasks (e.g., transforming text-based output from an AI application into natural language). Image processing tool module 908-4 can function to perform image processing tasks (e.g., generating natural language summaries of images). Time series processing tool module 908-5 can function to acquire and / or process (e.g., output from AI applications and sensors, etc.) time series data. For example, time series processing tool module 908-3 can be executed by one or more agents 906 to acquire and process time series data. API tool module 908-6 can function to perform API call tasks (e.g., performing API calls to trigger or access an AI application). For example, different agents 906 can use API tool module 906-8 whenever an agent needs to access or trigger another application.

[0147] Visualization tool module 908-7 can function to generate one or more visualizations and / or graphical user interfaces, such as dashboards. For example, visualization tool module 908-7 can generate dashboards based on information retrieved by other agents and / or information output by other agents (e.g., a natural language summary of associated tool outputs). Filter tool module 908-9 can function to filter data records and / or types, etc. For example, filter tool module 908-9 can filter projections (e.g., fields) identified by projection tool module 908-10 as part of a structured data retrieval process. In various embodiments, tool 908 can be executed in parallel or otherwise.

[0148] In some embodiments, filter tool module 908-9 can identify implicit filters based on queries or other inputs, and those identified implicit filters can be used as part of a structured data retrieval process. For example, a query may include "When was the last time premium towels sold out?". Filter tool module 908-9 can identify "premium towels" as a filter (e.g., based on associated type descriptions). Filter tool module 908-9 can also identify contextual date-time filters. For example, a query may include "How many systems were offline yesterday?" Filter tool module 908-9 can determine yesterday's date, while considering time zones and other relevant data to generate an accurate filter. In some embodiments, filter tool module 908-9 can validate the identified filters before using them (e.g., as part of a structured data retrieval process).

[0149] Projection tool module 908-10 can function to identify and select fields (e.g., type fields, object fields) that are relevant to determining the response to a query or other input. Grouping tool module 908-11 can function to group data (e.g., type and tool output, etc.) that can then be used (e.g., by structured data retrieval agent module 906-3 and / or structured data retrieval tool module 908-2) to generate structured query requests.

[0150] The sorting tool module 908-12 can function to sort data (e.g., type and tool output, etc.), which can then be used (e.g., by the structured data retrieval agent module 906-3 and / or the unstructured data retrieval tool module 906-2) to generate structured query requests. The limiting tool module 908-13 can function to limit the output of the structured data retrieval process. For example, it can limit the number of data records, types, groups, and / or filters retrieved.

[0151] Code generation tool module 908-14 can function to generate source code, machine code, and / or other computer code. For example, code generation tool module 908-14 can be configured to generate and / or execute SQL queries and / or JAVA code, etc. Code generation tool module 908-14 can be used to facilitate query generation for agents, other tools, and large language models, etc. In some embodiments, code generation tool module 908-15 can be configured to generate source code for an application or create an application.

[0152] Embedding generator module 912 can function to generate embeddings based on structured data records and / or segments, and unstructured data records and / or segments. Embedding generator module 912 can be identical to embedding generator module 128. Embedding generator module 912 can include one or more models (e.g., embedding models, deep learning models) that can transform and / or convert data records into vector representations, where vectors of semantically similar data records (e.g., the content of the data records) are close together in the vector space. This can facilitate retrieval operations utilizing agent 906 and tool 908.

[0153] In some embodiments, the embedding generator module 912 may generate embeddings using one or more embedding models (e.g., an implementation of the ColBERT embedding model). Embeddings may include digital representations of unstructured and / or structured data records that capture the semantic or contextual meaning of the data records. For example, an embedding may be represented by one or more vectors. Embeddings may be used when retrieving data records and performing similarity assessments or other retrieval operations. Embeddings may be stored in an embedding index (e.g., a vector data store 940). In some embodiments, the vector store 940 is a type of database specifically optimized for storing and retrieving embeddings using similarity heuristics (e.g., approximate nearest neighbor (ANN) algorithms) that may be implemented by agent 906 and / or tool 908. In one example, the vector store 940 may include an implementation of the FAISS vector store.

[0154] Chunking module 910 can function to process (e.g., chunk) a corpus of data records (e.g., one or more enterprise systems and / or external systems) for disposal by various systems (e.g., enterprise generative artificial intelligence system 802). Chunking module 910 can partition data records and insert or append corresponding headers for each chunk. Headers may include one or more attributes describing the chunk. A segment may include a header along with paragraphs of a text document, portions of a database table, and models or sub-models, etc. For simplicity, references to paragraphs may include the segment and / or other content of the segment (e.g., text). Segments may be stored in a segmented data store. Chunking can be rule-based.

[0155] In some implementations, the chunking module 910 can preprocess data records and / or segments to generate corresponding context information. In some embodiments, the context information may be included and / or represented in context metadata, and / or context metadata may be generated based on the context information. Context information can improve security and the accuracy and reliability of associated retrieval operations. In one example, the context information includes context metadata. The context information may include references between segments and / or data records. For example, references may indicate relationships that can be used (e.g., traversal) when performing similarity assessments or other aspects of retrieval operations (e.g., by one or more agents 906). The context information may also include information that can assist a large language model in generating plans and / or responses. For example, the chunking module 910 can generate context information for structured data chunks (or paragraphs) that include natural language descriptions of data records and the locations of related data records, etc.

[0156] Context information may include access control. In some implementations, context information provides user-based access control (e.g., role-based access control) to associated data records and / or segments. More specifically, context information may indicate user roles that can access corresponding segments and / or data records, and / or user roles that cannot access corresponding segments and / or data records. Context information may be stored in the headers of data records and / or data record segments. Context information may maintain references between data records and / or data record segments. The chunking module 910 may generate context information before, after, or simultaneously with the generation of associated embeddings. For example, context information may be used to create embeddings, or it may be used to enhance embeddings. Context information may be used by the chunking module 910 to map relationships between data records and / or segments of one or more enterprises or enterprise systems and store these relationships in a data model. In one example, the chunking module 910 implements the word2vec algorithm. In some implementations, the chunking module 910 utilizes a model trained on a domain-specific (or industry-specific) dataset.

[0157] In some embodiments, the chunking module 910 may perform some or all of the functionalities described herein periodically (e.g., in batches), on demand, and / or in real time. For example, the chunking module 910 may trigger periodically, on demand, manually, and / or automatically. In some implementations, subsequent chunking operations may only incorporate changes relative to previous chunking operations (e.g., "increments").

[0158] In some implementations, the embedding generator module 912 generates enhanced embeddings. For example, the chunking module 910 may generate enhanced embeddings based on context information, data records, and / or data record segments. Enhanced embeddings may include vector values ​​based on embedding vectors and context information. In some embodiments, enhanced embeddings include embedding vector values ​​along with context metadata including context information. Enhanced embeddings may be indexed in an enhanced embedding data store (e.g., vector data store 940). Agent 906 and / or tool 908 may retrieve unstructured and / or structured data records based on enhanced embeddings. In some embodiments, context information may be included and / or represented in context metadata, and / or context metadata may be generated based on context information.

[0159] The crawling module 914 can function to scan and / or crawl different data sources across different domains (e.g., enterprise data sources, external data sources). The crawling module 914 can identify existing data records, new data records, and / or updated data records. The crawling module 914 can notify the chunking module 910 of new and updated data records, and the chunking module 910 can chunk those data records. In some embodiments, the crawling module 914 can function periodically, on demand, and / or in real time and / or trigger operations.

[0160] In some embodiments, the crawling module 914 can function to scan and / or crawl different data sources (e.g., enterprise system 504, external system 506) across different domains (e.g., data domains). This can identify existing data records, new data records, and / or updated data records. The crawling module 914 can notify the chunking module 910 of new and updated data records, and the chunking module 910 can chunk those data records. In some embodiments, the crawling module 914 can function and / or trigger operations periodically, on demand, and / or in real time.

[0161] In some implementations, the information source may include a model registry storing various models (e.g., machine learning models, large language models, multimodal models). Models can be trained on general datasets and / or domain-specific datasets. The processes described herein can be applied to various model registries. For example, a model may be associated with embedding values ​​(e.g., generated by an embedding model) to facilitate model retrieval.

[0162] In some embodiments, the segmentation module 910, the embedding generator module 912, and / or the crawling module 914 can be implemented. Figure 2 The functionalities described in the text (e.g., extracting information, recognizing images and tables, segmenting text, extracting information, and generating infographics, etc.).

[0163] Enterprise generative artificial intelligence systems 802 can perform some or all of the functionalities described herein periodically (e.g., in batches), on demand, and / or in real time. For example, the system can trigger the intelligent crawling and indexing described herein periodically, on demand, manually, or automatically. In some implementations, subsequent crawling and indexing operations may only incorporate changes relative to previous crawling and indexing operations (e.g., "incrementally").

[0164] Understanding module 916 can function to process input to determine a result (e.g., an "answer"), determine the underlying principles of the result, and determine whether understanding module 916 needs more information to determine the result. Understanding module 916 can output information (e.g., results or additional queries) in natural language or machine language format. In some implementations, one or more features of the understanding module define the conditions or functions for determining whether more information is needed to satisfy the initial input or whether sufficient information exists to satisfy the initial input.

[0165] In some embodiments, the understanding module 916 includes one or more large language models. The large language models can be configured to generate and process context and other information described herein. The understanding module 916 may also include other language models that preprocess the input (e.g., a user query) before it is provided to the agent for disposal. The understanding module 916 may also include one or more large language models that process the output from other models and modules (e.g., the model of agent 906). The understanding module 916 may also include another large language model for processing responses from a large language model into a format more consistent with the final response that can be transmitted to various users and / or systems (e.g., the user or system that provided the initial query or other intended recipients of the response).

[0166] For example, the understanding module 916 can format responses based on various viewpoints. Viewpoints can be based on user type (e.g., human or machine), user role (e.g., data scientist, engineer, and manager), and access permissions. Therefore, viewpoints enable the understanding module 916 to generate and provide responses specifically tailored to the recipient. The understanding module 916 can also notify the user and the system whether it cannot find an answer (e.g., contrary to presenting potentially erroneous or biased responses).

[0167] In some implementations, the features of one or more large language models in the understanding module 916 define conditions or functions for determining whether more information is needed to satisfy the initial input or whether there is sufficient information to satisfy the initial input. The large language model in the understanding module 916 can also define stopping conditions that indicate a stopping threshold condition, which indicates the maximum number of iterations that can be performed before the iteration process terminates.

[0168] In some embodiments, the understanding module 916 may generate and (e.g., in one or more data stores) principles and context. Principles may be inferences used by the understanding module 916 to determine outputs (e.g., natural language output, indications that it needs more information, indications that it can satisfy the initial input). The understanding module 916 may generate context based on principles. In some implementations, context includes concatenations and / or annotations of one or more segments of a data record and / or embeddings associated with them, along with mappings of concatenations and / or annotations. For example, mappings may indicate relationships between different segments and / or weighted or relative values ​​associated with different segments. Principles and / or context may be included in prompts provided to a large language model.

[0169] In some embodiments, the understanding module 916 includes a query and principle generator that generates queries or other inputs for a model (e.g., a large language model, other machine learning models) and / or generates and stores principles and context (e.g., in data storage 960). The query and principle generator can function to process, extract, and / or transform different types of data (e.g., text, database tables, images, videos, and / or code, etc.). For example, the query and principle generator can accept a database table as input and transform it into natural language describing the database table, which can then be provided to one or more other models (e.g., a large language model) of the understanding module 916, which can then process the transformed input as an "answer" or otherwise satisfy the query. In some implementations, the query and principle generator includes a model that can convert (or transform) images, database tables, and / or other non-text inputs into text formats (e.g., natural language). It will be understood that although queries are used in the various examples herein, other types of input (e.g., instruction sets) can be processed in the same or similar manner as described for queries.

[0170] In some embodiments, the understanding module 916 can use different models for different domains. For example, different domains may correspond to different industries (e.g., aerospace, defense), different technological environments (e.g., local, air-gap, cloud-native), and / or different enterprises or organizations. Therefore, the understanding module 916 can use a specific model (e.g., a data model and / or a large language model) for a specific domain (e.g., a data model describing the properties and relationships of aerospace objects and a large language model trained on an aerospace-specific dataset), and use another data model and / or a large language model for another domain (e.g., a data model describing the properties and relationships of defense-specific objects and a large language model trained on a defense-specific dataset), and so on.

[0171] In some embodiments, the orchestrator module 904 includes some or all of the functionality and / or structure of the understanding modules 916 and / or 906, which are further described below. Similarly, in some embodiments, the understanding module 916 may include some or all of the functionality and / or structure of the orchestrator module 904.

[0172] In some embodiments, the understanding module 916 may function to generate large language model hints (or simply hints) and hint templates. For example, the understanding module 916 may generate a hint template for processing initial input, a hint template for processing iterative input (i.e., input received during the iteration process after processing the initial input), and another hint template for the output result stage (i.e., when the understanding module 916 has determined that it has sufficient information and / or meets the stopping condition). The understanding module 916 may modify appropriate hint templates according to the stage of the iteration process. For example, the hint template may be modified to generate hints that include principles and context, which can inform subsequent iterations.

[0173] Enterprise access control module 918 can function to provide enterprise access control (e.g., layers and / or protocols) for enterprise generative artificial intelligence system 802, associated systems (e.g., enterprise systems), and / or environments (e.g., enterprise information environments). Enterprise access control module 918 can provide functionality to enforce access control policies for generated results (e.g., preventing orchestrator module 904 and / or understanding module 916 from generating results containing sensitive information) and / or filtering results generated before providing the final result.

[0174] In some implementations, the enterprise access control module 918 can evaluate whether a user is authorized to access all or only a portion of the results (e.g., answers). For example, a user can provide a query associated with a first department or subunit of an organization. Members of that department or subunit can be restricted from accessing certain data, data types, data models, or other aspects of the data domain to be searched. In cases where the initial results include data that the user has restricted access to, the enterprise access control module 918 can determine how to handle such restricted data, such as completely omitting the restricted data, omitting the restricted data but indicating that the results include data that the user has restricted access to, or providing information related to all of the initial results, etc. In an example of completely omitting the restricted data, a final set of results can be returned for presentation to the user, where the final set of results does not inform the user that a portion of the initial results has been omitted. In an example of omitting the restricted data but providing the user with an indication that there is restricted data, the final results can include only those results that the user is authorized to access, but can include information indicating that there were X initial results but only Y results were output, where Y < X. In the third example above, all of the results, including those that the user has restricted access to, can be output to the user.

[0175] Additionally or alternatively, the enterprise access control module 918 can communicate with one or more other modules to obtain information that can be used to enforce access permissions / restrictions in conjunction with performing retrieval operations rather than for controlling the presentation of results to the user. For example, the enterprise access control module 918 can restrict the data sources to which retrieval operations are applied, such as not applying the retrieval operations to portions of the data source that are denied user access and applying the retrieval operations to portions of the data source that are permitted user access, etc. Note that the above exemplary techniques for enforcing access restrictions have been provided for illustrative purposes and not by way of limitation, and it should be understood that modules operating in accordance with embodiments of the present disclosure can implement other techniques to present results via an interface based on access restrictions.

[0176] In some embodiments, to facilitate the enforcement of access restrictions in conjunction with searches performed by the enterprise generative AI system 802, the enterprise access control module 918 may store information associated with access restrictions or permissions for each user. To retrieve relevant restriction data for a user, the enterprise access control module 918 may receive information identifying the user in conjunction with input or when the user logs into the system on which the enterprise access control module 918 is running. The enterprise access control module 918 may use the user identification information to retrieve appropriate restriction data to support the enforcement of access restrictions in conjunction with enterprise searches. In some embodiments, the enterprise access control module 918 may include credential management functionality of a model-driven architecture on which the enterprise generative AI system 802 is deployed, or it may be a remote credential management system communicatively coupled to the enterprise generative AI system 802 via a network.

[0177] The AI ​​traceability module 920 can function to provide traceability and / or explainability of responses generated by the enterprise generative AI system 802. For example, the AI ​​traceability module 920 can indicate portions of data records used to generate the responses and their associated data sources. The AI ​​traceability module 920 can also function to validate large language model outputs. For example, the AI ​​traceability module 920 can automatically and / or on-demand provide source references to validate or verify large language model outputs. The AI ​​traceability module 920 can also determine the compatibility of different sources (e.g., data records, paragraphs) used to generate large language model outputs. For example, the AI ​​traceability module 920 can identify contradictory data records (e.g., one data record indicates John Doe is an employee of Acme, while another indicates John Doe works for different companies) and provide outputs that are generated based on the contradictions in the conflicting information. The AI ​​traceability module 920 can collaborate with and / or include the functionality of the anti-hallucination and attribution module 934.

[0178] Parallelization module 922 can function to control the parallelization of various systems, modules, agents, models, and processes described herein. For example, parallelization module 922 can cause parallel execution of different agents and / or orchestrators. Parallelization module 922 can be controlled by orchestrator module 904.

[0179] The model generation module 924 can function to obtain, generate, and / or modify some or all of the different types of models (e.g., machine learning models, large language models, data models) described herein. In some implementations, the model generation module 924 can use various machine learning techniques or algorithms to generate models. As used herein, artificial intelligence and / or machine learning can include Bayesian algorithms and / or models, deep learning algorithms and / or models (e.g., artificial neural networks, convolutional neural networks), gap analysis algorithms and / or models, supervised learning techniques and / or models, unsupervised learning algorithms and / or models, semi-supervised learning techniques and / or models, random forest algorithms and / or models, similarity learning and / or distance algorithms, generative artificial intelligence algorithms and models, clustering algorithms and / or models, transformer-based algorithms and / or models, machine learning algorithms and / or models based on neural network transformers, and / or reinforcement learning algorithms and / or models, etc. Algorithms can be used to generate corresponding models. For example, algorithms can be executed on datasets (e.g., domain-specific datasets, enterprise datasets) to generate and / or output corresponding models.

[0180] In some embodiments, a large language model is a deep learning model (e.g., generated by a deep learning algorithm) that can identify, summarize, translate, predict, and / or generate text and other content based on knowledge gained from large-scale datasets. Large language models can include transformer-based models. Examples of large language models include Google's BERT, OpenAI's GPT-3, and Microsoft's Transformer. Large language models can process massive amounts of data, thereby improving accuracy in prediction and classification tasks. Large language models can use this information to learn patterns and relationships, which can help them make improved predictions and groupings relative to other machine learning models. Large language models can include artificial neural network transformers pre-trained using supervised and / or semi-supervised learning techniques. In some embodiments, large language models include deep learning models specifically designed for text generation. In some embodiments, large language models can be characterized by a large number of parameters (e.g., hundreds of billions or trillions of parameters) and a large text corpus used to train them.

[0181] Although the systems and processes described herein use large language models, it will be understood that, instead of large language models or in addition to large language models, other embodiments may use different types of machine learning models. For example, orchestrator 904 may use a deep learning model specifically designed to receive non-natural language input (e.g., images, videos, audio) and provide natural language output (e.g., summaries) and / or other types of output (e.g., video summaries).

[0182] The model deployment module 926 can function to deploy some or all of the different types of models described herein. In some implementations, the model deployment module 926 can deploy models before or after the deployment of the enterprise generative AI system. For example, the model deployment module 926 can collaborate with the model optimization module 928 to exchange or otherwise modify large language models of the enterprise generative AI system.

[0183] In some implementations, the model registry 950 can store various models (e.g., machine learning models, large language models, data models) and / or model configurations. Models can be trained on general datasets and / or domain-specific datasets. For example, the model registry can store different configurations of various large language models (e.g., which can be deployed or exchanged in an enterprise generative artificial intelligence system 802). In some embodiments, individual models can be associated with embedded values ​​or enhanced embedded values ​​to (e.g., to facilitate retrieval operations in the same or similar manner as data record retrieval).

[0184] Model optimization module 928 can function to enable tuning and learning using modules described herein (e.g., understanding module 916) and / or models (e.g., machine learning models, large language models). For example, model optimization module 928 can tune understanding module 916 and / or orchestrator module 904 (and / or its models) based on tracking user interactions within the system, capturing explicit and / or implicit feedback (e.g., by training a user interface), etc. In some example implementations, model optimization module 928 can use reinforcement learning to accelerate knowledge base bootstrapping. Reinforcement learning can be used for explicit bootstrapping of various systems (e.g., enterprise generative AI system 802) by instrumenting things like time spent and / or results clicked. Example aspects of model optimization module 928 include innovative learning frameworks that can bootstrap models for different enterprise environments.

[0185] In some embodiments, reinforcement learning is a machine learning training method that rewards desired behavior and / or punishes unwanted behavior. Typically, a reinforcement learning agent is able to perceive and interpret its environment, take actions, and learn through trial and error. Reinforcement learning uses algorithms and models to determine the optimal behavior in an environment to obtain the maximum reward. This optimal behavior is learned through interaction with the environment and observation of how to react. Without a supervisor, the learner must independently discover the sequence of actions that maximizes the reward. This discovery process is similar to trial-and-error search. The quality of actions is measured not only by the immediate reward they return but also by the delayed reward they may obtain. Because reinforcement learning can learn actions that ultimately lead to success in unseen environments without the help of a supervisor, it is a very powerful algorithm. ColBERT is an example retrieval model that enables scalable BERT-based search on large text collections (e.g., within tens of milliseconds). ColBERT uses a late-interaction architecture that uses BERT to independently encode queries and documents, then employs “cheap” but powerful interaction steps that model their fine-grained similarity. In addition to reducing the cost of re-ranking documents retrieved by traditional models, ColBERT's pruning-friendly interaction mechanism also enables end-to-end retrieval directly from large document collections using vector similarity indexes.

[0186] In some embodiments, the model optimization module 928 can periodically, on demand, and / or in real time retrain the model (e.g., a transformer-based natural language machine learning model). In some example implementations, a corresponding candidate model (e.g., a candidate transformer-based natural language machine learning model) can be trained based on user selection, and the model optimization module 928 can replace some or all of the model with one or more candidate models already trained on the received user selection.

[0187] In some embodiments, besides before or after runtime, the model optimization module 928 can also replace the models of the enterprise generative artificial intelligence system at runtime or during runtime. For example, the orchestrator module 904, the understanding module 916, and / or the agent 906 can use a specific set of machine learning models for one domain and other models for different domains. The model exchange module can select and use appropriate models for a given domain. This can even occur during iterative processing. For example, when a new query is generated by the understanding module 916, the domain may change, which may trigger the model exchange module to select and deploy different models suitable for that domain.

[0188] In some embodiments, the model optimization module 928 can train a generative artificial intelligence model to develop different types of responses (e.g., best results, ranking results, smart cards, chatbots, and / or new content generation, etc.).

[0189] In some embodiments, the model optimization module 928 can periodically, on demand, and / or in real time retrain the model (e.g., a large language model). In some example implementations, corresponding candidate models can be trained based on user selections, and the system can replace some or all of the model with one or more candidate models already trained on the received user selections.

[0190] Interface module 930 can be used to receive input (e.g., complex input) from a user and / or system. Interface module 930 can also generate and / or transmit output. Input can include system input and user input. For example, input can include instruction sets, queries, natural language input or other human-readable input, and / or machine-readable input, etc. Similarly, output can also include system output and human-readable output. In some embodiments, input (e.g., requests, queries) can be input in various natural forms that facilitate human interaction (e.g., basic text box interfaces, image processing, and / or voice activation, etc.) and processed to quickly find relevant and responsive information.

[0191] In some embodiments, interface module 930 may function to generate a graphical user interface component (e.g., a server-side graphical user interface component) that can be rendered as a complete graphical user interface on the enterprise generative artificial intelligence system 802 and / or other systems. For example, interface module 930 may be used to present an interactive graphical user interface for displaying and receiving information.

[0192] Communication module 932 can be used to send requests, transmit and receive communications, and / or otherwise provide communication with one or more of the systems, modules, engines, layers, devices, data stores, and / or other components described herein. In a particular implementation, communication module 932 can be used to encrypt and decrypt communications. Communication module 932 can be used to send requests to one or more systems and receive data from one or more systems via a network or part of a network (e.g., communication network 808). In a particular implementation, communication module 932 can send requests and receive data via a connection, which may be wholly or partially wireless. Communication module 932 can request messages and / or other communications from associated systems, modules, layers, etc., and receive messages and / or other communications from associated systems, modules, layers, etc. Communications can be stored in enterprise generative artificial intelligence system data storage 970.

[0193] In some embodiments, the configuration, coordination, and collaboration of the orchestrator module 904, agent 906, tool 908, and / or other modules of the enterprise generative AI system 802 (e.g., understanding module 916) enable the enterprise generative AI system 802 to provide a multi-hop architecture that allows for complex reasoning across multiple agents 906, tools 908, and data sources (e.g., vector data stores, feature data stores, data models, enterprise data stores, unstructured data sources, and structured data sources). In various embodiments, some or all of the modules of the enterprise generative AI system 802 can be configured manually (e.g., by a user) and / or automatically (e.g., without requiring user input). For example, large language model hints can be configured, tool 908 descriptions can be configured (e.g., to be utilized more efficiently by agent 906 and orchestrator module 904), and a maximum number of hops or iterations can be configured. In one example, orchestrator module 904 receives a query from a user, determines a plan for answering the query, and selects agent 906 and / or tool 908 to execute a prescribed set of tasks that formulate the plan to answer the query. Agent 906 and / or tool 908 execute the prescribed set of tasks, and orchestrator module 904 observes the results. Orchestrator module 904 determines whether to submit a final answer or whether it needs more information. If orchestrator module 904 has sufficient information, it can generate and / or provide a final answer. Otherwise, orchestrator module 904 can create another prescribed set of tasks, and the process can continue until orchestrator module 904 has sufficient information to answer or a stopping condition (e.g., maximum hop count) is met.

[0194] The anti-hallucination and attribution module 934 may be the same as the anti-hallucination and attribution module 110. The anti-hallucination and attribution module 934 may be an extension of the understanding module 916 and / or included in the understanding module 916.

[0195] Figure 10A A flowchart 1000 depicts an example iterative generative artificial intelligence process using unstructured data according to some embodiments. This example process can be implemented by an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802). In this flowchart and / or sequence diagram, as well as other flowcharts and / or other sequence diagrams, the flowcharts exemplify the sequence of steps by way of example. It should be understood that, where applicable, some or all of the steps may be repeated, reorganized to be executed in parallel, and / or reordered. Furthermore, for clarity, some steps that have been included may have been removed to avoid providing excessive information, and some steps that have been included may have been removed but may have been included for clarity of illustration.

[0196] In step 1002, the user query is provided to the retrieval model (e.g., the retrieval module of the retrieval agent module). In step 1004, the retrieval model receives the query and performs a similar search (e.g., an ANN-based search) in vector storage 1006. In step 1010, the retrieved information is returned to the retrieval model and provided to the large language model. In step 1012, the large language model (e.g., the large language model used in step 1010 and / or a different large language model) determines whether additional information is needed to answer the user query. If more information is needed, steps 1004-1012 can be iteratively repeated with updated large language model hints and / or queries (e.g., using the large language model in step 1008) until the large language model has sufficient information to answer the query or a stopping condition is met (e.g., the maximum number of iterations has been performed). In step 1014, an answer is generated and / or presented (e.g., a final result if sufficient information exists for the large language model to determine an answer, or "I don't know" if the stopping condition is met before sufficient information can be received). The final result may also include the principles used by the large language model to generate the answer. The large language model used in steps 1008 and 1010 may be the same large language model and / or different large language models.

[0197] Figure 10B A flowchart 1030 depicts an example non-iterative generative artificial intelligence process using unstructured data according to some embodiments. This example process can be implemented by an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802). In these flowcharts and / or sequence diagrams, and other flowcharts and / or other sequence diagrams, the flowcharts exemplify a sequence of steps by way of example. It should be understood that, where applicable, some or all of the steps may be repeated, reorganized to be executed in parallel, and / or reordered. Furthermore, for clarity, some steps that have been included may have been removed to avoid providing excessive information, and some steps that have been included may have been removed but may have been included for clarity of illustration.

[0198] In step 1032, a query is received. Query 1032 is executed against vector storage 1034, and relevant paragraphs 1036 are retrieved. In some embodiments, user query 1032 is preprocessed (e.g., by an editor) before being applied to vector storage 1034 to retrieve paragraphs 1036. For example, query 1032 may be translated and transformed. Since vector storage may struggle to handle complex input, the system can generate new queries or multiple shorter queries from user query 1032 that vector storage 1034 can handle efficiently and accurately. Query 1032 and paragraphs 1036 are provided to a large language model 1038, which can create extracts 1040 (e.g., paragraph summaries) for each paragraph. Extracts are combined (e.g., concatenated) in step 1042 and provided to the large language model 1044 along with query 1032. In some embodiments, the extraction step is optional, and instead of extraction, paragraphs may be concatenated and provided to the large language model 1044. The large language model 1044 can generate a final response based on query 1032 and combined extraction 1042. In some embodiments, the large language model 1044 can post-process the results before they are provided to the user (e.g., using an orchestrator). For example, it can be translated, formatted, including quotations and attributes, based on viewpoints, etc.

[0199] Figure 10C A flowchart 1060 depicts an example non-iterative generative artificial intelligence process using unstructured data according to some embodiments. This example process can be implemented by an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802). In these flowcharts and / or sequence diagrams, and other flowcharts and / or other sequence diagrams, the flowcharts exemplify the sequence of steps by way of example. It should be understood that, where applicable, some or all of the steps may be repeated, reorganized to be executed in parallel, and / or reordered. Furthermore, for clarity, some steps that have been included may have been removed to avoid providing excessive information, and some steps that have been included may have been removed but may have been included for clarity of illustration.

[0200] In step 1062, a query is received. For example, the query could be “How much wine do they produce?”. This query is difficult for traditional large language models to handle and typically leads to large language model illusion because it is unclear how to handle “they” in query 1062. To address this problem, enterprise generative AI systems can use context 1064 to generate improved queries. For example, a previous conversation 1064 (e.g., as part of a chat with a chatbot) may have already included a discussion about France. The system can provide France as contextual information 1064 to generate a new query 1066 (such as “How much wine does France produce?”). This prevents large language models from becoming delusional and allows them to provide accurate and reliable final results 1078.

[0201] More specifically, an enterprise generative artificial intelligence system can (e.g., using a large language model 1065) generate a rewritten query 1066 that can be executed against a vector store 1068 to retrieve paragraph 1070. In some embodiments, the rewritten query 1066 can be preprocessed (e.g., by an orchestrator) before being applied to the vector store 1068. For example, the rewritten query 1066 can be translated and transformed. Because the vector store struggles to handle complex input, the system can generate new queries or multiple shorter queries from the rewritten query 1066 that the vector store 1068 can handle efficiently and accurately. In some embodiments, this preprocessing can be performed during the generation of the rewritten query (e.g., rewriting the query includes a preprocessing step).

[0202] The rewritten query 1066 and paragraph 1070 are provided to a large language model 1069, which can create extracts 1072 (e.g., paragraph summaries) for each paragraph. Extracts 1072 are combined (e.g., cascaded) in step 1074 and provided to the large language model 1076 together with the rewritten query 1066. An enterprise generative AI system can use the combined extracts 1074 to generate a principle 1082 for determining the final response 1078. For example, the large language model 1076 can generate the final response 1078 based on principle 1082 and / or present principle 1082 (or a summary of principles) along with the final response 1078 (e.g., for citation or attribution purposes).

[0203] In some embodiments, the extraction step is optional, and instead of extraction or combined extraction, paragraphs can be concatenated and provided to the large language model 1076. The large language model 1044 can generate a final response 1078 based on the rewritten query 1066 and the combined extraction 1074. In some embodiments, the large language model 1076 can post-process the result before the final result 1078 is provided to the user (e.g., using an editor). For example, it can be translated, formatted based on viewpoint, including citations and attributes, etc.

[0204] Figure 11 A flowchart 1100 depicts an example iterative generative artificial intelligence process using unstructured data according to some embodiments. This example process can be implemented by an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802). In this flowchart and / or sequence diagram, as well as other flowcharts and / or other sequence diagrams, the flowcharts exemplify the sequence of steps by way of example. It should be understood that, where applicable, some or all of the steps may be repeated, reorganized to be executed in parallel, and / or reordered. Furthermore, for clarity, some steps that have been included may have been removed to avoid providing excessive information, and some steps that have been included may have been removed but may have been included for clarity of illustration.

[0205] exist Figure 11 In the examples, an enterprise generative AI system (e.g., enterprise generative AI system 802) includes one or more retrieval modules 1104 and one or more understanding modules 1106. For example, retrieval module 1104 may include one or more large language models, and understanding modules may include one or more other large language models. These modules, and the iterative interactions (e.g., communication) between them, can allow the enterprise generative AI system to achieve the technical features and benefits discussed herein.

[0206] exist Figure 11 In the example, the enterprise generative AI system may receive initial input 1102 from a user or another system. For example, an orchestrator module 1103 may receive input 1102. The enterprise generative AI system may provide this input to a retrieval module 1104 (e.g., corresponding to one or more of agents 906 and / or tools 908), which can then extract and “retrieve” information from the embedding storage 1108. For example, the retrieval module 1104 may retrieve paragraphs (or data records) related to the input by using one or more similarity heuristics (e.g., ANN algorithms) executed on the embedding storage 1108 (e.g., one or more vector stores).

[0207] The enterprise generative AI system can use the retrieved information to generate an initial prompt for understanding module 1106. Understanding module 1106 can process this initial prompt and determine, based on the initial input, whether it contains sufficient information to meet criteria (e.g., answering a question). See, for example, step 1107. If it contains sufficient information to meet the initial input, the understanding module can then provide the result to a recipient (e.g., see, for example, step 1113), such as the user or system that provided the initial input. However, if understanding module 1106 determines, based on the initial input, that it does not contain sufficient information to meet the criteria, it can further synthesize information via an iterative process that provides the core benefits of the system.

[0208] There may be many reasons why the understanding module 1106 might need additional information. For example, traditional systems use only a single-pass process that addresses only a portion of a complex input. Enterprise generative AI systems address this problem by triggering subsequent iterations to solve the rest of the complex input and by including context to further refine the process.

[0209] More specifically, if the understanding module 1106 determines that it needs additional information to satisfy the initial input, it can generate context-specific data (or simply "context") that informs future iterations of the process and helps the system satisfy the initial input more efficiently and accurately. The context is based on principles used by the understanding module 1106 while it is processing a query (or other input). For example, the understanding module 1106 may receive segments of information retrieved by the retrieval module 1104. These segments may be, for example, paragraphs of (one or more) data records, and may be associated with embeddings from the embedding data store 1108 that facilitate the processing of the understanding module 1106. The query and principle generator 1112 of the understanding module 1106 can process this information and generate principles explaining why it produces the results. These principles can be stored in the historical principle data store 1110 by the enterprise generative AI system and provide a basis for subsequent iterations of the context.

[0210] More specifically, subsequent iterations may include the understanding module 1106 generating a new query, request, or other output, which is then passed back to the retrieval module. The retrieval module 1104 can process the new query and retrieve additional information. The system then generates a new prompt based on the additional information and context. The understanding module 1106 can process the new prompt and again determine if it requires additional information. If it requires additional information, the enterprise generative AI system can repeat (e.g., iterate) this process until the understanding module 1106 can meet the criteria based on the initial input, at which point the understanding module 1106 can generate an output result 1114 (e.g., "Answer" or "I don't know"). For example, if (e.g., by applying rules from the understanding module 1106) no relevant paragraphs are generated or retrieved and / or not enough relevant paragraphs are generated, retrieved, and / or extracted, the answer "I don't know" is generated. The understanding module 1106 can prevent illusions and improve the performance of "I don't know" queries while preserving calls to models (e.g., large language models).

[0211] In some embodiments, the presence and / or association of sufficient information, but (e.g., using the understanding module 1106) not being extracted, can be determined based on the number of retrieved paragraphs. For example, a threshold number or percentage of retrieved paragraphs for which relevant information has been extracted may need to be met (e.g., a specific number or percentage of retrieved paragraphs) for the enterprise understanding module 1106 to determine that it has sufficient information to answer the query. In another example, a threshold number or percentage of retrieved paragraphs for which no relevant information has been extracted (e.g., 4 paragraphs or 80% of retrieved paragraphs) may cause the enterprise understanding module 1106 to determine that it does not have sufficient information to answer the query.

[0212] Enterprise generative AI systems can also implement supervisory functions, such as stopping conditions to prevent the system from generating illusions or otherwise providing incorrect answers. Stopping conditions can also prevent the system from executing infinite iterative loops. In one example, the enterprise generative AI system can limit the number of iterations that can be performed before the understanding module 1106 provides an output or indicates that no output can be found. Users can also receive feedback 1116, which can be stored in the feedback data store 1118. In some embodiments, the enterprise generative AI system can use feedback to improve the accuracy and / or reliability of the system. As discussed elsewhere herein, it will be understood that in some embodiments, the functionality of the understanding module can be included within an orchestrator.

[0213] Figure 12A flowchart 1200 depicts an example enterprise generative artificial intelligence method according to some embodiments. In this flowchart and / or sequence diagram, as well as other flowcharts and / or other sequence diagrams, the flowcharts exemplify a sequence of steps by way of example. It should be understood that, where applicable, some or all of the steps may be repeated, reorganized to be performed in parallel, and / or reordered. Furthermore, for clarity, some steps that have been included may have been removed to avoid providing excessive information, and some steps that have been included may have been removed but may have been included for clarity of illustration.

[0214] In step 1202, the enterprise system (e.g., enterprise system 804) displays a graphical user interface (GUI). In some embodiments, an interface module (e.g., interface module 930) of an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802) can facilitate the display of the GUI. For example, the interface module can generate (e.g., render) a server-side portion of the GUI, and the enterprise system can generate a client-side portion of the GUI.

[0215] In step 1204, the enterprise generative artificial intelligence system receives a question through a graphical user interface. For example, the question may be a natural language query. In some embodiments, an orchestrator module (e.g., orchestrator module 904) receives the query.

[0216] In step 1206, the enterprise generative AI system retrieves problem-related information from different enterprise systems (e.g., enterprise system 804). In some embodiments, an understanding module (e.g., understanding module 916) retrieves the information. Figure 15 An example information retrieval process is illustrated. For instance, the understanding module 916 can collaborate with agents and / or tools to retrieve information.

[0217] In step 1208, the enterprise generative AI system generates an answer to the question using information retrieved from different enterprise systems through generative AI models (or multiple models). In some embodiments, the understanding module generates the answer.

[0218] In step 1210, the enterprise system displays the answer to the question through a graphical user interface. In some embodiments, the interface module facilitates the display.

[0219] Figure 13A flowchart 1300 depicts an example enterprise generative artificial intelligence method using an intelligent agent orchestrator architecture according to some embodiments. In step 1302, an enterprise system (e.g., enterprise system 804) displays a graphical user interface (GUI). In some embodiments, an interface module (e.g., interface module 930) of the enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802) may facilitate the display of the GUI. For example, the interface module may generate (e.g., render) a server-side portion of the GUI, and the enterprise system may generate a client-side portion of the GUI.

[0220] In step 1304, the enterprise generative artificial intelligence system receives a question through a graphical user interface. For example, the question may be a natural language query. In some embodiments, an orchestrator module (e.g., orchestrator module 904) receives the query.

[0221] In step 1306, the enterprise generative AI system manages different agent programs (e.g., agent modules 906 to 910) through an orchestrator program (e.g., orchestrator module 904) to generate answers to the question. The orchestrator can use at least one multimodal model to transform the question into a series of instructions for enabling different agent programs to retrieve question-related information. Different agent programs can use one or more machine learning models based on the series of instructions to retrieve question-related information.

[0222] In step 1308, the enterprise generative artificial intelligence system retrieves problem-related information through different agent programs. In some embodiments, the retrieval agent modules 906-1 to 906-4 retrieve information (e.g., from vector storage, enterprise systems, and / or external systems).

[0223] In step 1310, the enterprise generative AI system uses relevant information through a generative AI model to generate an answer to the question. In some embodiments, an understanding module (e.g., understanding module 916) generates the answer. The generative AI model and the multimodal model can be the same model. Alternatively, the generative AI model and the multimodal model can be different models.

[0224] In step 1312, the enterprise system displays the answer to the question through a graphical user interface. In some embodiments, the interface module facilitates the display.

[0225] In various embodiments, managing different agent programs may include iteratively processing multiple instructions from an orchestrator. Retrieval may include retrieving time-series data, structured data, and unstructured data. Agent programs may instantiate tools to manipulate instructions, retrieved data, or intermediate data. Agent programs may perform operations such as computation, translation, formatting, and / or visualization. One or more agent programs may be trained on different domain-specific machine learning models. Agent programs may employ a type system to unify incompatible data from different data sources.

[0226] Figure 14 A flowchart 1400 depicts an example anti-hallucination and attribution method for an enterprise generative artificial intelligence system. In step 1402, the enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802) receives output from a generative artificial intelligence model that processes cues. In some embodiments, an anti-hallucination and attribution module (e.g., anti-hallucination and attribution module 110) receives the output. The generative artificial intelligence model may employ generative adversarial networks, variational autoencoders, autoregressive models, and / or recurrent neural networks.

[0227] In step 1404, an anti-illusion and attribution module (e.g., in an enterprise generative AI system) resolves the output from the generative model into chunks attributable to one or more source segments. The anti-illusion and attribution module may be an extension of and / or a component of the enterprise generative AI system. In some embodiments, a response parser (e.g., response parser 112) resolves the output into chunks.

[0228] In step 1406, the anti-hallucination and attribution module retrieves one or more source paragraphs based on a similarity assessment between the chunk and one or more source paragraphs. In some embodiments, a retrieval unit (e.g., anti-hallucination and attribution retrieval unit 114) retrieves source paragraphs. The retrieval unit may include one or more machine learning models and generative artificial intelligence models, etc. Retrieval unit 114 may collaborate with and / or be included as part of agents (e.g., agents 906-1, 906-2, 906-3, 906-4, etc.) and / or tools (e.g., tools 908-1, 908-2, 908-3, etc.) described herein. Similarity assessment may be performed by retrieval unit 114 and / or may include a similarity assessment for calculating similarity scores associated with the chunk and source paragraphs.

[0229] In some embodiments, an enterprise generative AI system (e.g., enterprise generative AI system 802) generates an infographic (e.g., infographic 220) for individual data records in a set of data records that include one or more source segments. The infographic may describe the relationships between the source segments and one or more other classes, including source images, source tables, and source code. Retrieval of one or more source segments may be based on the infographic (e.g., traversal). Figure 2 (Infographic 220 depicted in the image).

[0230] In step 1408, the anti-hallucination and attribution module attributes at least a portion of one or more source paragraphs to the chunk based on a similarity threshold. The anti-hallucination and attribution module can compare a similarity assessment result (e.g., a score) between the chunk and one or more source paragraphs with the similarity threshold. For example, if the score is equal to or higher than the similarity threshold, the associated (one or more) source paragraphs can be attributed to the chunk. In another example, if the score is lower than the threshold, the associated (one or more) source paragraphs are not attributed to the chunk.

[0231] In step 1410, the anti-hallucination and attribution module combines the chunks with source paragraphs attributed to those chunks. This may include stitching source references to the source paragraphs and / or part or all of the source paragraphs with the chunks.

[0232] In step 1412, the anti-hallucination and attribution module generates a response to the prompt based on the combination. The response may include output with an inline source identifier for identifying the source paragraph to which the attribution is made. The response may include output with at least a portion of the source paragraph to which the attribution is made. The response may be output (e.g., displayed) on one or more systems (e.g., one or more enterprise systems).

[0233] In some embodiments, the output includes sentences, and chunks may include one or more sentences.

[0234] Figure 15 A flowchart 1500 is depicted as an example information retrieval method for an enterprise generative artificial intelligence approach, according to some embodiments. Like other methods described herein, this flowchart 1500 can be combined with other flowcharts (e.g., flowchart 1200). In step 1502, the enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802) identifies enterprise datasets, artificial intelligence applications, and data models from different data domains of one or more enterprise systems based on a received question. The enterprise dataset may include any of documents, document segments, and insights generated by one or more artificial intelligence applications.

[0235] In step 1504, the enterprise generative AI system determines relevance scores associated with the enterprise dataset based on a data model. In some embodiments, an understanding module (e.g., understanding module 916) determines the relevance scores. Each relevance score may be associated with a corresponding portion of the enterprise dataset, and each relevance score is determined relative to other portions of the enterprise dataset.

[0236] In step 1506, the enterprise generative AI system determines information relevant to the problem using one or more generative AI models, based on relevance scores and one or more enterprise access control protocols. In some embodiments, the understanding module determines the information (e.g., based on data retrieved by one or more agents and / or tools).

[0237] Figure 16 A flowchart 1600 depicts an example method, according to some embodiments, for routing requests to different agent programs and verifying responses to enterprise generative artificial intelligence. Like other methods described herein, flowchart 1600 can be combined with other flowcharts (e.g., flowchart 1300). In step 1602, an orchestrator (e.g., orchestrator 904 of enterprise generative artificial intelligence system 802) directs the retrieval request to the different agent programs.

[0238] In step 1604, the enterprise generative AI system receives information from multiple data domains from different agent programs based on instructions from the orchestrator. In some embodiments, the orchestrator module and / or understanding module (e.g., understanding module 916) receive the information.

[0239] In step 1606, the orchestrator analyzes the information to formulate one or more answers to the question, wherein the orchestrator provides additional retrieval requests to at least one of the different agent programs to retrieve additional information for satisfying contextual validation criteria associated with the question.

[0240] In step 1608, the orchestrator outputs empirically validated responses that provide one or more answers to the question and satisfy the context validation criteria.

[0241] Figure 17A schematic diagram 1700 depicts an example computer system for implementing the features disclosed herein, according to some embodiments. Any of the systems, engines, data storage, and / or networks described herein may include instances of one or more computing devices 1702. In some embodiments, the functionality of computing device 1702 is modified to perform some or all of the functionalities described herein. Computing device 1702 includes a processor 1704 communicatively coupled to a communication channel 1716, a memory 1706, a storage unit 1708, an input device 1710, a communication network interface 1712, and an output device 1714. Processor 1704 is configured to execute executable instructions (e.g., programs). In some embodiments, processor 1704 includes a circuit system or any processor capable of processing executable instructions.

[0242] Memory 1706 stores data. Some examples of memory 1706 include storage devices such as RAM, ROM, RAM cache, virtual memory, etc. In various embodiments, working data is stored in memory 1706. Data in memory 1706 can be erased or eventually transferred to storage unit 1708.

[0243] Storage unit 1708 includes any storage unit configured to retrieve and store data. Some examples of storage unit 1708 include flash drives, hard disk drives, optical drives, cloud storage, and / or magnetic tape. Memory system 1706 and storage system 1708 each include a computer-readable medium that stores instructions or programs executable by processor 1704.

[0244] Input device 1710 is any device that inputs data (e.g., a mouse and keyboard). Output device 1714 outputs data (e.g., a speaker or a display). It should be understood that storage unit 1708, input device 1710, and output device 1714 may be optional. For example, a router / switch may include processor 1704 and memory 1706, as well as means for receiving and outputting data (e.g., a communication network interface 1712 and / or output device 1714).

[0245] The communication network interface 1712 can be coupled to a network (e.g., network 808) via link 1718. The communication network interface 1712 can support communication via Ethernet, serial, parallel, and / or ATA connections. The communication network interface 1712 can also support wireless communication (e.g., 802.11, WiMax, LTE, Wi-Fi). Clearly, the communication network interface 1712 can support many wired and wireless standards.

[0246] It should be understood that the hardware components of computing device 1702 are not limited to... Figure 17The computing device 1702 may include more or fewer hardware, software, and / or firmware components (e.g., drivers, operating system, touchscreen, and / or biometric analyzer, etc.) than those depicted herein. Furthermore, hardware components may share functionality and remain within the various embodiments described herein. In one example, encoding and / or decoding may be performed by processor 1704 and / or a coprocessor located on a GPU (i.e., NVIDIA).

[0247] Example types of computing devices and / or processing devices include one or more microprocessors, microcontrollers, reduced instruction set computers (RISC), complex instruction set computers (CISC), graphics processing units (GPUs), data processing units (DPUs), virtual processing units, associative processing units (APUs), tensor processing units (TPUs), vision processing units (VPUs), neuromorphic chips, AI chips, quantum processing units (QPUs), brain-on-a-chip (WSE) engines, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or discrete circuit systems.

[0248] It should be understood that "engine," "system," "database," and / or "database" can include software, hardware, firmware, and / or circuitry. In one example, one or more software programs including instructions executable by a processor can perform one or more of the functions of the engine, database, database, or system described herein. In another example, a circuitry can perform the same or similar functions. Alternative embodiments can include more, fewer, or functionally equivalent engines, systems, databases, or databases, and still remain within the scope of this embodiment. For example, the functionality of various systems, engines, databases, and / or databases can be combined or divided differently. Databases or databases can include cloud storage. It should also be understood that the term "or," as used herein, can be interpreted in an inclusive or exclusive sense. Furthermore, multiple instances may be provided for a resource, operation, or structure described herein as a single instance.

[0249] The data storage described herein can be any suitable structure (e.g., active database, relational database, self-referencing database, table, matrix, array, flat file, document-oriented storage system, and non-relational No-SQL system, etc.) and can be cloud-based or otherwise. The systems, methods, engines, data storage, and / or databases described herein can be implemented at least in part by processors, with one or more specific processors being examples of hardware. For example, at least some operations of the method can be performed by one or more processors or an engine implemented by a processor. Furthermore, one or more processors can also operate to support the performance of related operations in a “cloud computing” environment or as “Software as a Service” (SaaS). For example, at least some operations can be performed by a group of computers (as an example of a machine including processors), where these operations are accessible via a network (e.g., the Internet) and via one or more suitable interfaces (e.g., application programming interfaces (APIs)).

[0250] Some operations can be distributed across processors, residing not only within a single machine but also deployed across multiple machines. In some example embodiments, the processor or processor-implemented engine may reside in a single geographic location (e.g., within a home environment, office environment, or server cluster). In other example embodiments, the processor or processor-implemented engine may be distributed across multiple geographic locations.

[0251] Throughout this specification, multiple instances can implement components, operations, or structures described as a single instance. Although the individual operations of one or more methods are instantiated and described as separate operations, one or more of these operations can be performed concurrently, and the order in which they are instantiated is not required. Structures and functionalities presented as separate components in the example configuration can be implemented as composite structures or components. Similarly, structures and functionalities presented as single components can be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of this document's subject matter.

[0252] In the example implementation, the enterprise generative AI system described in this paper can connect across data stores to one or more virtual metadata repositories, abstract access to different data sources, and support granular data access control maintained by the enterprise AI system. The enterprise generative AI framework can manage virtual data lakes by leveraging enterprise directories connected to multiple data domains and industry-specific domains. The orchestrator of the enterprise generative AI framework can create embeddings for multiple data types across multiple vertical industries and knowledge domains, and even specific enterprise knowledge. Embedding of objects in the data domains of the enterprise information system enables rapid identification and sophisticated processing using relevance scores, as well as additional functionality to enforce access, privacy, and security protocols. In some implementations, the orchestrator module can employ various embedding methods and techniques understood by those skilled in the art. In the example implementation, the orchestrator module can use a model-driven architecture for conceptual representations of enterprise and external datasets, as well as optional data virtualization. For example, a model-driven architecture can be described as in U.S. Patent 10,817,530, issued October 27, 2020, by C3 AI Ltd., Serial No. 15 / 028,340, with priority date January 23, 2015, entitled "Systems, Methods, and Devices for an Enterprise Internet-of-Things Application Development Platform." The type system of a model-driven architecture can be used for objects embedded in a data domain.

[0253] Model-driven architecture handles the compatibility of system objects (e.g., components, functionality, data, etc.) that can be used by orchestrators to dynamically generate queries for searching across a wide range of data domains (e.g., documents, tabular data, insights derived from AI applications, web content, or other data sources). Type systems provide data accessibility, compatibility, and operability across different systems and data. Specifically, type systems address data operability across various programming languages, inconsistent data structures, and incompatible software application programming interfaces. Type systems provide data abstractions that define extensible type models, enabling the dynamic addition of new properties, relationships, and functions without costly development cycles. Type systems can be used as a platform-specific language (DSL) for developers, applications, or UIs to access data. Type systems provide the ability to interact with data to process, predict, or resolve it based on one or more type or function definitions within the type system. An orchestrator is a mechanism for enabling search functionality across various data domains relative to existing query modules, which are typically limited in their searchable data domains (e.g., web query modules are limited to web content, file system query modules are limited to searching the file system, etc.).

[0254] Type definitions can be canonical types declared in metadata using syntax similar to that used for types persisted in relational or NoSQL data stores. The canonical model in a type system is an application-agnostic (i.e., application-independent) model, enabling all applications to communicate with each other in a common format. Unlike standard types, canonical types consist of two parts: a canonical type definition and one or more transformation types. The canonical type definition defines the interface for integration, and the transformation type is responsible for transforming the canonical type into the corresponding type. Using the transformation type, the integration layer can transform the canonical type into the appropriate type.

[0255] Various embodiments of this disclosure include systems (e.g., having one or more processors and a memory storing instructions that, when executed by one or more processors, cause the system to perform the functionality described herein), methods, and nontransitory computer-readable media (or media) configured to: display a graphical user interface; receive a question through the graphical user interface; retrieve information related to the question from different enterprise systems; generate an answer to the question using the information retrieved from the different enterprise systems via a generative artificial intelligence model; and display the answer to the question through the graphical user interface. The systems, methods, and nontransitory computer-readable media (or media) may also be configured to: identify enterprise datasets, artificial intelligence applications, and data models from different data domains of enterprise systems based on the question; determine a relevance score associated with the enterprise dataset based on the data model; and determine information related to the question based on the relevance score and enterprise access control protocols via one or more generative artificial intelligence models. Questions include natural language queries. Enterprise datasets may include any of documents, document segments, and insights generated by one or more artificial intelligence applications. Each relevance score can be associated with a corresponding part of the enterprise dataset, and each relevance score is determined relative to the other parts of the enterprise dataset.

[0256] Various embodiments of this disclosure include systems (e.g., having one or more processors and a memory storing instructions that, when executed by one or more processors, cause the system to perform the functionality described herein), methods, and non-transitory computer-readable media (or media in total) configured to: display a graphical user interface; receive a question through the graphical user interface; manage different agent programs through an orchestrator program to generate answers to the question; retrieve information related to the question through the different agent programs; use the relevant information to generate an answer to the question through a generative artificial intelligence model; and display the answer to the question through the graphical user interface. The system, method, and non-transitory computer-readable medium (or media) may also be configured to: instruct different agent programs to make retrieval requests via an orchestrator; receive information from multiple data domains from the different agent programs based on instructions from the orchestrator; analyze the information via the orchestrator to formulate one or more answers to a question, wherein the orchestrator provides additional retrieval requests to at least one of the different agent programs to retrieve additional information for satisfying contextual validation criteria associated with the question; and output empirical responses that satisfy the contextual validation criteria for one or more answers to the question via the orchestrator.

[0257] The orchestrator can use at least one multimodal model to transform the problem into a series of instructions for enabling different agent programs to retrieve problem-related information. Different agent programs can use one or more machine learning models based on the series of instructions to retrieve problem-related information. The generative AI model and the multimodal model can be the same model. Alternatively, the generative AI model and the multimodal model can be different models. In some embodiments, the system, method, and non-transitory computer-readable medium are also configured to perform traceability analysis that generates natural language output, indicating any of the documents, document segments, and insights from a corresponding portion of one or more enterprise datasets.

[0258] Various embodiments of this disclosure include systems (e.g., having one or more processors and a memory storing instructions that, when executed by one or more processors, enable the system to perform the functionality described herein), methods, and nontransitory computer-readable media (or media in total) configured to: receive output from a generative artificial intelligence model processing a prompt; parse the output from the generative model into chunks to be attributed to one or more source segments; retrieve one or more source segments based on a similarity assessment between the chunks and one or more source segments; attribute at least a portion of one or more source segments to the chunks based on a similarity threshold; combine the chunks with the source segments attributed to those chunks; and generate a response to the prompt based on the combination, the output including an output having an inline source identifier for identifying the attributed source segments. The system, method, and non-transitory computer-readable medium (or media) may also be configured to: filter one or more source paragraphs based on similarity scores; and generate an information graph for each data record in a set of data records that includes one or more source paragraphs, wherein the information graph describes the relationship between the source paragraphs and one or more other classes, including source images, source tables, and source code.

[0259] Generative AI models can employ generative adversarial networks, variational autoencoders, autoregressive models, or recurrent neural networks. Outputs can include sentences, and chunks can consist of one or more sentences. Similarity assessment can include a similarity evaluation (e.g., cosine similarity assessment) used to compute similarity scores associated with chunks and source paragraphs. Responses can be generated by the generative AI model. Retrieval of one or more source paragraphs can be based on infographics.

Claims

1. A method comprising: receiving output from a generative artificial intelligence model processing a prompt; parsing the output from the generative model into chunks to be attributed to one or more source passages; retrieving the one or more source passages based on a similarity assessment between the chunks and the one or more source passages; attributing at least a portion of the one or more source passages to the chunks based on a similarity threshold; combining the chunks with the source passages attributed to the chunks; and generating a response to the prompt based on the combination, the response including the output with inline source identifiers to identify the attributed source passages. the generative artificial intelligence model employs a generative adversarial network, a variational autoencoder, an autoregressive model, or a recurrent neural network.

2. The method of claim 1, wherein, the output includes sentences, and the chunks include one or more of the sentences.

3. The method of claim 1, wherein, the similarity assessment includes a similarity assessment that computes a similarity score associated with the chunks and the source passages.

4. The method of claim 1, wherein, the response is generated by a generative artificial intelligence model.

5. The method of claim 1, wherein, the one or more source passages are filtered based on the similarity score.

6. The method of claim 4, further comprising:

7. The method of claim 6, further comprising: generating an information graph for each data record in a set of data records including the one or more source passages, wherein the information graph describes relationships between a source passage and one or more other classes, the other classes including a source image, a source table, and a source code. the retrieval of the one or more source passages is based on the information graph.

8. The method of claim 7, wherein, 9. A system comprising: one or more processors; and memory storing instructions that, when executed by the one or more processors, cause the system to: receive output from a generative artificial intelligence model processing a prompt; parse the output from the generative model into chunks to be attributed to one or more source passages; retrieving the one or more source passages based on a similarity assessment between the chunks and the one or more source passages; attributing at least a portion of the one or more source passages to the chunks based on a similarity threshold; combining the chunks with the source passages attributed to the chunks; and generating a response to the prompt based on the combination, the response including the output with inline source identifiers to identify the attributed source passages. the generative artificial intelligence model employs a generative adversarial network, a variational autoencoder, an autoregressive model, or a recurrent neural network. the output includes sentences, and the chunks include one or more of the sentences. the similarity assessment includes a similarity assessment that computes a similarity score associated with the chunks and the source passages.

10. The system of claim 9, wherein, the response is generated by a generative artificial intelligence model.

11. The system of claim 9, wherein, the instructions, when executed by the one or more processors, cause the system to filter the one or more source passages based on the similarity score.

12. The system of claim 9, wherein, ​ 13. The system of claim 9, wherein, ​ 14. The system of claim 12, wherein, ​ 15. The system of claim 14, wherein, The instructions, when executed by the one or more processors, cause the system to generate, for each data record in a set of data records comprising the one or more source paragraphs, an information graph, wherein the information graph describes relationships between a source paragraph and one or more other classes, the other classes comprising a source image, a source table, and a source code.

Citation Information

Patent Citations

  • Systems, methods, and devices for an enterprise internet-of-things application development platform

    US10817530B2