ENTREPRENEURIAL GENERATIVE ARTIFICIAL INTELLIGENCE ANTI-HALLUCINATION AND ASSIGNMENT ARCHITECTURE
The anti-hallucination and attribution architecture addresses the issue of inaccurate generative AI outputs by detecting and mitigating hallucinations, ensuring reliable and context-specific insights through chunking and attribution, enhancing enterprise AI systems' accuracy and efficiency.
Patent Information
- Application Number
- DE112024001874
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-04-30
- Publication Date
- 2026-03-05
AI Technical Summary
Conventional generative AI systems suffer from hallucinations, producing inaccurate and inconsistent outputs due to statistical inaccuracy, model overfitting, and data bias, and lack mechanisms to verify or confirm the accuracy of their responses, especially in enterprise environments with incomplete and inconsistent information.
An anti-hallucination and attribution architecture is introduced that detects, prevents, and mitigates hallucinations by breaking down responses into chunks, processing them with an anti-hallucination module to verify accuracy and provide attributions, using a combination of agents and tools to process diverse data sources efficiently, and leveraging domain-specific models for context-specific insights.
Enhances the reliability and accuracy of generative AI outputs by preventing and mitigating hallucinations, providing verifiable attributions, and ensuring context-specific, secure, and efficient information retrieval across enterprise systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL AREA
[0001] This disclosure relates to generative artificial intelligence and machine learning. More specifically, this disclosure relates to anti-hallucination and attribution architectures for enterprise-level generative artificial intelligence. BACKGROUND
[0002] Generative artificial intelligence (generative AI) refers to a subfield of machine learning that deals with algorithms capable of generating new instances of data. These algorithms are typically deep learning models trained on large datasets to learn the underlying statistical properties of the data. Generative AI approaches employ artificial neural networks to mimic the statistical properties of large training datasets. Unlike traditional AI approaches, which focus on analyzing existing data or performing specific tasks based on predefined rules, generative AI can generate entirely new outputs using its learned understanding of the data. Conventional generative AI approaches have several drawbacks, such as…Hallucinations and the inability of users to verify or confirm the responses of the generative artificial intelligence. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 shows a diagram of an example of the logical flow of a generative system for artificial intelligence in enterprises with an anti-hallucination and mapping architecture according to some embodiments. Fig. Figure 2 shows a diagram of an exemplary parsing and information diagram generation process for an enterprise-generating artificial intelligence system with an anti-hallucination and attribution architecture according to some embodiments. Fig. Figure 3 shows a flowchart of an exemplary anti-hallucination and attribution process for generative artificial intelligence systems in enterprises according to some embodiments. Fig. Figure 4 shows a diagram of an example structure of response segments generated by an anti-hallucination and attribution procedure for generative artificial intelligence systems in enterprises according to some embodiments. Fig. Figure 5 shows a diagram of an exemplary architecture and environment of a generative system for artificial intelligence in enterprises, as used in some implementations. Fig. Figure 6A shows an exemplary graphical user interface for enterprise search and the underlying architecture according to some embodiments. Fig. Figure 6B shows an exemplary graphical user interface for generative artificial intelligence in enterprises according to some embodiments. Fig. Figure 7 shows a diagram of an exemplary multi-layered architecture and environment of a generative system for artificial intelligence in enterprises according to some embodiments. Fig. Figure 8 shows a diagram of an exemplary network system for generative artificial intelligence in enterprises according to some embodiments. Fig. Figure 9 shows a flowchart of an exemplary iterative generative artificial intelligence process using unstructured data according to some embodiments. Fig. Figure 10A shows a flowchart of an exemplary iterative process of artificial intelligence using unstructured data according to some embodiments. Fig. Figure 10B shows a flowchart of a non-iterative generative process of artificial intelligence using unstructured data according to some embodiments. Fig. Figure 10C shows a flowchart of a non-iterative generative process of artificial intelligence that uses unstructured data, according to some examples. Fig. Figure 11 shows a flowchart of an exemplary iterative generative process of artificial intelligence using unstructured data according to some embodiments. Fig. Figure 12 shows a flowchart of an exemplary generative artificial intelligence process for businesses according to some embodiments. Fig. Figure 13 shows a flowchart of an exemplary generative method for artificial intelligence in enterprises, which uses an agent-orchestrator architecture according to some embodiments. Fig. Figure 14 shows a flowchart of an example of an anti-hallucination and attribution procedure for generative artificial intelligence systems in enterprises according to some embodiments. Fig. Figure 15 shows a flowchart of an exemplary procedure for information gathering for generative artificial intelligence systems in companies according to some embodiments. Fig. Figure 16 shows a flowchart of an example procedure for forwarding queries to different agent programs and for validating responses for generative artificial intelligence systems in enterprises according to some embodiments. Fig. Figure 17 shows a diagram of an exemplary computer system for implementing the features disclosed herein according to some embodiments. DETAILED DESCRIPTION
[0003] Generative AI is an artificial intelligence technology that uses machine learning algorithms to perform tasks that mimic human cognitive intelligence and generate content. This content can take the form of text, audio, video, images, and more. However, conventional generative AI processes often present distorted or erroneous information as a result of a hallucination (or simply a hallucination) by the machine learning model. In generative AI, an artificial hallucination, confabulation, or deception refers to a specific type of output generated by the model that deviates from factual accuracy, basic truth, or the intended context.Conventional generative AI can produce content that does not correspond to reality or verifiable information, typically due to statistical inaccuracy, model overfitting, and data bias. The generated content may appear plausible and internally coherent, but it is inaccurate and inconsistent with the given command or surrounding information. Corporate environments can also contain conflicting information, further increasing the likelihood of hallucinations. Hallucinations in corporate environments can be exacerbated by often incomplete and inconsistent information scattered across a variety of disparate and incompatible corporate systems. Current generative AI systems are unable to detect, prevent, or mitigate hallucinations.Conventional methods of generative artificial intelligence do not offer users a mechanism to verify, validate, or confirm the results of the generative AI. Consequently, users have no way of discerning whether the results of the generative AI are accurate or the product of hallucinations.
[0004] This document discloses an anti-hallucination and attribution architecture for generative AI systems in enterprise environments. This architecture enhances the accuracy and reliability of generative AI content (e.g., responses) by detecting, preventing, and mitigating hallucinations. Furthermore, the hallucination and attribution architecture can be added as a separate tool or module to existing generative AI systems, allowing it to work seamlessly with them without requiring any system modifications or redesigns. The anti-hallucination and attribution architecture can also be implemented with minimal impact on ongoing production systems.
[0005] The generative artificial intelligence systems for enterprises described here can transform the way we interact with enterprise data, fundamentally changing the human-computer interaction (HCI) model for enterprise software. Organizations running sensitive workloads in cloud-native, on-premises, or air-gapped environments can implement generative AI architectures to generate enterprise-wide insights. This is achieved by leveraging tools for rapidly discovering and retrieving information, with agents that develop and coordinate complex operations in response to simple, intuitive input.The enterprise generative AI architecture enables business users to ask open-ended, multi-stage, context-specific questions that are processed using generative AI and machine learning to understand the query, identify relevant information, and generate new context-specific insights with predictive analytics. The enterprise's generative AI architecture supports simplified human-computer interactions with an intuitive natural language interface and advanced accessibility features for adaptive input formats, including but not limited to text, audio, video, images, and more.
[0006] The anti-hallucination and mapping architecture includes an anti-hallucination and mapping module that can work with the generative artificial intelligence systems and components used in the enterprise (e.g., retriever tools, large language models, or other generative artificial intelligence models, etc.) to detect, prevent, and mitigate hallucinations caused by the large language models or other models used (e.g., generative artificial intelligence models, multimodal models). For example, a user might send a command (e.g., a question, a query, etc.) to the enterprise's generative artificial intelligence system, such as..."How many different engineers did John Doe work with in his engineering department?" This might require the company's generative artificial intelligence system to identify John Doe, identify John Doe's department, determine the engineers in that department in a third iteration, ascertain which of these engineers John Doe worked with, and finally combine these results to generate the answer to the query. With such complex queries, hallucinations can occur at any step in the process of generating the answer. For example, there might be multiple John Does within an organization, and the system could easily hallucinate to determine which John Doe the query refers to.The anti-hallucination and mapping architecture can prevent or mitigate such hallucinations by forwarding the response to an anti-hallucination and mapping module before generating a final response.
[0007] The anti-hallucination and attribution module can break down the response generated by the company's generative artificial intelligence system into multiple chunks (e.g., several sentences). The module can then process these chunks, along with the original passages retrieved by the AI, to identify relevant passages for each chunk. Finally, the module can combine these chunks and the retrieved relevant passages to generate an attribution. This attribution can include, for example, source identifiers that specify the documents and passages used to create the response. The system thus provides a verifiable attribution that confirms the reliability of the response and that the model did not hallucinate.If the anti-hallucination and mapping module cannot locate any relevant passages, it may determine that the model has hallucinated and either rerun the query to find a different result or notify the user that no reliable result could be obtained, rather than simply giving the user the answer.
[0008] The generative artificial intelligence systems for enterprises described here can also use a combination of agents and tools to efficiently process a variety of inputs from different data sources (e.g., with different data formats) and return results in a common data format (e.g., natural language). The enterprise generative artificial intelligence anti-hallucination and mapping architecture includes an orchestrator agent (or simply orchestrator) that monitors, controls, and / or otherwise manages many different agents, tools, and / or modules (e.g., an anti-hallucination and mapping module). Orchestrators can contain one or more machine learning models and perform monitoring functions, such as forwarding inputs (e.g.,Queries, sets of instructions, natural language input, or other human-readable or machine-readable input are sent to specific agents to perform a set of prescribed tasks (e.g., retrieval requests prescribed by the orchestrator to answer a query). Machine learning models can include some or all of the different types or modalities of models described here (e.g., multimodal machine learning models, grand language models, data models, statistical models, audio models, visual models, audiovisual models, etc.). Agents can contain one or more multimodal models (e.g., grand language models) to perform the prescribed tasks using a variety of different tools. Different agents can use different tools to make unstructured data retrievals, structured data retrievals, API calls (e.g.,Tools can be used to access and process insights from artificial intelligence applications, etc. They can include one or more specific functions and / or machine learning models to perform a particular task (or set of tasks).
[0009] Agents can adapt and perform differently depending on the context. A context might be a specific domain (e.g., an industry), and an agent might use a particular model (e.g., a large language model, a different machine learning model, and / or a data model) trained on industry-specific datasets, such as healthcare datasets. The agent might use a healthcare model when receiving healthcare input, and it can also easily and efficiently adapt to use a different model based on different input or context. In fact, some or all of the models described here can be trained for specific domains in addition to, or instead of, more general purposes.The generative artificial intelligence architecture for enterprises uses domain-specific models to achieve accurate, context-specific search results and insights.
[0010] The orchestrator manages the agents to efficiently process different inputs or different parts of an input. For example, an input might require the system to access and retrieve records from various data sources (e.g., unstructured data stores, structured data stores, time-series data stores, etc.), database tables from different database types, and machine learning insights from various machine learning applications. The different agents can process each of these requests separately and in parallel, significantly increasing computational efficiency.
[0011] The agents can process the diverse data returned by the various agents and / or tools. For example, large language models typically receive input in natural language format. The agents can receive information in a non-natural language format (e.g., database table, image, audio) from a tool and convert it into natural language that describes the tool's output in a format understood by large language models. A model (e.g., a large language model, a multimodal model) can then process this input to generate an initial response that the anti-hallucination and attribution module can validate and / or attribute before providing a final output.
[0012] Fig. Figure 1 shows a diagram illustrating an example of the logical flow of a generative artificial intelligence system in an enterprise with an anti-hallucination and attribution architecture, according to some implementations. General generative artificial intelligence models (e.g., large language models, multimodal models) have varying capabilities and follow instructions at different performance levels. In one example, an enterprise generative artificial intelligence system can generate attributions (e.g., source citations) for responses produced by the enterprise's generative artificial intelligence system (e.g., a large language model) by instructing the generative artificial intelligence model to cite its sources. However, this approach has several drawbacks and limitations.
[0013] More precisely, different models may or may not follow this instruction, and this can depend on the specific model used by the company's generative artificial intelligence system, as well as the task or request being received. The formatting of the hash definition of the source that the company's generative AI system uses to obtain citations can vary from client to client, which in turn affects citation performance. This approach can also limit the system's ability to cite only from the originally retrieved documents. This can lead to the system ignoring important or relevant documents (e.g.,(because the system did not choose a sufficiently high value for the number of documents retrieved, which can vary from query to query) and prevents the system from realizing the potential benefits of using a finely tuned model on a customer corpus in conjunction with the company's generative artificial intelligence system.
[0014] Furthermore, while this attribution approach can reduce hallucinations to some extent, it does not completely or optimally prevent them. Since generating attributions and avoiding hallucinations are critical for generative artificial intelligence models in business, the model's adherence to instructions (e.g., instructions to cite sources) cannot be left to chance. The generative artificial intelligence systems for business described here may be model-independent, and some models may implement such an approach very poorly. That said, this attribution approach is compatible with various types of models, including proprietary or open-source large language models (LLMs), small language models, image generation models, audio generation models, video generation models, omnimodal models, and so on.Therefore, an approach is needed that not only prevents hallucinations but is also model-independent.
[0015] For example, a company's generative artificial intelligence model does not prompt the model to cite its sources but performs the attribution in a separate step after the model has generated its answer. This addresses the limitations discussed above, and the system can also draw on information encoded in a finely tuned model within its corpus when using information from retrieved passages. The system can combine local and global corpus searches to provide better and more informative answers. This allows the system to use this new component (e.g., the anti-hallucination and attribution module 110) as a separate tool for verifying statements and providing corroborating evidence.
[0016] An overview of the current architecture for anti-hallucination and attribution is in Fig. 1 shown. In the example of Fig. 1. A query 102 is received (e.g., a question posed by a user via a graphical user interface). The query 102 is forwarded to the retriever 104 to retrieve passages (e.g., from a vector memory) relevant to the query (e.g., using a similarity score that generates a corresponding similarity value). The retriever 104 passes the retrieved passages through a context processor 106, which forms an instruction that is then provided to the generative artificial intelligence model 108 to generate the response. In the new version, the response is not only provided to the user, but the system also applies the anti-hallucination and mapping module 110 to this response. The response may first pass through a response parser 112, which breaks the response down into chunks (or segments) to be mapped.These chunks and the originally retrieved passages are processed by the anti-hallucination and attribution module 114 to find and retrieve relevant passages for each chunk (e.g., from a vector memory). The chunks, along with the corresponding retrieved passages, are then provided to the response renderer 116, which combines the retrieved passages and the chunks to generate the associated response 118.
[0017] The Anti-Hallucination and Attribution Module 110 can parse and chunk the response. To do this, the Anti-Hallucination and Attribution Module 110 can split, segment, or select a subset of the response into sentences and related parts. These pieces or parts are then combined by the Anti-Hallucination and Attribution Module 110 into larger chunks until a maximum number of tokens is reached in each chunk. To calculate the tokens, the Anti-Hallucination and Attribution Module 110 can use a tokenizer chosen to be the same tokenizer used by the embedding model in the Anti-Hallucination and Attribution Retriever 114. Each chunk is then processed by the Anti-Hallucination and Attribution Retriever 114 to generate relevant passages for each chunk.In this phase, the Anti-Hallucination and Attribution module 110 has the ability to restrict or filter the passages used for attribution to the original passages retrieved at the beginning of the pipeline (e.g., by setting a tag value). These passages are then forwarded, along with the chunks, to the Response Renderer 116 to generate the assigned response 118. During this phase, the Anti-Hallucination and Attribution module 110 can include multiple passages to support each chunk. This can be controlled by a parameter value. For example, if a value greater than 1 is set, the Anti-Hallucination and Attribution module 110 can only include additional passages for each chunk if their score exceeds a certain threshold set by the Anti-Hallucination and Attribution Retriever 114.The Anti-Hallucination and Attribution Module 110 can always contain at least one quote for each chunk, although it can only contain more quotes if the score assigned to the passage is greater than the threshold. The Anti-Hallucination and Attribution Module 110 can also provide a visualization, such as a color code, based on the maximum score of the quoted passages for the chunk.
[0018] For example, an initial response generated by the generative artificial intelligence model 108 might include the following: Parcel B1 is an approximately 40-50 acre site and adjacent shipyard infrastructure specifically constructed for maneuvering (with dredged berth pockets and turning basins) (“Parcel B1”). Development of Parcel A is expected to be completed in the second quarter of 2024, while development of the remainder of Phase 1 is expected to be completed in 2024 and 2025. Development of Phase 2 is scheduled to begin in 2024 and be completed by 2028. The cost of developing Phases 1 and 2 is estimated at approximately $550 million each, not including dredging to be carried out by the Department of Transportation. The Department is issuing bonds to help cover a portion of these costs.
[0019] The corresponding attributed response generated by the anti-hallucination and attribution module 110 may include the following: Parcel A is expected to be completed in the second quarter of 2024, while development of the remainder of Phase 1 is expected to be completed in 2024 and 2025. Development of Phase 2 is expected to begin in 2024 and be completed by 2028 (from
[143] with a score of 0.7580196261405945). The cost of developing Phases 1 and 2 is estimated at approximately $550 million each, excluding hydraulic dredging being carried out by the Department of Transportation. The Department is issuing bonds to help cover some of these costs (from
[143] with a score of 0.8340072631835938). Parcel B1 is a plot of land of approximately 40-50 hectares and an adjacent quay facility that was specially built for maneuvering (with dredged berth pockets and a turning basin) (“parcel B1”) (from
[143] with a score of 0.7503519654273987).The development of parcel A is scheduled for completion in the second quarter of 2024, with the development of the remainder of Phase 1 expected to be completed in 2024 and 2025. The development of Phase 2 is scheduled to begin in 2024 and be completed by 2028 (from
[143] with a score of 0.7610742449760437).
[0020] In some embodiments, the anti-hallucination and attribution module 110 can visually mark (e.g., highlight, color-code) the associated response. For example, the color yellow can be used to highlight chunks (e.g., sentences) that have a relatively low similarity value (e.g., relative to a threshold). The color red can be used to mark chunks for which the anti-hallucination and attribution retriever 114 has not found any relevant passages. The color green can be used to highlight chunks that have a relatively high similarity value (e.g., relative to the threshold), indicating that this chunk was not generated by any hallucinations of the generative artificial intelligence model 108.
[0021] In another example, an original answer generated by the generative artificial intelligence model 108 might be: No, you do not have to pay federal and state taxes on the interest from the 2023 Series A bonds, as it is not exempt from gross income for federal income tax purposes and is not considered gross income under the existing Gross Income Tax Act. However, it is recommended that you consult a tax advisor for further information.
[0022] The corresponding response generated by the Anti-Hallucination and Attribution Module 110 may include the following: No, you do not have to pay federal and state tax on the interest from the 2023 Series A bonds because it is not excluded from gross income for federal income tax purposes and is not considered gross income under the current Gross Income Tax Act (from
[161] with a score of 0.7880104780197144). However, it is recommended that you consult a tax advisor for further information (from
[161] with a score of 0.4502568244934082).
[0023] Fig. Figure 2 shows Diagram 200 of an exemplary parsing and information diagram generation process for an enterprise generative artificial intelligence system that uses an enterprise generative artificial intelligence anti-hallucination and mapping architecture according to some embodiments. The exemplary process of preprocessing data sets and retrieving information can be performed by one or more of the systems and / or subsystems described herein (e.g., the enterprise generative artificial intelligence system 802). The information diagram can be used to facilitate information retrieval, accurately mapping source code chunks to a response, and avoiding hallucinations.
[0024] In general, datasets can contain various modalities of information, such as plain text, tables, images, code, video, audio, and the like. To effectively and reliably retrieve information for a query, the preprocessing and information retrieval process offers a multimodal approach to extracting information from the datasets.
[0025] The preprocessing and information retrieval process can comprise three stages. The first stage can include parsing and extracting various modalities from these documents. This process can be performed in parallel (or essentially in parallel) for all the different modalities.
[0026] In one example, text and code can be parsed (step 202). Depending on the file format, extracting textual information—that is, plain text and code—can be challenging and may require the use of different libraries. Regardless of the dataset type, the process can involve various steps to prepare for subsequent stages. Depending on the file format, some of these steps can become more complex.
[0027] One of these steps may involve extracting text information (step 204) from the data records (step 201) so they can be further parsed. For example, everything that is not an image or a table is extracted. The output of this step may include all text information (e.g., information other than image and table captions) from the data records (it must include this), which can then be used for further separation of text and code. The parsed results can exhibit high fidelity (e.g., no introduction of random whitespace or strange characters that would break the meaning of sentences) and are robust with respect to font, size, color, and position relative to other page elements.
[0028] A further step can be the separation of text and code (step 205). The goal of this step is to recognize and separate code and plain text. This allows the system to process these modalities separately. Another step can be the chunking and parsing of text and code (step 206). After separating the text and code modalities, the purpose of this step is to identify, locate, and extract related pieces of code (step 207) and to chunk the text content (step 206) in a meaningful and coherent way (i.e., no cutting mid-sentence or mid-paragraph, especially due to page breaks, and where possible, chunks with related topics).
[0029] The system (e.g., the company's generative artificial intelligence system 802) can use an object-oriented structure in which there can be classes for plain text and code. These classes can have at least fields that track the content, its position in the document, and the number of tokens of the content (this means, for example, that the system should know the tokenizer for this purpose). For the text class, the system may already be able to track which code snippets have been removed from (or are associated with) the content of the chunk. This may already be implemented as part of the system.
[0030] In another example, tables can be parsed. Tables are an important feature in datasets. To perform effective information retrieval, the system can first fully locate and identify tables (e.g., because tables may span multiple pages / records / segments or appear in different structures on a single page). See Step 206. The system can identify libraries that provide these functions and measure their performance in fully identifying tables. Once the tables are identified, they can be extracted as an image, data frame, or similar format (Step 209). The system can also extract the table label or title, column headers, and possibly the row index(s), etc., as associated metadata (Step 209).Similar to the text and code classes, the system can also have a table class that tracks the extracted table content, its position, title / label, etc.
[0031] In another example, images can be parsed. Similar to the handling of tables, the system can also begin with the complete identification and localization of images (step 212). The system can include an image class to track the images in the documents, i.e., the content, the image's position in the document, and its caption or title as associated metadata. The system can also extract images as associated metadata (step 213).
[0032] At the end of the first stage, the system can have multiple instances of the classes Text, Code, Table, and Image, with different modalities represented in each data record. The system can then proceed to the second stage, namely the creation of an information diagram for each data record.
[0033] In the second phase, to facilitate effective information retrieval, the system can represent the information in each data record 201 by creating and / or using an information diagram 220. The nodes 230, 241, 243, and 245 of this graph correspond to the different instances of the four modality classes from the first stage. It is a (directed) two-sided graph whose edges lead from the text nodes to all other modality nodes. Creating these edges is the main objective of this phase.
[0034] In the example of Fig. An edge exists between a text node 230 and other modality nodes 241, 243, and / or 245 if a reference (e.g., a relationship) exists between them. This can be determined based on direct references to them in the text chunk, based on proximity, or even based on contextual similarity with their label or title. Once the edges have been identified, the system can track which edges exist between each text node 230 and other modality nodes 241, 243, and / or 245 as part of the text class. Once the system has fully specified this graph, it can design an information retrieval process that utilizes this graph 220. For example, agents (e.g., Agents 906) and tools (e.g., Tools 908) can use the graph to retrieve information.In another example, the anti-hallucination and mapping module 110 can use graph 220 to map sources to a response and generate a mapped response. For this purpose, information graph 220 can also contain access control protocol information (e.g., as defined by the enterprise access control layer 515 and / or access control module 918).
[0035] In the third stage, the system can outline a process for information retrieval using the information diagram 220. One approach begins with the initial embedding of the content of the text nodes (and / or other textual metadata associated with other modalities) (step 229) and the storage of these embeddings in a vector memory 226. Upon a query 222, the system can embed these (step 228) and find the most relevant chunks or text nodes 230 associated with it. This would be the entry in the graph. At this point, the system can follow the outgoing edges to other modality nodes. The classes associated with these modalities (e.g., code, image, and table) may have a method that enables the generation of relevant insights in response to the query.This method can be supported by various approaches, including multimodal models or other tools for understanding and querying the specific modalities. These insights can then be combined with the text chunks and the query in an aggregator 250 to form a text body or a command used to query a model (e.g., a multimodal model, a large language model, etc.).
[0036] In an implementation, the in Fig. The functionality shown and described in Figure 2 can be executed by a chunking module (e.g., chunking module 910) and / or an embedding generator module (e.g., embedding generator module 912). For example, steps 204–213 can be executed by the chunking module and steps 228–229 by the embedding generator module. In some embodiments, the aggregator 250 comprises part of an orchestrator module (e.g., orchestrator module 904) and / or an understanding module (e.g., understanding module 916).
[0037] Fig. Figure 3 shows a flowchart 300 of an exemplary anti-hallucination and matching process for generative artificial intelligence systems in enterprises. In general, the anti-hallucination and matching process (e.g., performed by an anti-hallucination and matching module 110) is a completely self-contained process (and module) that can be used to validate and confirm any response (e.g., a response from a generative artificial intelligence model or a human-generated response), whether or not it contains inline citations.
[0038] In step 302, a response is received (e.g., from an enterprise generative artificial intelligence system, a generative artificial intelligence model, a human user, etc.). The response can be received by an anti-hallucination and attribution module (e.g., anti-hallucination and attribution module 110). In step 304, the anti-hallucination and attribution module determines whether the response already contains quotations (e.g., provided by the enterprise generative artificial intelligence model during the response generation process). If the anti-hallucination and attribution module determines that the response is associated with quotations (e.g., inline quotations), the response can be segmented by extracting the segmentation based on the available quotations (step 306).If the Anti-Deception and Attribution module detects that there are no quotations in the response, it performs generic segmentation (step 308). More specifically, generic segmentation is generally performed based on the number of tokens and sentences in the response. More precisely, for each segment, the Anti-Deception and Attribution module can insert as many sentences into the segment as necessary until it reaches a maximum token limit for that segment.
[0039] In step 310, the anti-hallucination and matching module assigns a set of sources based on a relevance or similarity search. The module can filter the source texts (or sources) based on the resulting scores. This filtering can compare the resulting scores against one or more thresholds. For example, if a score is at or above the threshold, the corresponding source texts can be assigned to the segment. If the score is below the threshold, the corresponding source texts are not assigned to the segment. The similarity search can be an extension of a similarity value (e.g., the "classic" cosine) and / or a function that can be overridden by an end user.
[0040] The anti-hallucination and attribution module uses a mapping between the response segments and the sources with the relevance score. The anti-hallucination and attribution module can also quantify / extract a credibility score for each source. The anti-hallucination and attribution module can then perform source segmentation (step 312). More specifically, the anti-hallucination and attribution module can use the same segmentation logic that was used for response segmentation. The anti-hallucination and attribution module calculates the pairwise relevance or similarity score between the response and source segments (step 314).
[0041] In step 316, the anti-hallucination and attribution module extends this mapping by adding a confirmation rating / label for each pair within the four categories of contradiction, support, neutral, and implication. These can then be used to further refine keys and values in this mapping. The "supportive" label can be added to potentially compensate for shortcomings related to the generic segmentation.
[0042] At this point, the anti-hallucination and attribution module has established and quantified the relationship between the response and source segments. This is shown in Fig. This is illustrated in a diagram showing all the different scores that define the relationships between the response and source segments. These can, in turn, be used to quantify a singular score assigned to each response segment (e.g., to quantify its anchoring in the sources) and a singular score assigned to each source segment (e.g., in the context of the corresponding response segment). See Step 318. In some embodiments, this process can incorporate user-defined heuristics for calculating some or all of these scores and for calculating the source credibility score based on the source's available metadata, either during inference or when ingesting the sources.
[0043] Fig. Figure 4 shows a diagram 400 of an example structure of response segments generated by an anti-hallucination and attribution procedure for generative artificial intelligence systems in enterprises according to some embodiments. More specifically, it comprises Fig. 4 Response segments 402, Sources 404, Relevance ratings 406 for the sources 404, Credibility ratings 408, Source segments 410, Relevance ratings for source segments 412 and Confirmation ratings 414.
[0044] Fig. Figure 5 shows a diagram of an exemplary generative artificial intelligence system architecture and environment 500 for enterprises according to some embodiments. In the example of Fig. 5 comprises the system architecture and environment 500, a generative artificial intelligence system 502, enterprise systems 504, external systems 506, domain models 508, a vector data store 526, an embedding model data store 524, and an access control layer 515. In some embodiments, the query comprises a natural language query received through a graphical user interface. In some embodiments, the one or more enterprise datasets comprise any documents, document segments, and insights generated by the one or more artificial intelligence applications. In some embodiments, each of the relevance ratings is associated with a corresponding part of the one or more enterprise datasets, and each of the relevance ratings is determined relative to the other corresponding parts of the one or more enterprise datasets.
[0045] Each data model within the multitude of data models can correspond to a different data domain within the multitude of different data domains. In some embodiments, each data model represents the respective relationships and attributes of the corresponding different data domain from the multitude of different data domains. The respective relationships and attributes include any data types, data formats, and industry-specific information. In some embodiments, the natural language output includes a summary of at least one of the respective parts of the one or more enterprise datasets associated with a relevance score.
[0046] In general, an enterprise's generative artificial intelligence system 502 can be used to securely query and process enterprise data and applications across various areas of an enterprise information environment. This can be referred to as generative enterprise search (or simply enterprise search). As shown, the generative artificial intelligence system 502 can receive a query 516 (e.g., an input, a command, a natural language query, an instruction, etc.). In general, the enterprise's generative artificial intelligence system 502 can process the query using large language models 520 and retrieval models 522. More specifically, the enterprise's generative artificial intelligence system 502 can use the large language models 520 to interpret, understand, and / or parse queries.The query models 522 can interact with the large language models 520 and the domain models 508 to retrieve data sets (e.g., documents, images, application output, artificial intelligence insights, objects, etc.) across different domains, using specific data models 512 for each domain. Accordingly, the generative artificial intelligence system 502 can use the large language models 520, the query models 522, and the domain models 508 to produce an accurate, reliable, and secure enterprise search result.
[0047] The enterprise's generative artificial intelligence system 502 can facilitate the ingestion and storage of enterprise system data from enterprise systems 504 and / or external system data from external systems 506 (e.g., systems outside the enterprise's information environment) using connectors 510, data models 512, and various persistent storage mechanisms and techniques 514. Enterprise systems 504 may include CRM systems, EAM systems, ERP systems, and the like, and the connectors may facilitate data ingestion from various data sources (e.g., Oracle systems, SAP systems, and / or the like). In some embodiments, the data models 512 provide attributes, relationships, and / or functions associated with a particular domain. Domains may include, for example, aerospace domains 512-1, energy domains 512-2, defense domains 512-3, and / or the like.The domain models 508 can enable the generative artificial intelligence system 502 to deliver domain-specific results without compromising the security or integrity of the underlying enterprise data, systems, and applications.
[0048] Furthermore, the generative artificial intelligence system 502 utilizes or manages an enterprise-wide access control layer 515, which can offer numerous technological advantages. In some embodiments, the enterprise access control layer 515 facilitates the separation of the underlying enterprise information (e.g., enterprise data, applications, systems) from the large language models 520 and / or other machine learning models of the enterprise generative artificial intelligence system 502. Accordingly, the generative artificial intelligence system 502 can deliver domain-specific deterministic results without requiring the large language models 520 and / or other machine learning models to be trained on such enterprise information, which can lead to the numerous problems described above (e.g., information loss, hallucinations).
[0049] In some embodiments, the enterprise generative artificial intelligence system 502 can use the enterprise access control layer 515 to implement additional enterprise controls. For example, an enterprise information environment may include users and systems with varying levels of enterprise permissions. The enterprise access control layer 515 can ensure that responses or outputs comply with the access and security protocol. The enterprise generative artificial intelligence system 502 protects information by granting a user access based on permissions, profiles, and controls. In one example, the enterprise access control layer 515 can filter information constrained by the query model 522 before it is processed by the large language models 520 or before the response or other output is presented.More specifically, the access control layer 515 can filter data sources, records and / or other elements of an enterprise information environment so that query responses (or supporting traceability references) do not contain information that the user is not authorized to access.
[0050] The company's generative artificial intelligence system can also perform similar functions based on the context of the users and / or systems making the query. For example, a director and an engineer might make the same query (e.g., "Which projects are overdue?"), and the company's 502 generative artificial intelligence system can use contextual information (e.g., user role, permissions, domain associated with the user, etc.) to provide a response that is context-dependent both in terms of content (e.g., information about overdue projects relevant to the specific requester) and in terms of how the response is presented (e.g., an engineer might receive more detailed technical information, while a director might receive less technical detail).
[0051] In some embodiments, the enterprise generative artificial intelligence system (502) can crawl, index, and / or map a corpus of data records (e.g., records from one or more enterprise systems or environments), using contextual information (e.g., contextual metadata) along with record embeddings to enable access control (e.g., role-based access), improved identification and querying of records, and mapping of relationships between records. For example, contextual information can prevent some users from accessing (e.g., viewing, retrieving) certain records and improve the similarity assessments used in retrieval operations (e.g., in a generative artificial intelligence process).
[0052] In some implementations, the enterprise's generative artificial intelligence system 502 can generate embeddings based on the embedding models of the embedding model data store 524 and the content of the data records. In some implementations, the embeddings can be represented by one or more vectors that can be stored in the vector data store 526. In some implementations, the retrieval models 522 can use the embeddings to retrieve relevant data records and perform similarity or relevance assessments or other aspects of retrieval operations. As used here, data records can be unstructured data records (e.g., documents and text data stored in a file system in a format such as PDF, DOCX, MD, HTML, TXT, PPTX, image files, audio files, video files, application output, and the like), structured data records (e.g., structured data records, e.g.,Database tables or other data sets stored according to a data model or type system), time series data sets (e.g., sensor data, insights from artificial intelligence applications) and / or other types of data sets (e.g., access control lists).
[0053] Fig. Figure 6A shows an exemplary graphical user interface 600 for enterprise search and the underlying architecture 602-608 according to some embodiments. In general, the features described in Figure 6A include: Fig. 6A and Fig. The graphical user interfaces shown in Figure 6B represent a human-computer interface for receiving queries in natural language and displaying relevant information from the company's information environment in response to those queries. Although in Fig. 6A or Fig. If not shown in 6B, the relevant information may also include visualizations, predictive analyses, control commands and / or other information obtained or generated by the company's artificial intelligence generative systems.
[0054] In the example of Fig. Figure 6A of the diagram includes an enterprise query input interface (600) and a framework (602-608) for unifying information access methods and application operations across legacy and new enterprise applications and a growing number of data sources in various enterprise environments. The enterprise generative artificial intelligence systems described here can harmonize access to information and increase the availability of complex application operations while maintaining the organization's security and privacy controls. This framework uses machine learning techniques (e.g., generative artificial intelligence algorithms and models) to navigate enterprise information and applications, understand organization-specific context queues (e.g., acronyms, nicknames, jargon, and the like), and determine the appropriate response for a request (e.g., a query, question, command, etc.).) to locate the most relevant information. This can, for example, lower the learning curve and reduce the steps a user has to take to access information, thereby democratizing the use of information that is currently blocked by the complexity and expertise required by traditional enterprise information systems.
[0055] In the example of Fig. 6A comprises the enterprise query input interface 600, a graphical user interface element configured to receive various inputs, such as a natural language query. For example, a user might ask the system, "How can my breaks be scheduled?", and an enterprise generative artificial intelligence system can process the query using a natural language processing component 602 and one or more pre-trained generative transformers 604 (e.g., large language models and / or other machine learning models) to generate a secure, accurate, and reliable response based on a variety of different enterprise data stores and applications 608.More specifically, the natural language processing component 602 and the generative, pre-trained transformers 604 can generate one or more new queries 606 to process the original user query. For example, the first new query might be an SQL query to retrieve records from relational database management systems, while the second query might be a different type of query (such as a set of commands) to execute one or more applications so that the result of this application execution can be returned and used by the system to generate the enterprise search response. An example interface for the enterprise search response is shown in [reference to relevant documentation]. Fig. 6 is shown and is described below.
[0056] In some embodiments, the in Fig. 6A and Fig. The search interfaces for enterprises shown and described in 6B are generated by the generative systems for artificial intelligence in enterprises described herein, and the framework 602-608 can represent the generative system architectures and environments for artificial intelligence in enterprises described herein.
[0057] Fig. Figure 6B shows an exemplary graphical user interface 650 for the generative artificial intelligence of an enterprise according to some embodiments. In some embodiments, the graphical user interface 650 for generative artificial intelligence in enterprises can be generated, at least partially, by the generative artificial intelligence systems in enterprises described here. In the example of Fig. 6B contains the graphical user interface 650 for enterprise generative artificial intelligence, an input section 652 for an enterprise search query, a generative section 654 for an enterprise search result, and an interactive query section 656.
[0058] The input part 652 for enterprise search presents an enterprise search query 658. In some implementations, the query can be entered via input part 652, although it can also be entered via another interface (e.g., the one in Fig. 6A (interface 600 shown) can be entered and displayed in input section 652 as part of the generative enterprise search response.
[0059] Part 654 of the generative enterprise search result contains a type of AI-generated response 660, a status of the AI-generated response 662, a generative AI-generated enterprise search result 666, source data sections 668 used to generate the response, source identifications 669, and feedback elements of the AI-generated response 670. In the example of Fig. 6B is the response type 660, a summary, although other response types may exist that generative artificial intelligence systems in businesses can produce. The response status indicates the status of the response. The response status can be, for example: query processed, metric evaluated, documents searched, response generated (e.g., result), visualization generated (e.g., a time series visualization for display in a graphical user interface of the response), and / or fully generated (e.g., as shown in [reference to relevant document / document]. Fig. 6B shown).
[0060] Source data sections 668 contain at least some of the information from the source data used to generate the response. This can, for example, allow the response to be confirmed without having to independently verify it. Source identifiers 669 identify the source data records used to generate the response. Source identifiers can include, for example, an entity name, a domain type, a description or name of the record (e.g., service manual, user manual, technical manual, and / or similar), and / or a record type (e.g., a document, more specifically, a PDF document), and / or similar information. This can also ensure traceability and allow the user to trust the response.
[0061] The response feedback section 670 allows users to provide feedback on the response (e.g., positive or negative feedback). Generative AI systems in businesses can, for example, use the received feedback to improve themselves (e.g., through reinforcement learning).
[0062] The interactive query section 656 allows users to enter additional related queries (e.g., "follow-up questions") via an interactive input section 657. In the example of Fig. 6B The interactive query section 656 includes a chat interface, although other interfaces may use different interactive query sections. The interactive query section 656 also includes a system-generated message 674 that prompts the user to ask follow-up questions, and users can ask additional related queries 676. The interactive query section 656 may also include a status section 678 that indicates the processing status of the additional related query 676. The status may include, for example, information such as: query processed (e.g., as in Fig. 6B shown), processed query, evaluated metric, searched documents, generated response (e.g. result), generated visualization (e.g. a time series visualization for display in a graphical user interface for the response) and / or fully generated.
[0063] Time series refers to a list of data points in chronological order that can represent the change in value of data relevant to a specific problem over time, such as inventory levels, equipment temperatures, financial values, or customer transactions. Time series provide the historical information that can be analyzed by generative and machine learning algorithms to build and test predictive models. In example implementations, time series data is cleaned, normalized, aggregated, and combined to represent the state of a process over time and to identify patterns and correlations that can be used to create and evaluate predictions applicable to future behavior.
[0064] Fig. Diagram 700 shows an exemplary multi-layered architecture and environment of a generative artificial intelligence system for enterprises (e.g., generative artificial intelligence system 802 for enterprises) according to some embodiments. In the example of Fig. Section 7 describes the architecture and environment of the generative system for artificial intelligence in enterprises as a hierarchy of layers. More precisely, the hierarchy of layers includes an input layer (702), a monitoring layer (710), an agent layer (720), an agent and tool layer (730), a tool and data model layer (750), and an external layer (780). It is clear that these layers are only examples, and other examples may include any number of such layers (e.g., any number of layers 720 and 730).
[0065] The input layer 702 represents a layer of the generative system architecture for artificial intelligence in enterprises, which receives input (e.g., a query, a complex input, a set of instructions, and / or similar) from a user or a system. For example, an interface module of the enterprise's generative artificial intelligence system can receive the input.
[0066] The monitoring layer 710 represents a layer of the enterprise's generative artificial intelligence system architecture that contains one or more large language models (e.g., of an orchestrator module) capable of developing a plan to respond to input received in the input layer 702. A plan may include a set of prescribed tasks (e.g., fetch tasks, API calls, and the like). For example, monitoring layer 710 may provide the preprocessing and postprocessing functions described herein, as well as the functions of the orchestrators and understanding modules described herein. Monitoring layer 710 may coordinate with one or more of the subsequent layers 720–780 to execute the prescribed set of tasks.
[0067] Agent layer 720 represents a layer of the system architecture for generative artificial intelligence in enterprises, containing agents that can execute the prescribed set of tasks. In the example of Fig. Layer 7 comprises agent layer 720, a machine learning agent 722, an information gathering agent 724, a dashboard agent 726, and an optimization agent 728. Each of the agents 724–728 can contain an extensive language model that provides inferential functionality for performing its assigned portion of the prescribed task set. Specifically, agents 724–728 can instruct the agents and tools of subsequent layers (e.g., layer 730), of which there can be any number, to perform the tasks. For example, the machine learning agent 722 can instruct the word processing tool 732 to perform a word processing task (e.g., converting the output of an artificial intelligence application into natural language), an image processing tool 734 to perform an image processing task (e.g., creating a digital image).B, generating a natural language summary of an image output by an artificial intelligence application), a time series tool 736 to obtain a summary of time series data (e.g., time series data output by an artificial intelligence application), and an API tool 738 to perform an API call task (e.g., making an API call to trigger or access an artificial intelligence application).
[0068] The Information Retrieval Agent 724 can work with and / or coordinate several different agents to perform query tasks. For example, the Information Retrieval Agent 724 can instruct an Agent 740 for retrieving unstructured records, an Agent 742 for retrieving structured records, and an Agent 744 for retrieving type systems to retrieve one or more data models (or subsets of data models) and / or types from a type system. The type system ensures compatibility between different data formats, protocols, operating systems, different systems, etc. Types can encapsulate data formats for some or all of the different types or modalities described here (e.g., multimodal, text, coded, speech, statistics, audio, visual, audiovisual, etc.). Thus, a data model can contain a variety of different types (e.g.,in a tree or graph structure), and each type can describe data fields, operations, functions, and the like. Each type can represent a different object (e.g., a real-world object like a machine or sensor in a factory) or system (e.g., a computing cluster, enterprise data storage, file systems), and each type can contain a large language model context that provides the large language model with a context for creating or updating a plan. The context might include, for example, a natural language summary or description of the type (e.g., a description of the represented object, its relationships to other types or objects, its associated methods and functions, and so on). Types can be defined in a natural language format to enable efficient processing by large language models.The Type System Retriever Agent 744 can traverse Data Model 754 to retrieve a subset of Data Model 754 and / or types of Data Model 754. The Agent 742 for retrieving structured data can then use this retrieved information to efficiently retrieve structured data from a structured data source (e.g., a structured data source that is structured or modeled according to Data Model 754).
[0069] Dashboard agent 726 can be configured to generate one or more visualizations and / or graphical user interfaces, such as dashboards. For example, dashboard agent 726 can run tools 752-5 and 752-6 to generate dashboards based on information retrieved by the other agents and / or information output by the other agents (e.g., natural language summaries of related tool outputs).
[0070] The optimization agent 728 can be configured to perform a variety of different prescriptive analysis functions and mathematical optimizations 752-7 to assist in calculating answers to various problems. For example, the large language model 706 can use the optimization agent 728 to generate plans, determine a set of prescribed tasks, ascertain whether further information is needed to generate a final result, and the like.
[0071] The tool and data model layer 750 is intended to represent a layer of the company's generative artificial intelligence system architecture, comprising tools 752 and the data model 754. Agents 740-742 can execute tools 752 to retrieve information from various applications and data stores 782 in the external layer 780 (e.g., external in relation to the company's generative artificial intelligence system). Tools 752 may contain connectors that can establish a connection to systems and data stores located outside the company's generative artificial intelligence system.
[0072] Fig. Figure 8 shows a diagram 800 of an example of a network system for generative artificial intelligence in enterprises according to some embodiments. In the example of Fig. 8 The network system comprises a generative artificial intelligence system 802, enterprise systems 804-1 to 804-N (individually the enterprise system 804, collectively the enterprise systems 804), external systems 806-1 to 806-N (individually the external system 806, collectively the external systems 806) and a communications network 808.
[0073] The company's generative artificial intelligence system 802 can be used to iteratively and non-iteratively generate inputs and outputs of the machine learning model to determine a final output (e.g., an "answer" or "result") in response to an initial input (e.g., from a user or another system). In some embodiments, the functionality of the company's generative artificial intelligence system 802 can be executed by one or more servers (e.g., a cloud-based server) and / or other computing devices. The company's generative artificial intelligence system 802 can be implemented using a type system and / or a model-driven architecture. The company's generative artificial intelligence system 802 can also include the anti-hallucination and attribution module 110 (e.g.,added to the company's generative artificial intelligence system 802 and / or connected to the company's generative artificial intelligence system 802 after the company's generative artificial intelligence system 802 has been deployed).
[0074] In various implementations, the generative System 802 for artificial intelligence in enterprises can provide a wide range of technical features, such as the effective processing and generation of complex natural language inputs and outputs, the generation of synthetic data (e.g., supplementing customer data gathered during an onboarding process or otherwise filling data gaps), the generation of source code (e.g., application development), the generation of applications (e.g., artificial intelligence applications), the provision of cross-domain functionality, and a variety of other technical features not provided by traditional systems. Synthetic data can refer to content generated on-the-fly as part of the processes described here (e.g., through large language models). Synthetic data can also refer to unretrieved ephemeral content (e.g.,This includes temporary data that is not present in a database, as well as combinations of retrieved information, queried information, model outputs and / or similar.
[0075] In some embodiments, the generative 802 AI system for enterprise use can provide and / or enable an intuitive, non-complex interface to quickly execute complex user requests with improved access, privacy, and security enforcement. The generative 802 AI system for enterprise use can include a human-computer interface for receiving natural language requests and presenting relevant information with predictive analytics from the enterprise's information environment in response to those requests. For example, the generative 802 AI system can understand the language, intent, and / or context of a natural language user request.The company's generative artificial intelligence system 802 can execute the user's natural language query to identify relevant information from a corporate information environment and present it to the human computer interface (e.g., in the form of an "answer").
[0076] In some embodiments, generative artificial intelligence models (e.g., large language models of an orchestrator) of the company's generative artificial intelligence system 802 can interact with agents (e.g., retrieve agents) to retrieve and process information from various data sources. For example, data sources can store records and / or segments of records that can be identified by the company's generative artificial intelligence system 802 based on embedding values (e.g., vector values associated with records and / or segments). Records can contain tables, text, images, audio, video, code, application output (e.g., predictive analytics and / or other insights generated by artificial intelligence applications), and / or similar data.
[0077] In some embodiments, the company's 802 generative artificial intelligence system can generate context-based synthetic outputs based on retrieved information from one or more retriever models. For example, retriever models (such as a retrieval agent) can provide additional retrieved information to the large language models to generate additional context-based synthetic outputs until the context validation criteria are met. Once the validation criteria are met, the company's 802 generative artificial intelligence system can output the additional context-based synthetic output as a result or set of instructions (collectively referred to as "responses"). The context validation criteria may include a threshold for identifying source material from an enterprise data system that confirms the response.
[0078] In various configurations, the generative artificial intelligence system 802 delivers transformative, context-based, intelligent generative results for businesses. For example, the generative artificial intelligence system 802 can process input from business users via a natural language interface to quickly find, retrieve, and present relevant data across the entire corpus of an organization's information systems.
[0079] As previously described, the generative System 802 for enterprise artificial intelligence can process both machine-readable input (e.g., compiled code, structured data, and / or other types of formats that can be processed by a computer) and human-readable input. Input can also include complex input, such as input containing "and" or "or," input containing different types of information to satisfy the input (e.g., data records, text documents, database tables, and AI insights), and / or similar.For example, a complex input might be: “How many different engineers did John Doe work with within his engineering department?” This might require the company’s 802 generative artificial intelligence system to identify John Doe in a first iteration, identify John Doe’s department in a second iteration, determine the engineers in that department in a third iteration, then determine which of those engineers John Doe interacted with in a fourth iteration, and finally combine these results, or parts of them, to generate the final answer to the query. More precisely, the company’s 802 generative artificial intelligence system can use parts of the results from each iteration to generate contextual information (or simply context) that can then inform subsequent iterations.
[0080] Enterprise Systems 804 can include enterprise applications (e.g., artificial intelligence applications), enterprise data storage, client systems, and / or other systems within an enterprise information environment. As used herein, an enterprise information environment can comprise one or more networks (e.g., cloud, on-premises, air-gapped, or otherwise) of enterprise systems (e.g., enterprise applications, enterprise data storage) and client systems (e.g., computer systems used to access enterprise systems). Enterprise Systems 804 can include different computer systems, applications, and / or data storage, as well as enterprise-specific requirements and / or characteristics. For example, Enterprise Systems 804 can include access and privacy controls. For instance, an organization's private network can comprise an enterprise information environment that includes various Enterprise Systems 804.Enterprise 804 systems can include, for example, CRM systems, EAM systems, ERP systems, FP&A systems, HRM systems, and SCADA systems. Enterprise 804 systems can contain or utilize artificial intelligence applications, and the artificial intelligence applications can utilize enterprise systems and data. Enterprise 804 systems can encompass the flow of data and the management of various processes (e.g., of one or more organizations) and allow access to enterprise systems and users while preventing access by other systems and / or users. In some embodiments, references to enterprise information environments can also include enterprise systems, and references to enterprise systems can also include enterprise information environments. In various embodiments, the functionality of the enterprise 804 systems can be accessed from one or more servers (e.g.,a cloud-based server) and / or other computing devices.
[0081] External systems 806 can include applications, data storage, and systems that reside outside the enterprise's information environment. For example, enterprise systems 804 might be part of an organization's enterprise information environment that is inaccessible to users or systems outside that enterprise information environment and / or organization. Accordingly, the example of external systems 806 might include internet-based systems, such as news media systems, social media systems, and / or similar systems, that reside outside the enterprise's information environment. In various embodiments, the functionality of external systems 806 can be performed by one or more servers (e.g., a cloud-based server) and / or other computing devices.
[0082] The Communication Network 808 can be one or more computer networks (e.g., LAN, WAN, air-coupled network, cloud-based network, and / or the like) or other transmission media. In some embodiments, the Communication Network 808 can enable communication between the systems, modules, machines, generators, layers, agents, tools, orchestrators, data stores, and / or other components described herein. In some embodiments, the Communication Network 808 includes one or more computer devices, routers, cables, buses, and / or other network topologies (e.g., mesh, and the like). In some embodiments, the Communication Network 808 can be wired and / or wireless.In various configurations, the communication network 808 can include local area networks (LANs), wide area networks (WANs), the Internet and / or one or more networks, which can be public, private, IP-based, non-IP-based, air-gapped, and so on.
[0083] Fig. Figure 9 shows a diagram 900 of an example of a generative artificial intelligence system 802 for enterprises according to some embodiments. In the example of Fig. 9 comprises the generative system for artificial intelligence in the enterprise 802, an administration module 902, an orchestrator module 904, a retrieval agent module 906-1, a retrieval agent module 906-2 for unstructured data, a retrieval agent module 906-3 for structured data, a retrieval agent module 906-4 for writing systems, a module 906-5 for machine learning, an agent 906-6 for time series processing, an API agent module 906-7, a math agent module 906-8, a visualization agent module 906-9, a code generation agent module 906-10, a tool for retrieving unstructured data 908-1, a tool for retrieving structured data 908-2, a text processing tool module 908-3, and an image processing tool module 908-4, a time series processing tool module; 908-5, an API tool module; 908-6, a visualization tool module; 908-7, an optimization tool module; 908-8, a filter tool module; 908-9, a projection tool module; 908-10,a group tool module 908-11, an order tool module 908-12, a limit tool module 908-13, a code generation tool module 908-14, a chunking module 910, an embeddings generator module 912, a crawling module 914, a comprehension module 916, an enterprise access control module 918, an artificial intelligence traceability module 920, a parallelization module 922, a model generation module 924, a model deployment module 926, a model optimization module 928, an interface module 930, a communication module 932, an anti-hallucination and attribution module 934, vector data store 940, model registration data store 950, feature data store 960, and generative artificial Intelligence System Data Storage 970.
[0084] In some embodiments, the chunking module 910, the embedding generator module 912, the crawling module 914, the vector data store 940 (which stores, for example, embeddings) and parts of the generative artificial intelligence system data store 970 (for example, a segment data store) may include an intelligent crawling and chunking subsystem (for example, the intelligent crawling and chunking subsystem 120).
[0085] The Management Module 902 can manage (e.g., create, read, update, delete, or otherwise access) data associated with the enterprise's generative artificial intelligence system 802. The Management Module 902 can store, manage, or otherwise store data in any of the data stores 940-970 and / or in one or more other local and / or remote data stores. It is understood that the data stores can be a single local data store of the enterprise's generative artificial intelligence system 802 and / or multiple remote data stores of the enterprise's generative artificial intelligence system 802. In some embodiments, the data stores described herein include one or more local and / or remote data stores. The Management Module 902 can perform operations manually (e.g., by a user interacting with a graphical user interface) and / or automatically (e.g., by a script).triggered by one or more of modules 904–930). Like other modules described here, some or all of the functions of the 902 management module may be contained in and / or interact with one or more other modules, systems, and / or data stores.
[0086] The Orchestrator module 904 can be used to create and / or execute one or more Orchestrator agents (or simply Orchestrators). An Orchestrator can orchestrate, monitor, and / or otherwise control Agents 906. In some implementations, the Orchestrator includes one or more large language models. The Orchestrator can interpret inputs, select appropriate Agents to process queries and other inputs, and forward the interpreted input to the selected Agents. The Orchestrator can also perform a variety of monitoring functions. For example, the Orchestrator can implement halt conditions to prevent the Understanding module from getting stuck in an infinite loop during an iterative, context-based, generative artificial intelligence process. The Orchestrator can also include one or more other types of models for processing (e.g.,This includes transformations of non-text inputs. Other models (e.g., other machine learning models, translation models) may be used in addition to or instead of the major language models for some or all of the agents and / or modules described herein.
[0087] In some embodiments, an orchestrator can process data received from a variety of data sources in various formats, which can be processed using natural language processing (NLP) (e.g., tokenization, stemming, lemmatization, normalization, and the like) with vectorized data. It can generate pre-trained transformers that are fine-tuned or retrained on specific data tailored to an associated data domain or application (e.g., SaaS applications, legacy enterprise applications, artificial intelligence applications). Further processing may include checking data modeling features and / or simulating machine learning models to select one or more appropriate analysis channels.Examples of data objects include accounts, products, employees, suppliers, opportunities, contracts, locations, digital portals, geolocation providers, SCADA (Supervisory Control and Data Acquisition) information, OMS (Open Manufacturing System) information, inventory levels, supply chains, bills of materials, transportation services, maintenance logs, and service logs.
[0088] In some embodiments, the Orchestrator Module 904 can, as needed, use a variety of components to inventory or generate objects (e.g., components, functions, data, and / or the like) using extensive and descriptive metadata to dynamically create embeddings for knowledge development across a wide range of data domains (e.g., documents, tabular data, insights from artificial intelligence applications, web content, or other data sources). In a sample implementation, the Orchestrator Module 904 might, for example, utilize some or all of the components described here. Accordingly, the Orchestrator Module 904 can, for example, facilitate the storage, transformation, and communication of data to simplify its processing and embedding.In some implementations, the Orchestrator module can create embeddings for multiple data types across multiple industries and knowledge domains, and even specific enterprise knowledge. Knowledge can be explicitly modeled and / or learned by the Orchestrator module 904, the Agents 906, and / or the Tools 908. In one example, the Orchestrator module 904 (and / or the Chunking module 910) creates embeddings that are translated or transformed for compatibility with the Understanding module 916.
[0089] In some embodiments, the Orchestrator 904 can be configured to interact with or interface to various data domains within the components of the generative artificial intelligence system 802. For example, the Orchestrator module 904 can embed objects from specific data domains, as well as from various data domains, applications, data models, analytical byproducts, artificial intelligence predictions, and knowledge bases, to provide robust search functionality without requiring specific programming for each data domain or data source. For instance, the Orchestrator module 904 can create multiple embeddings for a single object (e.g., an object can be embedded in a domain-specific or application-specific context).In some embodiments, the Chunking Module 910, together with the Orchestrator Module 904, can curate the data domains for embedding objects of the data domains in the enterprise's information systems and / or environments. In some embodiments, the Orchestrator 904 can work together with the Chunking Module 910 to provide the embedding functionality described here.
[0090] In some embodiments, the Orchestrator Module 904 can cause an Agent 906 to perform data modeling to translate raw data formats into target embeddings (e.g., objects, types, and / or the like). Data formats can include some or all of the various types or modalities described herein (e.g., multimodal, text, coded, speech, statistics, audio, visual, audiovisual, etc.). In an example implementation, the Orchestrator Module 904 and / or the company's generative artificial intelligence system 802 generally use a type system of a model-driven architecture to perform the data modeling and translate raw data formats into target types. A knowledge base of the company's generative artificial intelligence system 802 and generative artificial intelligence models can create the ability to integrate or combine insights from different artificial intelligence applications.
[0091] As described elsewhere, the generative system 802 for enterprise artificial intelligence can process machine-readable input in addition to human-readable input (e.g., compiled code, structured data, and / or other types of formats that can be processed by a computer). Input can also include complex input, such as input with "and" or "or," or input containing different types of information to satisfy the input (e.g., text documents, database tables, and AI insights). The orchestrator 904 can split this complex input (e.g., by using a large language model) so that it can be processed by multiple agents 906 (e.g., in parallel).
[0092] As previously mentioned, the Orchestrator Module 904 can perform various monitoring functions and / or otherwise process data. In some implementations, the Orchestrator Module 904 can enforce conditions (e.g., halt conditions, resource allocation, prioritization, and / or the like). A halt condition, for example, might specify a maximum number of iterations (or jumps) that can be performed before the iterative process terminates. The halt condition and / or other features managed by the Orchestrator Module 904 can be included in large language model requests and / or in the large language models of the Orchestrator and / or Understanding Module 916, which are discussed later. In some embodiments, the halt conditions can ensure that the enterprise's generative artificial intelligence system 802 does not get stuck in an infinite loop.This feature can also give the company's generative artificial intelligence system 802 the flexibility to have a different number of iterations for different inputs (e.g., as opposed to a fixed number of jumps). In another example, the orchestrator module 904 can perform resource allocation, such as virtualization or load balancing, based on computational conditions. In some embodiments, the orchestrator module 904 and / or the agents 906 include models that can convert (or transform) an image, a database table, and / or other non-text input into a text format (e.g., natural language).
[0093] In some embodiments, the Orchestrator Module 904 can work in conjunction with Agents 906 (e.g., Retrieval Agent Module 906-1, Agent Module 906-2 for unstructured data, Agent Module 906-3 for structured data) to process inputs iteratively and non-iteratively to determine output results or responses, to determine the context and reasons for the information in subsequent iterations, and to determine whether large language models (e.g., of the Orchestrator 904 and / or the Understanding Module 916) require additional information to determine responses. For example, the Orchestrator Module 904 can receive a request and instruct Agent 906-1 to retrieve the associated information.Agent module 906-1 can then select agent module 906-2 for retrieving unstructured data and / or agent module 906-3 for retrieving structured data, depending on whether orchestrator module 904 wants to retrieve structured or unstructured data records. The corresponding agents 906 can select the appropriate tools and forward the tool outputs to orchestrator module 904 and / or understanding module 916 to determine a final result.
[0094] The Orchestrator 904 can also select and exchange models as needed. For example, the Orchestrator 904 can exchange the models (e.g., data models, large language models, machine learning models) of the enterprise's generative artificial intelligence system 802 at or during runtime, as well as before or after runtime. For example, the Orchestrator 904, the Agents 906, and the Understanding Engine 916 can use specific sets of machine learning models for one domain and different models for other domains. The Orchestrator 904 can select and use the appropriate models for a given domain and / or input.
[0095] In some embodiments, the Orchestrator 904 can combine (e.g., merge) outputs / results from different agents to create a unified output. For example, one or more of the Agent Modules 906 can receive / output a document (or segments thereof) or related information (e.g., a text summary or translation), another Agent Module 906 can receive / output a database table, and so on. The Orchestrator 904 can then use one or more machine learning models (e.g., a large language model and / or another machine learning model) to combine the outputs / results into a unified output (e.g., with a common data format, such as natural language).
[0096] In some implementations, the Orchestrator 904 pre-processes inputs (e.g., initial inputs) before sending them to one or more Agents 906 for processing. For example, the Orchestrator 904 might transform the first part of an input into an SQL query and send it to Agent Module 906-2 to retrieve unstructured data, transform the second part of the input into an API call and send it to API Agent Module 906-7, and so on. In another example, such transformation functionality might be performed by Agents 906 instead of, or in addition to, the Orchestrator 904.
[0097] The Orchestrator Module 904 can process, extract, and / or transform various data types (e.g., text, database tables, images, video, code, and / or the like). For example, the Orchestrator Module 904 can take a database table as input and convert it into natural language describing the database table. This natural language can then be provided to the Understanding Module 916, which can then process this transformed input to "answer" a query or otherwise fulfill a request. In some embodiments, one large language model can be used to process text, while another model can be used to convert (or transform) an image, a database table, and / or other non-text input into a text format (e.g., natural language).
[0098] In some embodiments, the Orchestrator Module 904 may include some or all of the functions of the Understanding Module 916. For example, the Understanding Module 916 may be a component of the Orchestrator Module 904. Similarly, in some embodiments, the Understanding Module 916 may include some or all of the functions of the Orchestrator Module 904.
[0099] In the example of Fig. 9. Agent modules 906 comprise a variety of different example agent modules 906-1 to 906-N. It is clear that these are only examples and that different embodiments may include other agents instead of, or in addition to, agents 906-1 to 906-N. In some embodiments, each of the agents 906 consists of hardware and / or software and includes one or more large language models, one or more other machine learning models, and / or functions to provide inferential functionality for handling a prescribed set of tasks. A reference to an agent module may refer to the agent itself and / or the component that creates and / or executes the agent. In some embodiments, the orchestrator is a type of agent and may be referred to as the orchestrator agent.Accordingly, the term "orchestrator" can refer to the orchestrator itself and / or to the component that creates and / or executes the orchestrator.
[0100] In various embodiments, some or all Agent 906 modules can process data of different data types and / or formats. For example, the Agent modules 906 can receive a database table or an image as input (e.g., from a Tool 908) and translate the table or image into natural language that describes it. This translation can then be output for processing by other modules, models, and / or systems (e.g., the Orchestrator module 904 and / or the Understanding module 916). In one example, a large language model can be used to process text, while another model can be used to convert (or transform) an image, a database table, and / or other non-text input into a text format (e.g., natural language).
[0101] The Retrieval Agent Module 906-1 can retrieve structured and unstructured data records. In some embodiments, the Search Agent Module 906-1 can coordinate / instruct the Search Agent Module 906-2 for unstructured data to retrieve unstructured data records, and coordinate / instruct the Search Agent Module 906-3 for structured data and the Search Agent Module 906-4 for type systems to retrieve structured data records. For example, the Query Agent Module 906-1 can work in conjunction with other Agents 906 and Tools 908 to construct SQL queries for querying an SQL database.
[0102] The Agent Module 906-2 for retrieving unstructured data records can retrieve unstructured data records (e.g., from an unstructured data store) and / or passages (or segments) of these data records. Unstructured data records can include, for example, text data stored in a file system in a format such as PDF, DOCX, MD, HTML, TXT, PPTX, or similar.
[0103] In some embodiments, the Agent 906-2 can use embeddings (e.g., vectors stored in vector memory 940) when retrieving information. For example, the Agent 906-2 can use a similarity score or a search in vector data memory 940 to find relevant records based on k-nearest neighbor, where embeddings that are closer together are more likely to be relevant.
[0104] In some embodiments, the 906-2 agent module implements a Read-Extract-Answer (REA) and / or a Read-Answer (RA) data retrieval process for unstructured data. REA and RA can be particularly useful when the System 802 needs to process large amounts of data. For example, a query might identify many different records and / or passages (e.g., hundreds or thousands of records and passages). For simplicity, the reference to records can encompass records and / or passages.
[0105] More specifically, the 906-2 Unstructured Data Retrieval Agent module can determine whether each record is relevant to answering the query and filter out irrelevant records. For example, the 906-2 Agent can calculate and assign relevance scores (e.g., using a machine learning relevance model) to each of the retrieved records. The relevance score can be relative to the other retrieved records. For instance, the least relevant record can be assigned a minimum value (e.g., 0), and the most relevant record a maximum value (e.g., 100). The 906-2 Unstructured Data Retrieval Agent module can also filter out relevant (or irrelevant) documents. For example, the 906-2 Unstructured Data Retrieval Agent module can filter out records whose relevance score is below a configurable threshold (e.g., 90).In some embodiments, the number of records that the 906-2 agent module can retrieve for unstructured data for a given input or query can be user- or system-defined and can also be configurable. For example, a system can specify that a maximum of 50 records can be returned.
[0106] In some embodiments, a large language model (e.g., of the Agent Module 906-2 for retrieving unstructured data) can identify key points of the relevant documents and passages and then pass these key points to another large language model (e.g., a large language model of the Orchestrator 904). The large language model can provide a summary that can be used to generate the query response (e.g., the summary can be the query response). This can allow the System 802, for example, to consider a large variety of concepts and documents (as opposed to an iterative process). If the number of documents or passages is below a certain threshold, the Agent Module 906-2 for retrieving unstructured data can skip the extract step (e.g., summarizing key points) and pass the passages directly to the large language model. This can be referred to as an RA process.
[0107] The Agent Module 906-3 for Retrieving Structured Records can retrieve structured records and / or passages (or segments) thereof from various structured data stores. Structured records can, for example, contain tabular data stored in a relational database, a key-value store, or an external database, and can be modeled or accessed using entity types (or simply types). Structured records can include records that are structured according to one or more data models (e.g., complex data models) and / or records that can be retrieved based on one or more data models. Structured records can include records stored in a structured data store (e.g., a data store structured according to one or more data models).
[0108] In certain implementations, data models may contain a graph structure of objects or types, and the agents 906 and / or tools 908 can traverse the graph along different paths to identify relevant types of the data model (e.g., depending on the query and a plan to answer the query provided by the orchestrator module 904), and can combine multiple tables with complex joins (e.g., as opposed to simply passing a single data source and performing operations on that single table). The paths can be stored in a data store (e.g., a vector data store 940) for efficient retrieval.
[0109] In some embodiments, the 906-3 Agent Module for Retrieving Structured Data can use a variety of different tools (e.g., the 908-2 Tool for Retrieving Structured Data, the 908-9 Filter Tool, the 908-10 Projection Tool, the 908-11 Grouping Tool, the 908-12 Job Tool, the 908-13 Boundary Tool, and the like). In some embodiments, once the 906-3 Agent Module for Retrieving Structured Data has traversed the data model and retrieved the relevant type(s) and / or subsets of the data model, it can use this information, along with the outputs of the agent and / or tool, to construct a structured query specification that it can execute against one or more structured data stores to retrieve the structured data records.
[0110] The type system's agent module 906-4 can be used to retrieve types, data models, and / or subsets of data models. For example, a data model can contain a variety of different types, and each type can describe data fields, operations, and functions. Each type can represent a different object (e.g., an object with real-world words, such as a machine or a sensor in a factor), and each type can contain a large language model context, providing context for a large language model. Types can be defined in a natural language format for efficient processing by large language models.
[0111] In some embodiments, a type system is designed to be used by various computer systems, application developers, data scientists, operations personnel, and / or other users to build applications, develop and run machine learning algorithms, and manage and monitor the state of jobs running on a type system (e.g., an enterprise generative artificial intelligence system in some embodiments). The type system is a framework that enables systems, application developers, data scientists, and other users to communicate effectively with each other using the same language. Accordingly, an application developer can interact with the enterprise's 802 generative artificial intelligence system in the same way as a data scientist.For example, you can use the same types, the same methods, and the same functions.
[0112] In some implementations, a type system can abstract the complex infrastructure within the enterprise's 802 generative artificial intelligence system. In one example, developers might never need to write SQL, CQL, or any other query processing language to access data. When a user reads data, the enterprise's 802 generative artificial intelligence system can generate the correct query for the underlying data store, submit the query to the database, and present the results to the user as a collection of objects or outcomes.
[0113] In some embodiments, a type can resemble a class in a programming language (e.g., a Java class) and describe data fields, operations, and functions (e.g., static functions) that can be called on the type or by one or more applications, but the type is not tied to a specific programming language. A type can be a definition of one or more complex objects that System 802 can understand. For example, a type can represent a variety of objects, such as a water pump. Types can be used to model systems as well as objects (e.g., computer clusters, key-value data stores, file systems, file repositories, enterprise data stores, and the like). In some embodiments, complex relationships, such as "when were which light bulbs in which light fixtures?", can be modeled as a type.
[0114] The 906-5 machine learning agent module can be used to receive and / or process outputs from artificial intelligence applications (e.g., insights from artificial intelligence applications). For example, the 906-5 machine learning module can instruct the 908-3 text processing tool to perform a text processing task (e.g., converting an artificial intelligence application into natural language), the 908-4 image processing tool to perform an image processing task (e.g., generating a natural language summary of an image output by an artificial intelligence application), the 908-3 time series tool to summarize time series data (e.g., time series data output by an artificial intelligence application), and the 908-6 API tool to perform an API call task (e.g.,Executing an API call to trigger or access an artificial intelligence application).
[0115] The Time Series Processing Agent 906-6 can be used to receive and / or process time series data, such as time series data output by various applications (e.g., artificial intelligence applications), machines, sensors, etc. The Time Series Processing Agent 906-6 can instruct and / or work in conjunction with the Time Series Processing Tool Module 908-3 to receive time series data from one or more artificial intelligence applications and / or other data sources.
[0116] The API Agent module 906-7 can coordinate and manage communication with other applications. For example, the API Agent module 906-7 can instruct the API Tool module 908-6 to make various API calls and then process the tool's output (e.g., convert it into a natural language summary).
[0117] The mathematical agent module 906-8 can be used to determine whether an agent 906 or a large language model needs additional information to produce a response or result. In some embodiments, the mathematical agent 906-8 can instruct the optimization tool module 908-8 to perform a variety of different prescriptive analysis functions and mathematical optimizations to assist in computing answers to various problems. For example, the orchestrator module 904 can use the mathematical agent module 906-8 to create plans, determine whether the orchestrator module 904 needs additional information to generate a final result, and similar tasks.
[0118] The Visualization Agent module 906-9 can be used to generate one or more visualizations and / or graphical user interfaces such as dashboards, charts, and the like. For example, the Visualization Agent module 906-9 can run the Visualization Tool module 908-7 to generate dashboards based on information retrieved by other agents and / or information output by other agents (e.g., natural language summaries of the associated tool outputs). The Visualization Agent module 906-9 can also be used to generate summaries (e.g., natural language summaries) of visual elements such as charts, tables, images, and the like.
[0119] The Code Generation Agent module 906-10 can be used to instruct the Code Generation Tool module 908-14 to generate source code, machine code, and / or other computer code. For example, the Code Generation Agent module 906-10 can be configured to determine what code is needed (e.g., to satisfy a query, create an application, etc.) and instruct the tool 908-14 to generate that code in a specific language or format.
[0120] In some embodiments, the Tools 908 are specific functions that agents (e.g., Agents 906, Orchestrator Module 904) can access or execute while attempting to perform prescribed tasks (e.g., from a set of prescribed tasks in a plan specified by Orchestrator Module 904). The Tools 908 may include software and / or hardware. The Tools 908 may also include one or more machine learning models, or they may include functions without a machine learning model. In some embodiments, the Tools 908 do not include large language models, although in other embodiments they may. In some embodiments, some or all of the Agents 906 and / or Tools 908 can be manually configured (e.g., by a user). The Agents 906 and the Tools 908 can also normalize data (e.g.,to a common data format) before outputting the data.
[0121] The Unstructured Data Retrieval Tool 908-1 can be used to retrieve unstructured data records from an unstructured data store. In some embodiments, the Agent 906-2 can use embeddings (e.g., vectors stored in the Vector Memory 940) when retrieving information. For example, the Agent 906-2 can use a similarity score or search to find relevant data records based on k-nearest neighbor, where embeddings that are closer together are more likely to be relevant. The Structured Data Query Tool 908-2 can query and retrieve structured data records from a structured data store (e.g., structured or modeled according to a data model). The Structured Data Retrieval Tool 908-2 can be executed by the Structured Data Retrieval Agent Module 906-3.
[0122] The 908-3 text processing module can be used to retrieve and / or transform text (e.g., from unstructured datasets) and perform other text processing tasks (e.g., converting text-based output from an artificial intelligence application into natural language). The 908-4 image processing module can perform an image processing task (e.g., creating a natural language summary of an image). The 908-5 time series processing tool module can be used to obtain and / or process time series data (e.g., output from artificial intelligence applications, sensors, and the like). For example, the 908-3 time series processing tool module can be executed by one or more of the 906 agents to obtain and process time series data. The 908-6 API tool module can be used to perform an API call task (e.g.,(to execute an API call to trigger or access an artificial intelligence application). For example, different 906 agents can use the 906-8 API tool module when the agent needs to access or trigger another application.
[0123] The Visualization Tool Module 908-7 can be used to generate one or more visualizations and / or graphical user interfaces, such as dashboards. For example, the Visualization Tool Module 908-7 can generate dashboards based on information retrieved by other agents and / or information output by other agents (e.g., natural language summaries of related tool outputs). The Filter Tool Module 908-9 can filter data records, types, and / or the like. For example, the Filter Tool Module 908-9 can filter projections (e.g., fields) identified by the Projection Tool Module 908-10 as part of a structured data retrieval process. In various embodiments, the 908 tools can be executed in parallel or in other ways.
[0124] In some embodiments, the Filter Tool Module 908-9 can identify implicit filters based on a query or other input, and these identified implicit filters can be used as part of a structured data retrieval process. For example, a query might be, "When was a premium towel last dispensed?" The Filter Tool Module 908-9 can identify "premium towel" as a filter (for example, based on an associated type description). The Filter Tool Module 908-9 can also identify context-aware date and time filters. For example, a query might be, "How many systems were offline yesterday?" The Filter Tool Module 908-9 can determine yesterday's date, taking into account the time zone and other relevant data to produce an accurate filter. In some embodiments, the Filter Tool Module 908-9 can validate identified filters before using them (for example, by checking if the system was offline).as part of a structured data retrieval process).
[0125] The Projection Tool Module 908-10 can be used to identify and select fields (e.g., type fields, object fields) that are relevant for determining a response to a query or other input. The Grouping Tool Module 908-11 can group data (e.g., types, tool outputs, etc.) that can then be used to generate structured queries (e.g., by the Agent Module 906-3 and / or the Tool Module 908-2 for retrieving structured data).
[0126] The Ordering Tool Module 908-12 can be used to organize data (e.g., types, tool outputs, and the like), which can then be used to generate structured query queries (e.g., by the Agent Module 906-3 for structured data queries and / or the Tool Module 906-2 for unstructured data queries). The Limiting Tool Module 908-13 can be used to limit the output of a structured data retrieval process. For example, it can limit the number of retrieved records, types, groups, filters, or similar parameters.
[0127] The Code Generation Tool Module 908-14 can generate source code, machine code, and / or other computer code. For example, the Code Generation Tool Module 908-14 can be configured to generate and / or execute SQL queries, Java code, and / or similar code. The Code Generation Tool Module 908-14 can be used to enable query generation for agents, other tools, large language models, and the like. The Code Generation Tool Module 908-15 can, in some embodiments, be configured to generate source code for an application or to build an application.
[0128] The embedding generator module 912 can be used to create embeddings based on structured and unstructured datasets and / or segments. The embedding generator module 912 can be the same as the embedding generator module 128. The embedding generator module 912 can contain one or more models (e.g., embedding models, deep learning models) that can convert and / or transform datasets into a vector representation where the vectors for semantically similar datasets (e.g., the content of the datasets) are close together in the vector space. This can facilitate retrieval operations by the agents 906 and tools 908.
[0129] In some cases, the embedding generator module 912 can generate embeddings using one or more embedding models (e.g., an implementation of the ColBERT embedding model). The embeddings can contain a numerical representation for unstructured and / or structured data records that captures the semantic or contextual meaning of the data records. For example, the embeddings can be represented by one or more vectors. The embeddings can be used when retrieving data records and when performing similarity assessments or other aspects of retrieval operations. The embeddings can be stored in an embedding index (e.g., in the vector data store 940). In some embodiments, the vector data store 940 is a database type specifically designed for storing embeddings and retrieving embeddings using a similarity heuristic (e.g.,an ANN (approximate nearest neighbor) algorithm is optimized, which can be implemented by the agents 906 and / or tools 908. In one example, the vector memory 940 can include an implementation of a FAISS vector memory.
[0130] The Chunking Module 910 can be used to process a corpus of data records (e.g., from one or more enterprise systems and / or external systems) for processing by different systems (e.g., the enterprise's generative artificial intelligence system 802). The Chunking Module 910 can partition data records and insert or append a corresponding header to each chunk. The header can contain one or more attributes that describe the chunk. Segments can contain the header along with a passage from a text document, a portion of a database table, a model, or a submodel, etc. For simplicity, a reference to a passage can encompass the segment and / or other content (e.g., text) within a segment. Segments can be stored in a segment data store. Chunking can be rule-based.
[0131] In some implementations, the chunking module 910 can preprocess the records and / or segments to generate relevant contextual information. In some embodiments, the contextual information can be contained and / or represented in contextual metadata, and / or the contextual metadata can be generated from the contextual information. The contextual information can improve the security, accuracy, and reliability of the associated retrieval operations. In one example, the contextual information includes contextual metadata. The contextual information can contain references between segments and / or records. These references can, for example, indicate relationships that can be used (e.g., traversed) when performing similarity assessments or other aspects of retrieval operations (e.g., by one or more of the agents 906).Contextual information can also include information that can assist a large language model in creating a plan and / or answers. For example, the chunking module 910 can generate contextual information for structured data chunks (or passages) that contain natural language descriptions of the data records, locations of related data records, and similar information.
[0132] Contextual information can contain access controls. In some implementations, the contextual information provides user-based access controls (e.g., role-based access controls) for associated records and / or segments. Specifically, the contextual information can specify user roles that are allowed to access a corresponding segment and / or record, and / or user roles that are not allowed to access a corresponding segment and / or record. The contextual information can be stored in the headers of the records and / or record segments. The contextual information can contain references between records and / or record segments. The chunking module 910 can generate contextual information before, after, or simultaneously with the generation of the associated embeds.For example, embeddings can be created using contextual information, or embeddings can be enriched with contextual information. The contextual information can be used by the Chunking Module 910 to map relationships between records and / or segments from one or more companies or enterprise systems and to store these relationships in a data model. In one example, the Chunking Module 910 implements a word2vec algorithm. In some implementations, the Chunking Module 910 uses models trained on domain-specific (or industry-specific) datasets.
[0133] In some embodiments, the Chunking Module 910 can execute some or all of the functions described herein periodically (e.g., in batches), on demand, and / or in real time. For example, the Chunking Module 910 can trigger the chunking described herein periodically, on demand, manually, and / or automatically. In some implementations, subsequent chunking operations can contain only changes compared to previous chunking operations (e.g., the "delta").
[0134] In some implementations, the embedding generator module 912 produces enriched embeddings. For example, the chunking module 910 can generate enriched embeddings based on context information, records, and / or record segments. An enriched embedding can include a vector value based on an embedding vector and the context information. In some embodiments, an enriched embedding includes the value of the embedding vector along with the contextual metadata, which contains the contextual information. Enriched embeddings can be indexed in an enriched embedding data store (for example, a vector data store 940). The agents 906 and / or tools 908 can retrieve unstructured and / or structured records based on enriched embeddings.In some embodiments, the contextual information may be contained and / or represented in contextual metadata and / or the contextual metadata may be generated from the contextual information.
[0135] The Crawling Module 914 can be used to scan and / or crawl various data sources (e.g., enterprise data sources, external data sources) across different domains. The Crawling Module 914 can identify existing records, new records, and / or records that have been updated. The Crawling Module 914 can inform the Chunking Module 910 about new records and records that have been updated, and the Chunking Module 910 can then chunk these records. In some implementations, the Crawling Module 914 can operate at regular intervals, on demand, and / or in real time, and / or trigger operation.
[0136] In some embodiments, the crawling module 914 can operate by scanning and / or crawling different data sources (e.g., enterprise systems 504, external systems 506) across various domains (e.g., data domains). It can identify existing records, new records, and / or updated records. The crawling module 914 can inform the chunking module 910 about new records and records that have been updated, and the chunking module 910 can chunk these records. In some embodiments, the crawling module 914 can operate at regular intervals, on demand, and / or in real time, and / or trigger operation.
[0137] In some implementations, the information sources may contain model registers that store different models (e.g., machine learning models, large language models, multimodal models). The models can be trained on general datasets and / or domain-specific datasets. The processes described here can be applied to various model registers. For example, models can be associated with an embedding value (e.g., generated by an embedding model) to facilitate model retrieval.
[0138] In some embodiments, the chunking module 910, the embedding generator module 912 and / or the crawling module 914 can perform the functions described in Fig. 2. Implement the described functionality (e.g., extracting information, identifying images and tables, chunking text, extracting information, generating an information diagram, and the like).
[0139] The company's generative artificial intelligence system 802 can perform some or all of the functions described herein periodically (e.g., in batches), on demand, and / or in real time. For example, the system can trigger the intelligent crawling and indexing described herein periodically, on demand, manually, or automatically. In some implementations, subsequent crawling and indexing operations may only consider changes compared to a previous crawling and indexing operation (e.g., the "delta").
[0140] The Understanding Module 916 can process inputs to determine results (e.g., "answers"), determine reasons for results, and ascertain whether the Understanding Module 916 needs more information to determine the results. The Understanding Module 916 can output information (e.g., results or additional queries) in a natural language or machine language format. In some implementations, features of one or more models of the Understanding Module define conditions or functions that determine whether more information is needed to satisfy the original input or whether enough information is available to satisfy the original input.
[0141] In some embodiments, the Understanding Module 916 includes one or more large language models. The large language models can be configured to generate and process context and other information described herein. The Understanding Module 916 may also include other language models that preprocess inputs (such as a user request) before they are provided to the agents for processing. The Understanding Module 916 may also include one or more large language models that process outputs from other models and modules (such as models of the Agent 906). The Understanding Module 916 may also include another large language model to transform responses from a large language model into a format more closely resembling a final response that can be delivered to various users and / or systems (such as the users or systems that made the original request, or other intended recipients of the response).
[0142] For example, the Understanding Module 916 can format responses according to different viewpoints. Viewpoints can be based on user type (e.g., human or machine), user roles (e.g., data scientist, engineer, director, etc.), access permissions, and so on. Accordingly, the Understanding Module 916 can generate and deliver a viewpoint-based response specifically tailored to the recipient. The Understanding Module 916 can also notify users and systems if it cannot find a suitable response (e.g., as opposed to presenting a response that is likely to be erroneous or biased).
[0143] In some implementations, features of one or more of the large language models of the Understanding Module 916 define conditions or functions that determine whether more information is needed to satisfy the initial input, or whether sufficient information is available to satisfy the initial input. The large language models of the Understanding Module 916 can also define termination conditions, which specify a termination threshold condition that indicates a maximum number of iterations that can be performed before the iterative process terminates.
[0144] In some embodiments, the Understanding Module 916 can generate and store reasons and contexts (e.g., in one or more data stores). The reason can be the argument that the Understanding Module 916 uses to determine an output (e.g., natural language output, a statement that it needs more information, a statement that it can satisfy the original input). The Understanding Module 916 can generate context based on the basic principle. In some implementations, the context includes a concatenation and / or annotation of one or more segments of data records and / or associated embeddings, along with a mapping of the concatenations and / or annotations. The mapping can, for example, specify relationships between different segments, a weighted or relative value associated with the different segments, and / or similar information.The rationale and / or the context can be incorporated into the commands that are provided to the large language models.
[0145] In some embodiments, the Understanding Module 916 includes a query and rationale generator that generates queries or other inputs for models (e.g., large language models, other machine learning models) and / or generates and stores the rationale and context (e.g., in the data store 960). The query and rationale generator can process, extract, and / or transform various data types (e.g., text, database tables, images, video, code, and / or the like). For example, the query and rationale generator can take a database table as input and convert it into natural language describing the database table, which can then be provided to one or more other models (e.g., large language models) of the Understanding Module 916, which can then process this transformed input to "answer" a query or otherwise fulfill it.In some implementations, the query and rationale generator includes models that can convert (or transform) an image, a database table, and / or other non-text input into a text format (e.g., natural language). Although queries are used in various examples in this document, other types of input (e.g., command sets) can also be processed in the same or a similar way as described for queries.
[0146] In some embodiments, the Understanding Module 916 can use different models for different domains. For example, different domains might correspond to different industries (e.g., aerospace, defense), different technological environments (e.g., on-premises, in the air, in the cloud), different companies or organizations, and / or the like. Accordingly, the Understanding Module 916 can use certain models (e.g., data models and / or large language models) for one particular domain (e.g., a data model that describes properties and relationships of aerospace objects and a large language model trained on aerospace-specific datasets) and a different data model and / or large language model for another domain (e.g.,a data model that describes properties and relationships of defense-specific objects, and a large language model that has been trained on defense-specific datasets), and so on.
[0147] In some embodiments, the orchestrator module 904 includes some or all of the functions and / or structures of the understanding module 916 and / or 906 described below. Similarly, in some embodiments, the understanding module 916 may include some or all of the functions and / or structures of the orchestrator module 904.
[0148] In some embodiments, the Understanding Module 916 can be used to generate large language model prompts (or simply prompts) and prompt templates. For example, the Understanding Module 916 can generate a command template for processing an initial input, a command template for processing iterative inputs (i.e., inputs received during the iteration process after the initial input has been processed), and another command template for the output result phase (i.e., when the Understanding Module 916 has determined that it has sufficient information and / or a stop condition has been met). The Understanding Module 916 can modify the appropriate command template depending on a phase of the iterative process. For example, command templates can be modified to generate commands that include justifications and contexts that can be used for subsequent iterations.
[0149] The Enterprise Access Control Module 918 can be used to provide enterprise-wide access controls (e.g., layers and / or protocols) for the Generative Artificial Intelligence System 802, its associated systems (e.g., enterprise systems), and / or environments (e.g., enterprise information environments). The Enterprise Access Control Module 918 can provide functions for enforcing access control policies regarding the generation of results (e.g., to prevent the Orchestrator Module 904 and / or the Understanding Module 916 from generating results containing sensitive information) and / or filtering results that have already been generated before a final output is delivered.
[0150] In some implementations, the enterprise's Access Control Module 918 can evaluate (for example, using access control lists) whether a user is authorized to access all or only part of a result (such as a response). For instance, a user might submit a query associated with an initial department or subunit of an organization. Members of that department or subunit might be restricted from accessing certain data, data types, data models, or other aspects of a data domain in which a search is to be performed. If the initial results include data to which the user has restricted access, the enterprise's Access Control Module 918 can determine how to handle such restricted data, for example, by...The options are: to omit the restricted data entirely, to omit the restricted data but indicate that the results contain data to which the user has restricted access, or to provide information about all initial results. In the example where the restricted data is completely omitted, a final set of results can be returned for presentation to the user, with the final set of results not informing the user that some of the original results were omitted. In the example where the restricted data is omitted but the user is given a hint about the presence of the restricted data, the final results can include only the results to which the user has access, but include information indicating that there were X number of initial results, but only Y results are returned, where Y < X.In the third example described above, all results can be displayed to the user, including those results for which the user has restricted access.
[0151] Additionally or alternatively, the enterprise access control module 918 can communicate with one or more other modules to obtain information that can be used to enforce access permissions / restrictions in connection with the execution of retrieval operations, rather than controlling the presentation of the results to the user. For example, the enterprise access control module 918 can restrict the data sources to which retrieval operations are applied, such as preventing a retrieval operation from being applied to parts of the data sources for which user access is denied and applying retrieval operations to parts of the data sources for which user access is permitted.It should be noted that the exemplary access restriction enforcement techniques described above are provided for illustration purposes only and are not intended as a limitation, and it should be understood that modules operating in accordance with embodiments of this disclosure may implement other techniques to present results via an interface based on access restrictions.
[0152] To facilitate the enforcement of access restrictions in conjunction with searches performed by the enterprise's generative artificial intelligence system 802, the enterprise's access control module 918, in some embodiments, can store information associated with access restrictions or permissions for each user. To retrieve the relevant restriction data for a user, the enterprise access control module 918 can obtain user-identifying information in conjunction with the user's input or login to the system on which the enterprise access control module 918 is running. The enterprise access control module 918 can use the user-identifying information to retrieve appropriate restriction data to support the enforcement of access restrictions in conjunction with an enterprise search.In some embodiments, the enterprise access control module 918 may include the credential management functionality of a model-driven architecture employing the enterprise's generative artificial intelligence system 802, or it may be a remote credential management system that communicates with the enterprise's generative artificial intelligence system 802 via a network.
[0153] The AI Traceability Module 920 can be used to ensure the traceability and / or explainability of responses generated by the company's AI Generative System 802. For example, the AI Traceability Module 920 can specify portions of datasets used to generate the responses and their respective data sources. The AI Traceability Module 920 can also be used to validate the outputs of large language models. For example, the AI Traceability Module 920 can automatically and / or on demand provide source citations to confirm or validate the outputs of large language models. The AI Traceability Module 920 can also ensure the compatibility of different sources (e.g.,Identify data records (e.g., passages) used to generate a large language model output. For example, AI Traceability Module 920 can identify conflicting data records (e.g., one record indicates that John Doe is an employee of Acme, and another indicates that John Doe works for a different company) and provide a notification that the output was generated based on conflicting information. AI Traceability Module 920 can work with and / or incorporate the functionality of Hallucination Protection and Assignment Module 934.
[0154] The parallelization module 922 can be used to control the parallelization of the various systems, modules, agents, models, and processes described here. For example, the parallelization module 922 can trigger parallel execution of different agents and / or orchestrators. The parallelization module 922 can be controlled by the orchestrator module 904.
[0155] The Model Generation Module 924 can be used to obtain, generate, and / or modify some or all of the different model types described herein (e.g., machine learning models, large language models, data models). In some implementations, the Model Generation Module 924 can use a variety of machine learning techniques or algorithms to generate models. As used here, artificial intelligence and / or machine learning can employ Bayesian algorithms and / or models, deep learning algorithms and / or models (e.g.,Artificial neural networks (including convolutional neural networks), gap analysis algorithms and / or models, supervised learning techniques and / or models, unsupervised learning algorithms and / or models, semi-supervised learning techniques and / or models, random forest algorithms and / or models, similarity learning and / or distance learning algorithms, generative artificial intelligence algorithms and models, clustering algorithms and / or models, transformer-based algorithms and / or models, transformer-based machine learning algorithms and / or models for neural networks, reinforcement learning algorithms and / or models, and / or similar. The algorithms can be used to generate the corresponding models. For example, the algorithms can be run on datasets (e.g., domain-specific datasets, enterprise datasets) to generate and / or output the corresponding models.
[0156] In some implementations, a large language model is a deep learning model (e.g., generated by a deep learning algorithm) that can recognize, summarize, translate, predict, and / or generate text and other content based on knowledge gained from large datasets. Large language models can include Transformer-based models. Examples of large language models include Google's BERT, OpenAI's GPT-3, and Microsoft's Transformer. Large language models can process large amounts of data, resulting in improved accuracy in prediction and classification tasks. These models can use this information to learn patterns and relationships, helping them make better predictions and groupings compared to other machine learning models.Large language models can include artificial neural network transformers pre-trained using supervised and / or semi-supervised learning techniques. In some embodiments, large language models include deep learning models specialized in text generation. Large language models may also be characterized, in some embodiments, by a substantial number of parameters (e.g., in the range of tens or hundreds of billions) and by large text corpora used to train these models.
[0157] Although the systems and processes described here use large language models, other types of machine learning models may be used in other embodiments instead of, or in addition to, large language models. For example, an Orchestrator 904 may use deep learning models specifically designed to receive non-natural language input (e.g., images, video, audio) and provide natural language output (e.g., summaries) and / or other types of output (e.g., a video summary).
[0158] The Model Deployment Module 926 can be used to deploy some or all of the different model types described here. In some implementations, the Model Deployment Module 926 can deploy models before or after the deployment of an enterprise generative artificial intelligence system. For example, the Model Deployment Module 926 can work in conjunction with the Model Optimization Module 928 to exchange or otherwise modify large language models of an enterprise generative artificial intelligence system.
[0159] In some implementations, a model registry can store 950 different models (e.g., machine learning models, large language models, data models) and / or model configurations. The models can be trained on general-purpose and / or domain-specific datasets. For example, the model registry can store different configurations of various large language models (which can be used or exchanged, for example, in an enterprise generative artificial intelligence system). In some embodiments, each of the models can be associated with an embed value or an enriched embed value to facilitate retrieval operations (e.g., in the same or a similar way as querying datasets).
[0160] The Model Optimizer Module 928 can function by enabling tuning and learning by the modules described here (e.g., the Understanding Module 916) and / or the models (e.g., machine learning models, large language models). For example, the Model Optimizer Module 928 can tune the Understanding Module 916 and / or the Orchestrator Module 904 (and / or their models) based on tracking user interactions within systems, capturing explicit feedback (e.g., via a training user interface), implicit feedback, and / or the like. In some example implementations, the Model Optimizer Module 928 can use reinforcement learning to accelerate knowledge base bootstrapping. Reinforcement learning can be used for explicit bootstrapping of various systems (e.g.,of the company's generative artificial intelligence system 802) with instrumentation of time spent, results clicked, and / or similar metrics. Example aspects of the model optimization module 928 include an innovative learning framework that can boot models for various enterprise environments. Example aspects of the model optimization module 928 may include an innovative learning framework that can boot models for various enterprise environments.
[0161] In some implementations, reinforcement learning is a machine learning training method based on rewarding desired behaviors and / or punishing undesired behaviors. Generally, a reinforcement-learning agent is able to perceive and interpret its environment, take action, and learn through trial and error. Reinforcement learning uses algorithms and models to determine the optimal behavior in an environment to obtain maximum reward. This optimal behavior is learned through interactions with the environment and by observing its responses. In the absence of a supervisor, the learner must independently discover the sequence of actions that maximize the reward. This discovery process resembles a trial-and-error approach.The quality of actions is measured not only by the immediate reward but also by the delayed reward they might bring. Because the algorithm is capable of learning the actions that lead to success in an unseen environment without the help of a supervisor, it is very efficient. ColBERT is an example of a retriever model that enables scalable BERT-based searching of large text collections (e.g., within ten milliseconds). ColBERT uses a late interaction architecture that independently BERT-encodes a query and a document and then employs a "cheap" but powerful interaction step that models their fine-grained similarity.In addition to reducing the costs of reclassifying documents retrieved using a conventional model, the ColBERT-friendly interaction mechanism enables the use of vector similarity indices for end-to-end querying directly from a large document collection.
[0162] In some embodiments, the Model Optimization Module 928 can retrain models (e.g., transformer-based natural language machine learning models) periodically, on demand, and / or in real time. In some example implementations, appropriate candidate models (e.g., candidate models for transformer-based natural language machine learning) can be trained based on user selections, and the Model Optimization Module 928 can replace some or all of the models with one or more candidate models trained on the received user selections.
[0163] In some embodiments, the Model Optimizer Module 928 can exchange the models of the enterprise's generative artificial intelligence system not only before or after runtime, but also during runtime. For example, the Orchestrator Module 904, Module 916, and / or the Agents 906 may use specific sets of machine learning models for one domain and different models for other domains. The Model Exchange Module can select and use the appropriate models for a given domain. This can even occur during the iteration process. For instance, if new queries are generated by the Understanding Module 916, the domain may change, prompting the Model Exchange Module to select and use a different model appropriate for that domain.
[0164] In some embodiments, the model optimization module 928 can train generative artificial intelligence models to develop different types of responses (e.g., best results, ranked results, smart cards, chatbot, new content generation and / or similar).
[0165] In some embodiments, the model optimization module can retrain 928 models (e.g., large language models) periodically, on demand, and / or in real time. In some example implementations, appropriate candidate models can be trained based on user selections, and the system can replace some or all models with one or more candidate models trained on the received user selections.
[0166] The Interface Module 930 can receive input (e.g., complex input) from users and / or systems. The Interface Module 930 can also generate and / or transmit output. Input can include system input and user input. Examples of input can include command sets, queries, natural language input, or other human-readable input, machine-readable input, and / or the like. Similarly, output can also include system output and human-readable output. In some embodiments, an input (e.g., a request, query) can be entered and processed in various natural forms for simple human interaction (e.g., a simple text-based interface, image processing, speech activation, and / or the like) to quickly retrieve relevant and engaging information.
[0167] In some embodiments, the interface module 930 can be used to generate graphical user interface components (e.g., server-side graphical user interface components) that can be rendered as complete graphical user interfaces on the company's generative artificial intelligence system 802 and / or other systems. For example, the interface module 930 can provide an interactive graphical user interface for displaying and receiving information.
[0168] The Communications Module 932 can be used to send requests, transmit and receive communications, and / or otherwise enable communication with one or more of the systems, modules, machines, layers, devices, data storage devices, and / or other components described herein. In a particular implementation, the Communications Module 932 can be used to encrypt and decrypt communications. The Communications Module 932 can be used to send requests to and receive data from one or more systems over a network or part of a network (e.g., the Communications Network 808). In a particular implementation, the Communications Module 932 can send requests and receive data over a connection that may be wholly or partially wireless.The communication module 932 can request and receive messages and / or other communications from associated systems, modules, layers, and / or the like. These messages can be stored in the data storage unit 970 of the company's generative artificial intelligence system.
[0169] In some embodiments, the configuration, coordination, and cooperation of the Orchestrator Module 904, the Agents 906, the Tools 908, and / or other modules of the Enterprise Generative Artificial Intelligence System 802 (e.g., Understanding Module 916) enables the Enterprise Generative Artificial Intelligence System 802 to provide a multi-hop architecture that allows complex inferences across multiple Agents 906, Tools 908, and data sources (e.g., vector data stores, feature data stores, data models, enterprise data stores, unstructured data sources, structured data sources, and the like). In various embodiments, some or all modules of the Generative Artificial Intelligence System 802 can be configured manually (e.g., by a user) and / or automatically (e.g., without user input).For example, large language model commands can be configured, Tool 908 descriptions can be configured (e.g., for more efficient use by Agents 906 and Orchestrator Module 904), and a maximum number of jumps or iterations can be configured. In one example, Orchestrator Module 904 receives a request from a user and determines a plan to respond to the request, selecting Agents 906 and / or Tools 908 to perform a prescribed set of tasks that execute the plan. Agents 906 and / or Tools 908 perform the prescribed set of tasks, and Orchestrator Module 904 observes the result. Orchestrator Module 904 determines whether to submit a final response or whether it needs more information.If the Orchestrator Module 904 has sufficient information, it can generate and / or provide the final response. Otherwise, the Orchestrator Module 904 can create another prescribed set of tasks, and the process can continue until the Orchestrator Module 904 has sufficient information to respond or until a halt condition is met (e.g., a maximum number of jumps).
[0170] The Anti-Hallucination and Allocation Module 934 may be the same as the Anti-Hallucination and Allocation Module 110. The Anti-Hallucination and Allocation Module 934 may be an extension of the Comprehension Module 916 and / or be included in the Comprehension Module 916.
[0171] Fig. Figure 10A shows a flowchart 1000 of an example of an iterative generative artificial intelligence process that uses unstructured data according to some embodiments. This example process can be implemented by an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802). In this and other flowcharts and / or sequence diagrams, the flowchart illustrates an exemplary sequence of steps. It is understood that some or all of the steps may be repeated, reorganized for parallel execution, and / or rearranged as needed. Furthermore, some steps that could have been included may have been omitted to avoid providing too much information, and some steps that were included could be omitted but were included for clarity.
[0172] In step 1002, a user query is made to a search model (e.g., a search module of a search agent module). In step 1004, the query model receives the query and performs a similar search (e.g., an ANN-based search) in vector memory 1006. The retrieved information is returned to the query model and passed to a large language model in step 1010. In step 1012, a large language model (e.g., the large language model used in step 1010 and / or another large language model) determines whether additional information is needed to answer the user query. If more information is needed, steps 1004–1012 can be iterated with updated commands from the large language model and / or queries (e.g.,Using a large language model from step 1008, the process is repeated until the large language model has enough information to answer the query or a termination condition is met (e.g., when a maximum number of iterations have been performed). In step 1014, the answer (e.g., the final result if the large language model has enough information to determine an answer, or "I don't know" if the termination condition was met before it could obtain enough information) is generated and / or presented. The final result may also include the reasoning the large language model used to generate the answer. The large language models used in steps 1008 and 1010 may be the same large language model and / or a different large language model.
[0173] Fig. Figure 10B shows a flowchart (1030) of an exemplary non-iterative generative artificial intelligence process using unstructured data according to some embodiments. This example process can be implemented by an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802). In this and other flowcharts and / or sequence diagrams, the flowchart illustrates an exemplary sequence of steps. It is understood that some or all of the steps may be repeated, reorganized for parallel execution, and / or rearranged as needed. Furthermore, some steps that could have been included may have been omitted to avoid providing too much information, and some steps that were included could have been omitted but were included for clarity.
[0174] In step 1032, a query is received. The query 1032 is executed in a vector memory 1034, and relevant passages 1036 are retrieved. In some embodiments, the user query 1032 is preprocessed (e.g., by an orchestrator) before being applied to the vector memory 1034 to retrieve passages 1036. For example, the query 1032 can be translated, transformed, and the like. Since complex inputs can be difficult for vector memories to process, the system can generate a new query or several shorter queries from the user query 1032, which the vector memory 1034 can process efficiently and accurately. The query 1032 and the passages 1036 are provided to a large language model 1038, which can create an extract 1040 (e.g., a summary of the passage) for each passage. The extracts are combined in step 1042 (e.g.,The extracted passages are concatenated and, together with the query 1032, transmitted to a large language model 1044. In some embodiments, the extraction steps are optional, and the passages can be concatenated and fed to the large language model 1044 instead of being extracted. The large language model 1044 can generate a final response based on the query 1032 and the combined extracts 1042. In some embodiments, the large language model 1044 can post-process the result (e.g., by the orchestrator) before it is provided to the user. For example, it can be translated, formatted according to viewpoints, annotated with citations and attributions, etc.
[0175] Fig. Figure 10C shows a flowchart (1060) of an exemplary non-iterative generative artificial intelligence process using unstructured data according to some embodiments. This example process can be implemented by an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802). In this and other flowcharts and / or sequence diagrams, the flowchart illustrates an exemplary sequence of steps. It is understood that some or all of the steps can be repeated, reorganized for parallel execution, and / or rearranged as needed.Furthermore, some steps that could have been included may have been removed to avoid providing too much information for the sake of clarity, and some steps that were included could be removed, but may have been included for the sake of clarity.
[0176] In step 1062, a query is received. The query might be, for example, "How much wine do they produce?". This query would be difficult for conventional large language models to process and would typically lead to large language model hallucinations because it is unclear how to handle the "they" in query 1062. To solve this problem, a generative AI system used in enterprises can use context 1064 to generate an improved query. For example, a previous conversation 1064 (e.g., as part of a chat with a chatbot) might have included a conversation about France. The system can provide France as context information 1064 to generate a new query 1066, such as "How much wine is produced in France?", which can prevent the large language model from hallucinating and allow it to deliver an accurate and reliable final result 1078.
[0177] More specifically, the company's generative artificial intelligence system can generate a rewritten query 1066 (e.g., using the large language model 1065) that can be executed against the vector memory 1068 to retrieve passages 1070. In some embodiments, the rewritten query 1066 can be preprocessed (e.g., by an orchestrator) before being applied to the vector memory 1068. For example, the rewritten query 1066 can be translated, transformed, and so on. Since complex inputs are difficult for vector memory to process, the system can generate a new query or several shorter queries from the rewritten query 1066 that the vector memory 1068 can process efficiently and accurately. In some embodiments, this preprocessing can be performed when the rewritten query is generated (e.g., rewriting the query includes the preprocessing steps).
[0178] The rewritten query 1066 and the passages 1070 are fed to a large language model 1069, which can generate an extract 1072 (e.g., a summary of the passage) for each passage. The extracts 1072 are combined (e.g., concatenated) in step 1074 and, along with the rewritten query 1066, submitted to a large language model 1076. The company's generative artificial intelligence system can use the combined extracts 1074 to generate a rationale 1082 for determining the final answer 1078. For example, the large language model 1076 can generate the final answer 1078 based on the rationale 1082 and / or present the rationale 1082 (or a summary of the rationale) along with the final answer 1078 (e.g., for citation or attribution purposes).
[0179] In some embodiments, the extraction steps are optional, and the passages can be concatenated and provided to the large language model 1076 instead of extracts or combined extracts. The large language model 1044 can generate a final response 1078 based on the rewritten query 1066 and the combined extracts 1074. In some embodiments, the large language model 1076 can post-process the result (e.g., by the orchestrator) before the final result 1078 is provided to the user. For example, it can be translated, formatted according to viewpoints, annotated with citations and attributions, etc.
[0180] Fig. Figure 11 shows a flowchart (1100) of an example of an iterative generative artificial intelligence process using unstructured data according to some embodiments. This example process can be implemented by an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802). In this and other flowcharts and / or sequence diagrams, the flowchart illustrates an exemplary sequence of steps. It is understood that some or all of the steps can be repeated, reorganized for parallel execution, and / or rearranged as needed. Furthermore, some steps that could have been included may have been omitted to avoid providing too much information for the sake of clarity, and some steps that could have been included may have been omitted but were included for clarity.
[0181] In the example of Fig. Section 11 comprises a generative system for artificial intelligence in enterprises (e.g., the generative system for artificial intelligence in enterprises 802), one or more query modules 1104, and one or more understanding modules 1106. The query module 1104 can, for example, contain one or more large language models, and the understanding module can contain one or more other large language models. These modules and the iterative interactions (e.g., communication) between them can enable generative artificial intelligence systems in enterprises to achieve the technical features and advantages discussed here.
[0182] In the example of Fig. 11. The enterprise generative artificial intelligence system can receive an initial input 1102 from a user or another system. For example, an orchestrator module 1103 can receive the input 1102. The enterprise's generative artificial intelligence system can then forward this input to a retrieval module 1104 (e.g., corresponding to one or more agents 906 and / or tools 908), which can then "retrieve" information from the embedding memory 1108. For example, the retrieval module 1104 can retrieve passages based on the input by using one or more similarity heuristics (e.g., ANN algorithms) executed on the embedding memory 1108 (e.g., one or more vector stores) to retrieve passages (or data sets) relevant to the input.
[0183] The company's generative artificial intelligence system can use this retrieved information to generate an initial command for the Understanding Module 1106. The Understanding Module 1106 can process this initial command and determine if it contains enough information to meet the criteria based on the initial input (e.g., answering a question). See, for example, step 1107. If it has enough information to fulfill the original input, the Understanding Module can pass the result to a recipient (see, for example, step 1113), such as the user or the system that provided the original input. However, if the Understanding Module 1106 determines that it does not have enough information to meet the criteria based on the initial input, it can synthesize more information through the iterative process, which is the system's primary benefit.
[0184] There can be many reasons why the understanding module 1106 requires additional information. For example, conventional systems use only a single pass that addresses only a portion of a complex input. The generative AI system for enterprises tackles this problem by triggering subsequent iterations to resolve the remaining parts of the complex input and also incorporate the context to further refine the process.
[0185] More specifically, if the Understanding Module 1106 determines that it needs additional information to satisfy the original input, it can generate context-specific data (or simply "context") that informs future iterations of the process and helps the system satisfy the original input more efficiently and accurately. The context is based on the logic that the Understanding Module 1106 applies when processing queries (or other inputs). For example, the Understanding Module 1106 can receive segments of information retrieved by the Retrieval Module 1104. These segments might be passages of data records, and they might be associated with embeddings from an Embedded Data Store 1108, which facilitates processing by the Understanding Module 1106.A query and justification generator 1112 of the communication module 1106 can process the information and provide a justification for why the result turned out the way it did. This justification can be stored by the company's generative artificial intelligence system in a historical rational data store 1110 and form the basis for subsequent iterations of the context.
[0186] More specifically, subsequent iterations may involve the Understanding module 1106 generating a new query, request, or other output, which is then returned to the Retrieval module. The Query module 1104 can process this new request and retrieve additional information. The system then generates a new command based on the additional information and context. The Understanding module 1106 can process the new command and again determine if it needs additional information. If it does, the company's generative artificial intelligence system can repeat (e.g., iterate) this process until the Understanding module 1106 can meet the criteria based on the initial input, at which point the Understanding module 1106 can generate the output result 1114 (e.g., "Answer" or "I don't know").For example, the comprehension module 1106 can generate the response "I don't know" if no relevant passages have been generated or retrieved (e.g., by applying a rule of the comprehension module 1106) and / or not enough relevant passages have been generated, retrieved, and / or extracted to prevent hallucinations and increase performance on the "I don't know" queries, while saving a call to the models (e.g., large language models).
[0187] In some embodiments, the number of retrieved passages from which no relevant information was extracted (e.g., by the Understanding Module 1106) can be used to determine and / or correlate whether sufficient information is available. For example, a threshold number or percentage of retrieved passages from which relevant information was extracted may need to be met (e.g., a specific number or percentage of retrieved passages) for the Enterprise Understanding Module 1106 to determine that it has enough information to answer the query. In another example, a specific number or percentage of retrieved passages from which no relevant information was extracted (e.g.,4 passages or 80% of the retrieved passages), cause the business understanding module 1106 to determine that it does not have enough information to answer the request.
[0188] The company's generative artificial intelligence system can also implement monitoring functions, such as a stop condition that prevents the system from hallucinating or otherwise giving an incorrect answer. The termination condition can also prevent the system from executing an infinite iteration loop. For example, the company's generative artificial intelligence system can limit the number of iterations that can be performed before the understanding module 1106 either provides an output result or indicates that an output result cannot be found. The user can also provide feedback 1116, which can be stored in a feedback data store 1118. In some embodiments, the company's generative artificial intelligence system can use the feedback to improve the system's accuracy and / or reliability.As explained elsewhere, the functionality of the understanding modules can be integrated into the orchestrator in some embodiments.
[0189] Fig. Figure 12 shows a flowchart (1200) of an example of a generative artificial intelligence process in enterprises, according to some embodiments. In this and other flowcharts and / or sequence diagrams, the flowchart illustrates an exemplary sequence of steps. It is understood that some or all of the steps can be repeated, reorganized for parallel execution, and / or rearranged as needed. Furthermore, some steps that could have been included may have been omitted to avoid providing too much information, and some steps that were included could be omitted but were included for clarity.
[0190] In step 1202, an enterprise system (e.g., enterprise system 804) displays a graphical user interface (GUI). In some embodiments, an interface module (e.g., interface module 930) of an enterprise generative artificial intelligence system (e.g., enterprise generative artificial intelligence system 802) can facilitate the display of the GUI. For example, the interface module can generate server-side parts of the GUI (e.g., render), and the enterprise system can generate client-side parts of the GUI.
[0191] In step 1204, the company's generative artificial intelligence system receives a question via the graphical user interface. This question could, for example, be a natural language query. In some implementations, an orchestrator module (e.g., orchestrator module 904) receives the query.
[0192] In step 1206, the enterprise's generative artificial intelligence system retrieves information related to the question from various enterprise systems (e.g., enterprise systems 804). In some embodiments, an understanding module (e.g., understanding module 916) retrieves the information. An example of an information retrieval process is shown in Fig. 15 is shown. The understanding module 916 can, for example, work with agents and / or tools to retrieve the information.
[0193] In step 1208, the enterprise's generative artificial intelligence system generates an answer to the question using a generative artificial intelligence model (or multiple models) and information retrieved from the various enterprise systems. In some embodiments, the understanding module generates the answer.
[0194] In step 1210, the enterprise system displays the answer to the question via the graphical user interface. In some implementations, the interface module facilitates the display.
[0195] Fig. Figure 13 shows a flowchart (1300) of an example of an enterprise generative artificial intelligence process that uses an agent-orchestrator architecture according to some embodiments. In step 1302, an enterprise system (e.g., enterprise system 804) displays a graphical user interface (GUI). In some embodiments, an interface module (e.g., interface module 930) of an enterprise generative artificial intelligence system (e.g., generative artificial intelligence system 802) can facilitate the display of the GUI. For example, the interface module can generate server-side parts of the GUI (e.g., render), and the enterprise system can generate client-side parts of the GUI.
[0196] In step 1304, the company's generative artificial intelligence system receives a question via the graphical user interface. This question could, for example, be a natural language query. In some implementations, an orchestrator module (e.g., orchestrator module 904) receives the query.
[0197] In step 1306, the company's generative artificial intelligence system manages various agent programs (e.g., agent modules 906–910) through an orchestrator program (e.g., orchestrator module 904) to generate answers to questions. The orchestrator can use at least one multimodal model to transform the question into a series of instructions for the different agent programs, enabling them to retrieve the information associated with the question. The various agent programs can use one or more machine learning models to retrieve the information related to the question based on the series of instructions.
[0198] In step 1308, the company's generative artificial intelligence system retrieves the information associated with the question through its various agent programs. In some embodiments, agent modules 906-1 to 906-4 retrieve the information (e.g., from a vector memory, company systems, and / or external systems).
[0199] In step 1310, the company's generative artificial intelligence system uses a generative artificial intelligence model and the associated information to generate an answer to the question. In some embodiments, an understanding module (e.g., understanding module 916) generates the answer. The generative artificial intelligence model and the multimodal model may be the same model. The generative artificial intelligence model and the multimodal model may be different models.
[0200] In step 1312, the enterprise system displays the answer to the question via the graphical user interface. In some implementations, the interface module facilitates the display.
[0201] In various implementations, the management of the different agent programs can involve iterative processing or multiple instructions from the orchestrator. Retrieval can include the retrieval of time-series data, structured data, and unstructured data. The agent programs can instantiate a tool to perform an operation on the instruction, the retrieved data, or intermediate data. The agent programs can perform operations such as computation, translation, formatting, and / or visualization. The one or more agent programs can be trained with different domain-specific machine learning models. The agent programs can use a type system to unify incompatible data from different data sources.
[0202] Fig. Figure 14 shows a flowchart (1400) of an exemplary anti-hallucination and attribution procedure for generative systems for enterprise artificial intelligence. In step 1402, an enterprise generative system (e.g., the enterprise generative system 802) receives output from a generative AI model that processes a command. In some embodiments, an anti-hallucination and attribution module (e.g., the anti-hallucination and attribution module 110) receives the output. The generative AI model may use a generative adversarial network, a variance autoencoder, an autoregressive model, and / or a recurrent neural network.
[0203] In step 1404, the anti-hallucination and attribution module (e.g., of the enterprise's generative artificial intelligence system) decomposes the output of a generative artificial intelligence model into chunks that can be attributed to one or more source texts. The anti-hallucination and attribution module can be an extension of the enterprise's generative artificial intelligence system and / or a component of the enterprise's generative artificial intelligence system. In some embodiments, a response parser (e.g., response parser 112) decomposes the output into chunks.
[0204] In step 1406, the anti-hallucination and attribution module retrieves the one or more source texts based on a similarity assessment between the chunks and the one or more source texts. In some embodiments, a retriever (e.g., the anti-hallucination and attribution retriever 114) retrieves the source texts. The retriever may include one or more machine learning models, generative artificial intelligence models, and the like. The retriever 114 may cooperate with and / or be included as part of the agents (e.g., agents 906-1, 906-2, 906-3, 906-4, etc.) and / or tools (e.g., tool 908-1, 908-2, 908-3, etc.) described herein. The similarity assessment can be performed by the Retriever 114 and / or include a similarity assessment that calculates similarity values associated with the chunks and the source code.
[0205] In some embodiments, a generative artificial intelligence system (e.g., the generative artificial intelligence system 802) generates an information diagram (e.g., the information diagram 220) for each record in a set of records containing one or more source codes. The information diagram can describe relationships between source codes and one or more other classes, where the other classes include source images, source tables, and source code. Retrieval of the one or more source codes can be performed based on the information diagram (e.g., by iterating through the in Fig. 2 of the diagram shown (220).
[0206] In step 1408, the anti-hallucination and matching module assigns at least some of the one or more source texts to the chunks based on a similarity threshold. The anti-hallucination and matching module can compare the result of the similarity assessment (e.g., the score) between the chunks and the one or more source texts to the similarity threshold. For example, if the score is at or above the similarity threshold, the corresponding source text(s) can be assigned to the chunk. Conversely, if a score is below the threshold, the corresponding source text(s) will not be assigned to the chunk.
[0207] In step 1410, the anti-hallucination and matching module combines the chunks with the source code associated with those chunks. This may involve merging source citations of the source code and / or some or all of the source code with the chunks.
[0208] In step 1412, the anti-hallucination and attribution module generates a response to the command based on the combination. The response can include output with inline source identifiers that identify the attributed source code. The response can include output containing at least a portion of the attributed source code. The response can be output (e.g., displayed) on one or more systems (e.g., one or more enterprise systems).
[0209] In some embodiments, the output includes sentences, and the chunks may contain one or more of these sentences.
[0210] Fig. Figure 15 shows a flowchart 1500 of an exemplary information retrieval procedure for a generative enterprise artificial intelligence (AI) process according to some embodiments. Like the other methods described here, this flowchart 1500 can be combined with other flowcharts (e.g., flowchart 1200). In step 1502, a generative enterprise AI system (e.g., the generative enterprise AI system 802) identifies enterprise data records, AI applications, and data models from various data domains of one or more enterprise systems based on a received query. The enterprise data records may include documents, document segments, and insights generated by one or more AI applications.
[0211] In step 1504, the enterprise's generative artificial intelligence system determines relevance scores based on the data models associated with the enterprise datasets. In some embodiments, an understanding module (e.g., understanding module 916) determines the relevance scores. Each relevance score can be associated with a corresponding portion of the enterprise datasets, and each relevance score is determined relative to the other portions of the enterprise datasets.
[0212] In step 1506, the enterprise's generative artificial intelligence system determines information related to the question through one or more generative artificial intelligence models, based on relevance assessments and one or more of the enterprise's access control protocols. In some embodiments, the understanding module determines the information (e.g., based on data retrieved by one or more agents and / or tools).
[0213] Fig. Figure 16 shows a flowchart 1600 of an example procedure for routing requests to different agent programs and validating responses for generative artificial intelligence enterprises according to some embodiments. Like the other methods described here, this flowchart 1600 can also be combined with other flowcharts (e.g., flowchart 1300). In step 1602, an orchestrator (e.g., the enterprise's orchestrator 904 of the generative artificial intelligence system 802) directs retrieval requests to different agent programs.
[0214] In step 1604, the company's generative artificial intelligence system receives information from various agent programs across multiple data domains, based on instructions from the orchestrator. In some embodiments, the orchestrator module and / or the understanding module (e.g., understanding module 916) receives the information.
[0215] In step 1606, the orchestrator analyzes the information to formulate one or more answers to the question, with the orchestrator making additional retrieval requests to at least one of the different agent programs to retrieve additional information to satisfy a context validation criterion associated with the question.
[0216] In step 1608, the orchestrator outputs a validated answer from one or more of the answers to the question that meets the context validation criteria.
[0217] Fig. Figure 17 shows a diagram 1700 of an example computer system for implementing the features disclosed herein according to some embodiments. Each of the systems, machines, data storage devices, and / or networks described herein may include an instance of one or more computing devices 1702. In some embodiments, the functionality of the computing device 1702 is enhanced to perform some or all of the functions described herein. The computing device 1702 comprises a processor 1704, a memory 1706, a storage device 1708, an input device 1710, a communication network interface 1712, and an output device 1714, which is communicatively connected to a communication channel 1716. The processor 1704 is configured to execute instructions (e.g., programs).In some embodiments, the 1704 processor comprises a circuit or any processor capable of processing the executable instructions.
[0218] Memory 1706 stores data. Some examples of memory 1706 include storage devices such as RAM, ROM, RAM cache, virtual memory, etc. In various configurations, the working data is stored in memory 1706. The data in memory 1706 can be erased or eventually transferred to memory 1708.
[0219] Memory 1708 encompasses any storage device configured to retrieve and store data. Some examples of Memory 1708 include flash drives, hard disk drives, optical drives, cloud storage, and / or magnetic tapes. Both Memory System 1706 and Memory System 1708 comprise a computer-readable medium that stores instructions or programs that can be executed by the Processor 1704.
[0220] The input device 1710 is any device that inputs data (e.g., a mouse and keyboard). The output device 1714 outputs data (e.g., via a speaker or a display). The memory 1708, the input device 1710, and the output device 1714 are, of course, optional. For example, the routers / switches may include the processor 1704 and the memory 1706, as well as a device for receiving and outputting data (e.g., the communication network interface 1712 and / or the output device 1714).
[0221] The 1712 communication network interface can be connected to a network (e.g., the 808 network) via the 1718 connection. The 1712 communication network interface can support communication via Ethernet, serial, parallel, and / or ATA connections. It can also support wireless communication (e.g., 802.11, WiMAX, LTE, Wi-Fi). It is clear that the 1712 communication network interface can support many wired and wireless standards.
[0222] It becomes clear that the hardware elements of the 1702 data processing device do not correspond to those in Fig.The data processing device 1702 is limited to the components shown in Figure 17. It may include more or fewer hardware, software, and / or firmware components than those shown (e.g., drivers, operating systems, touchscreens, biometric analyzers, and / or the like). Furthermore, hardware elements may have the same functionality and be included in various embodiments described herein. For example, encoding and / or decoding may be performed by the processor 1704 and / or a coprocessor on a graphics processor (e.g., Nvidia).
[0223] Examples of computing and / or processing devices include one or more microprocessors, microcontrollers, reduced instruction set computers (RISCs), complex instruction set computers (CISCs), graphics processing units (GPUs), data processing units (DPUs), virtual processing units, associative processing units (APUs), tensor processing units (TPUs), vision processing units (VPUs), neuromorphic chips, AI chips, quantum processing units (QPUs), Cerebras wafer-scale engines (WSEs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or discrete circuits.
[0224] An "engine," "system," "data store," and / or "database" can consist of software, hardware, firmware, and / or circuits. In one example, one or more software programs containing instructions that can be executed by a processor can perform one or more of the functions of the machines, data stores, databases, or systems described herein. In another example, circuits can perform the same or similar functions. Alternative embodiments can include more, less, or functionally equivalent machines, systems, data stores, or databases and still fall within the scope of the present embodiments. For example, the functionality of the various systems, machines, data stores, and / or databases can be combined or divided differently. The data store or database can include cloud storage.The term "or" used here can be understood in both an inclusive and an exclusive sense. Furthermore, multiple instances of resources, processes, or structures can be provided, which are described here as a single instance.
[0225] The data stores described herein may have any suitable structure (e.g., an active database, a relational database, a self-referential database, a table, a matrix, an array, a flat file, a document-oriented storage system, a non-relational NoSQL system, and the like) and may be cloud-based or otherwise. The systems, methods, machines, data stores, and / or databases described herein may be at least partially processor-implemented, with a particular processor or processors being an example of hardware. For example, at least some of the operations of a procedure may be performed by one or more processors or processor-implemented machines.Furthermore, one or more processors can also be operated in such a way as to support the execution of the relevant operations in a cloud computing environment or as Software as a Service (SaaS). For example, at least some of the operations can be performed by a group of computers (as examples of machines with processors), with these operations being accessible via a network (e.g., the internet) and through one or more suitable interfaces (e.g., an application programming interface (API)).
[0226] The execution of certain operations can be distributed across processors that are not located on a single computer, but rather across a number of computers. In some embodiments, the processors or processor-implemented machines may be located in a single geographical location (e.g., within a home environment, an office environment, or a server farm). In other embodiments, the processors or processor-implemented machines may be distributed across a number of geographical locations.
[0227] In this description, multiple instances can implement components, operations, or structures that are described as a single instance. Although individual operations of one or more methods are represented and described as separate operations, one or more of the individual operations can be executed concurrently, and nothing requires that the operations be executed in the order shown. Structures and functions represented as separate components in the example configurations can be implemented as a combined structure or component. Likewise, structures and functions represented as a single component can be implemented as separate components. These and other variations, modifications, additions, and enhancements fall within the scope of this topic.
[0228] In a sample implementation, the enterprise generative AI systems described here can connect to one or more virtual metadata repositories across data stores, abstract access to disparate data sources, and support granular data access controls maintained by the enterprise AI system. The enterprise generative AI framework can manage a virtual data lake with an enterprise catalog connected to multiple data domains and industry-specific domains. The enterprise generative AI framework's orchestrator is capable of creating embeddings for multiple data types across multiple industries and knowledge domains, and even specific enterprise knowledge.Embedding objects in enterprise information system data domains enables rapid identification and complex processing with relevance assessment, as well as additional functions for enforcing access, data protection, and security protocols. In some implementations, the orchestrator module can employ a variety of embedding methods and techniques that are understood by a professional. In one example implementation, the orchestrator module can use a model-driven architecture for the conceptual representation of enterprise and external datasets and optional data virtualization. A model-driven architecture might be, for example, as described in U.S. Patent 10817530, serial number 15 / 028340, granted on October 27, 2020, with priority dated January 23, 2015, entitled "Systems, Methods, and Devices for an Enterprise Internet-of-Things Application Development Platform" by C3 AI, Inc.A type system of a model-driven architecture can be used to embed objects of the data domains.
[0229] The model-driven architecture ensures the compatibility of system objects (e.g., components, functions, data, etc.) that the orchestrator can use to dynamically generate queries for performing searches across a wide range of data domains (e.g., documents, spreadsheets, insights from AI applications, web content, or other data sources). The type system provides data accessibility, compatibility, and usability across different systems and data. In particular, the type system addresses data usability across various programming languages, inconsistent data structures, and incompatible software application programming interfaces. The type system offers data abstraction that defines extensible type models, allowing new properties, relationships, and functions to be added dynamically without requiring costly development cycles.The type system can be used as a domain-specific language (DSL) within a platform, which is used by developers, applications, or user interfaces to access data. The type system provides the ability to interact with data to perform processing, predictions, or analyses based on one or more type or function definitions within the type system. The orchestrator is a mechanism for implementing search functionality across a wide variety of data domains, in contrast to existing query modules, which are typically limited to their searchable data domains (e.g., web query modules are limited to web content, file system query modules are limited to searching the file system, etc.).
[0230] Type definitions can be a canonical type declared in metadata, using a syntax similar to that used for types persisted in relational or NoSQL data stores. A canonical model in the type system is application-independent, allowing all applications to communicate with each other in a common format. Unlike a standard type, canonical types consist of two parts: the canonical type definition and one or more transformation types. The canonical type definition specifies the interface used for integration, and the transformation type is responsible for converting the canonical type into a corresponding type. Using the transformation types, the integration layer can convert a canonical type into the corresponding type.
[0231] Various embodiments of the present disclosure include systems (e.g., with one or more processors and memory instructions which, when executed by the one or more processors, cause the system to perform the functionality described herein), methods, and non-transitory computer-readable media (or media) configured to perform the following: displaying a graphical user interface; receiving a question through the graphical user interface; retrieving question-related information from various enterprise systems; generating an answer to the question by a generative artificial intelligence model using the information retrieved from the various enterprise systems; and displaying the answer to the question through the graphical user interface.The systems, procedures, and non-transitory computer-readable media (or media) can further be configured to identify, based on the question, enterprise data records, artificial intelligence applications, and data models from various data domains of the enterprise systems; determine relevance scores associated with the enterprise data records based on the data models; and, through one or more generative artificial intelligence models, determine the information related to the question based on the relevance scores and the enterprise's access control protocols. The question comprises a natural language query. The enterprise data records may include documents, document segments, and insights generated by one or more artificial intelligence applications.Each of the relevance ratings can be associated with a corresponding part of the company data records, and each of the relevance ratings is determined relative to the other parts of the company data records.
[0232] Various embodiments of the present disclosure include systems (e.g., various embodiments of the present disclosure include systems (e.g.with one or more processors and memory instructions which, when executed by the one or more processors, cause the system to perform the functionality described herein), methods and non-transitory computer-readable media (or media) configured to display a graphical user interface, receive a question through the graphical user interface, manage various agent programs by an orchestrator program to generate answers to questions; retrieve information related to the question by the various agent programs; generate an answer to the question by a generative artificial intelligence model using the related information; and display the answer to the question through the graphical user interface.The systems, procedures, and non-transitory computer-readable media (or media) may further be configured to perform the following: instructing the orchestrator to make retrieval requests to the various agent programs; receiving information from multiple data domains from the various agent programs based on the orchestrator's instructions; analyzing the information to formulate one or more answers to the question, with the orchestrator making additional retrieval requests to at least one of the various agent programs to retrieve additional information to satisfy a context validation criterion associated with the question; and outputting a validated answer of the one or more answers to the question that satisfies the context validation criteria.
[0233] The orchestrator can use at least one multimodal model to transform the question into a series of instructions for the various agent programs to retrieve the information associated with the question. The various agent programs can use one or more machine learning models to retrieve the information related to the question based on the series of instructions. The generative artificial intelligence model and the multimodal model can be the same model. The generative artificial intelligence model and the multimodal model can be different models.In some embodiments, the systems, methods and non-transitory computer-readable media are further configured to generate a traceability analysis of the natural language output, wherein the traceability analysis specifies one of the documents, the document segments and the findings of the respective parts of the one or more enterprise data sets.
[0234] Various embodiments of the present disclosure include systems (e.g., with one or more processors and memory instructions which, when executed by the one or more processors, cause the system to perform the functionality described herein), methods, and non-transitory computer-readable media (or media) configured to: receive the output from a generative artificial intelligence model processing source text; parse the output from a generative model into chunks to be mapped to one or more source passages; retrieve the one or more source texts based on a similarity assessment between the chunks and the one or more source texts; map at least a portion of the one or more source texts to the chunks based on a similarity threshold;Combining the chunks with the source code associated with those chunks; and generating the response to the command based on the combination, the output containing inline source identifiers that identify the associated source code. The systems, procedures, and non-transitory computer-readable media (or media) may further be configured to perform filtering of the one or more source codes based on similarity values; and generating an information graph for each record in a set of records containing the one or more source codes, the information graph describing relationships between source codes and one or more other classes, the other classes being source images, source tables, and source code.
[0235] The generative artificial intelligence model can use a generative adversarial network, a variance autoencoder, an autoregressive model, or a recurrent neural network. The output can contain sentences, and the chunks can contain one or more of these sentences. The similarity assessment can include a similarity score (e.g., a cosine similarity score) that calculates similarity values for the chunks and the source code. The response can be generated by a generative artificial intelligence model. Retrieval of the one or more source codes can be based on the information graph. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 10817530
[0228] US 15 / 028.340
[0228]
Claims
[1] A procedure, comprising: Receiving an output from a generative artificial intelligence model that processes an input prompt; Parsing the output of the generative model into sections that are to be assigned to one or more source sections; Retrieving one or more source sections based on a similarity assessment between the sections and the one or more source sections; Assigning at least one part of the one or more source sections to the sections based on a similarity threshold; Combining the sections with the source sections that are assigned to those sections; and Generating a response to the input prompt based on the combination, where the response includes the output with inline source identifiers that identify the associated source sections. [2] The method according to claim 1, wherein the generative artificial intelligence model uses a generative adversarial network, a variational autoencoder, an autoregressive model or a recurrent neural network. [3] The method according to claim 1, wherein the output comprises sentences and the sections contain one or more of the sentences. [4] The method according to claim 1, wherein the similarity assessment comprises a similarity assessment that calculates similarity values that are assigned to the sections and the source sections. [5] The method according to claim 1, wherein the response is generated by a generative artificial intelligence model. [6] The method according to claim 4, further comprising filtering the one or more source sections based on the similarity values. [7] The method according to claim 6, further comprising: Generating an information graph for each record in a set of records comprising one or more source sections, wherein the information graph describes relationships between source sections and one or more other classes, the other classes comprising source images, source tables, and source code. [8] The method according to claim 7, wherein the retrieval of one or more source sections is based on the information graph. [9] A system, comprehensive: one or more processors; and Memory that stores instructions which, when executed by one or more processors, cause the system to: Receiving an output from a generative artificial intelligence model that processes an input prompt; Parsing the output of the generative model into sections that are to be assigned to one or more source sections; Retrieving one or more source sections based on a similarity assessment between the sections and the one or more source sections; Assigning at least one part of the one or more source sections to the sections based on a similarity threshold; Combining the sections with the source sections that are assigned to those sections; and Generating a response to the input prompt based on the combination, where the response includes the output with inline source identifiers that identify the associated source sections. [10] The system according to claim 9, wherein the generative artificial intelligence model uses a generative adversarial network, a variational autoencoder, an autoregressive model or a recurrent neural network. [11] The system according to claim 9, wherein the output comprises sentences and the sections contain one or more of the sentences. [12] The system according to claim 9, wherein the similarity assessment comprises a similarity assessment that calculates similarity values that are assigned to the sections and the source sections. [13] The system according to claim 9, wherein the response is generated by a generative artificial intelligence model. [14] The system according to claim 12, wherein the instructions, when executed by the one or more processors, cause the system to filter the one or more source sections based on the similarity values. [15] The system according to claim 14, wherein the instructions, when executed by the one or more processors, cause the system to generate an information graph for each record in a set of records comprising the one or more source sections, wherein the information graph describes relationships between source sections and one or more other classes, the other classes comprising source images, source tables and source code.
Citation Information
Patent Citations
Systems, methods, and devices for an enterprise internet-of-things application development platform
US10817530B2
15/028.340
US-PATENT10817530