Automating the efficient deployment of artificial intelligence models
The unified metadata graph system addresses inefficiencies in accessing siloed data by using natural language processing and large language models to optimize data retrieval and AI deployment, reducing resource waste and downtime.
Patent Information
- Application Number
- JP2024225656
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-09-17
- Filing Date
- 2024-12-20
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-12-20
AI Technical Summary
Existing systems face inefficiencies and resource wastage due to data silos across computing systems, requiring significant computational effort to maintain data integrity and access, often leading to system downtime and memory waste.
A unified metadata graph system that uses natural language processing and large language models to access and generate domain-specific metadata graphs, reducing the need for new data silos and reconfiguring existing systems, thereby optimizing data retrieval and AI model deployment.
This approach minimizes computational resource utilization and data retrieval time while maintaining data integrity, enabling efficient access to siloed data across disparate locations and automating AI model deployment.
Smart Images

Figure 0007783397000001 
Figure 0007783397000002 
Figure 0007783397000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Patent Application No. 18 / 888,151, entitled "AUTOMATING EFFICIENT DEPLOYMENT OF ARTIFICIAL INTELLIGENCE MODELS," filed September 17, 2024, which is a continuation-in-part of U.S. Patent Application No. 18 / 390,916, entitled "ACCESSING SILHOED DATA ACROSS DISPARATE LOCATIONS VIA A UNIFIED METADATA GRAPH SYSTEMS AND METHODS," filed December 20, 2023, and U.S. Patent Application No. 18 / 627,332, entitled "GENERATING A UNIFIED METADATA GRAPH VIA A RETRIEVAL-AUGMENTED GENERATION (RAG) FRAMEWORK SYSTEMS AND METHODS," filed April 4, 2024. This application also claims priority to U.S. patent application Ser. No. 18 / 617,305, entitled "ACCESSING SILHOED DATA ACROSS DISPARATE LOCATIONS VIA A UNIFIED METADATA GRAPH SYSTEMS AND METHODS," filed March 26, 2024, and to U.S. patent application Ser. No. 18 / 888,088, entitled "ARTIFICIAL INTELLIGENCE SANDBOX FOR AUTOMATING DEVELOPMENT OF AI MODELS," filed September 17, 2024. The contents of the foregoing applications are incorporated herein by reference in their entireties. [Background technology]
[0002] As computing systems become increasingly complex, the data used by such computing systems often requires its own data silo (database, data warehouse, or data lake) to process the data efficiently. Yet, because each computing system has its own data silo, copies of data may exist between different silos for each computing system. As a result, significant computational capacity is required to read and maintain a single version of the truth among copies of data residing in different data silos distributed across one or more computer systems. Moreover, data in one data silo may be similar to data in other data silos. For example, each computing system may require its own unique variable names, sequencing keys, and integrity constraints, so variable names and technical implementations may differ from data silo to data silo, but the underlying data is the same. New applications are built using the latest technologies and techniques, which quickly become outdated with the emergence of newer and better performing systems. These new and emerging silos compromise time and complexity in business processes, and building a coordinated and consolidated silo that can replace all existing silos requires a significant effort to build and validate. Left unchecked, this can result in similar data residing in multiple data silos that could otherwise be used to store new information, further contributing to large amounts of wasted computing resources. [Brief explanation of the drawings]
[0003] [Figure 1] FIG. 1 is an illustrative representation of a Graphical User Interface (GUI) for reducing computational resource usage when accessing siloed data across disparate locations via a unified metadata graph, in accordance with some implementations of the present technology. [Figure 2]FIG. 1 is a block diagram illustrating some of the components typically incorporated in at least some of the computer systems and other devices on which the disclosed systems operate, in accordance with some implementations of the present technology. [Figure 3] FIG. 1 is a system diagram illustrating an example computing environment in which the disclosed system operates in some implementations of the present technology. [Figure 4] 1 is a flow diagram illustrating a process for reducing computational resource usage when accessing siloed data across disparate locations through a unified metadata graph, according to some implementations of the present technology. [Figure 5] 1 is a diagram of an illustrative representation of a metadata graph, in accordance with some implementations of the present technology. [Figure 6] FIG. 1 is a close-up view of a metadata graph, in accordance with some implementations of the present technology. [Figure 7] FIG. 1 is a diagram of an artificial intelligence model, in accordance with some implementations of the present technology. [Figure 8] 1 is a flow diagram illustrating a process for generating a unified metadata graph via a retrieval-augmented generation (RAG) framework, according to some implementations of the present technology. [Figure 9A] 1 is an illustrative diagram of a Large Language Model (LLM) prompt, according to some implementations of the present technology. [Figure 9B] 1 is an illustrative diagram of a Large Language Model (LLM) prompt, according to some implementations of the present technology. [Figure 10A] FIG. 1 is a subsystem diagram illustrating an example RAG framework environment for generating a unified metadata graph, in accordance with some implementations of the present technology. [Figure 10B] FIG. 1 is a subsystem diagram illustrating an example RAG framework environment for generating a unified metadata graph, in accordance with some implementations of the present technology. [Figure 10C]FIG. 1 is a subsystem diagram illustrating an example RAG framework environment for generating a unified metadata graph, in accordance with some implementations of the present technology. [Figure 10D] FIG. 1 is a subsystem diagram illustrating an example RAG framework environment for generating a unified metadata graph, in accordance with some implementations of the present technology. [Figure 11] 1 is an illustrative representation of a generated metadata graph according to some implementations of the present technology. [Figure 12] FIG. 12 is a block diagram illustrating components within an AI sandbox 1200, according to some implementations. [Figure 13] FIG. 1 is a block diagram illustrating components within a data processor according to some implementations. [Figure 14] FIG. 1 is a block diagram illustrating components within a model generator, according to some implementations. [Figure 15] FIG. 1 is a block diagram illustrating the functionality of a model governor, according to some implementations. [Figure 16A] 1 is a flowchart illustrating a process for automatically generating an artificial intelligence (AI) model, according to some implementations. [Figure 16B] 1 is a flowchart illustrating a process for automatically generating a data pipeline for an AI model, according to some implementations. [Figure 17A] FIG. 1 is a diagram of an example chat interface. [Figure 17B] FIG. 1 is a diagram of an example chat interface. [Figure 17C] FIG. 1 is a diagram of an example chat interface. [Figure 17D] FIG. 1 is a diagram of an example chat interface. [Figure 18] FIG. 1 is a schematic diagram illustrating the operation of a model automaton, according to some implementations. [Figure 19] 1 is a flowchart illustrating a process for automating deployment of AI models, according to some implementations. DETAILED DESCRIPTION OF THE INVENTION
[0004] In the drawings, for purposes of discussion of some implementations of the present technology, some components and / or operations may be separated into different blocks or combined into a single block. Moreover, while the present technology is susceptible to various modifications and alternative forms, specific implementations are shown by way of example in the drawings and are described in detail below. Nevertheless, the intention is not to limit the present technology to the specific implementations described. On the contrary, the present technology is intended to cover all modifications, equivalents, and alternatives falling within the scope of the present technology, as defined by the appended claims.
[0005] To maintain data integrity between computing systems, modern computing systems may have data silos created to store data for a given computing system or software application. For example, each data silo may be configured with unique variable names, access protocols (e.g., SQL, AMQP, etc.), data formats (e.g., relational, non-relational, etc.), or other unique characteristics. Having a data silo specifically configured for a given computing system or software application not only enables the computing system / software application to communicate with the given data silo, but also maintains data integrity of the data within the given data silo so that only data within the silo can be modified, thereby protecting data stored in other data silos.
[0006] While data silos offer such benefits, they also result in many drawbacks. For example, one such drawback is that data silos prevent computing systems / software applications that are not configured to communicate with a given data silo from obtaining or receiving data from that data silo. Because each data silo may be configured for a specific software application or computing system, data scientists must reconfigure the data silo or software application / computing system when a new software application is built or when a computing system is scaled. Another drawback is that data silos may store the same or similar information as other data silos. For example, depending on the configuration of such data silos (e.g., variable names, access protocols, data formats, or other characteristics), one data silo may store information associated with a first variable name, and another data silo may store the same information associated with a second variable name, where the first and second variable names are different. Although the variable names are different, the underlying data may be the same (or similar). This results in a large amount of wasted computer memory across computing systems as various copies of the data exist across different data silos. Yet another drawback is that searching for data that may be stored within data silos is often difficult due to their configuration. For example, because each data silo is isolated from the others, there is no common interface for searching all available data silos at once, forcing users to manually search each and every data silo iteratively until they find the data they need to retrieve. Not only is this time-consuming, but such iterative searches require hundreds, if not thousands, of queries to be submitted to each and every data silo, thereby wasting a large amount of computational resources. Retrieving data from distributed silos becomes increasingly complex when you don't know where the data is stored.
[0007] Existing systems have previously attempted to address such shortcomings by utilizing computers and data scientists to create new data silos that (i) can eliminate copies of data and (ii) are capable of communicating with all computing systems / software applications that utilize such data. Yet, manual creation of new data silos is largely impractical to implement. For example, due to the sheer scale of modern computing systems, there may be hundreds, if not thousands, of data silos and corresponding computing systems / software applications that would need to be modified to communicate and utilize such data. Because such computing systems / software applications rely on large amounts of data stored within such data silos to be processed in real time (or near real time), reconfiguring such systems, applications, or data silos can result in significant computing system downtime, thereby impacting user experience.
[0008] Furthermore, even if computer and data scientists manually create new data silos, there is a threat of impacting the data integrity of the data that the data silo stores. For example, when creating a new data silo, computer / data scientists not only must remove copies of the data, but may also need to reformat the data so that the intended computing system / software application can effectively communicate with the data in the data silo. Such modifications to the data may corrupt the data, rendering such valuable data unusable. Even when data scientists create copies of the data silo if the data stored in a given data silo is corrupted, this further exacerbates the problem of wasted computer memory, as even more copies of the data must be made.
[0009] Moreover, creating new data silos or reconfiguring existing computing systems / software applications creates the additional problem of wasting computational resources (e.g., computer processing and computer memory resources) of a given system. For example, when each data silo, computing system, or software application must be reconfigured / created, computational resources are wasted because each new data silo or new computing system / software application occupies a large amount of memory. Therefore, creating these new data silos, computing systems, or software applications further exacerbates these problems.
[0010] For these and other reasons, there is a need to eliminate data copying and simplify data access patterns when accessing siloed data across disparate locations through a unified metadata graph. There is a further need to access siloed data across disparate locations to enable user access to such siloed data without creating new data silos, databases, or reconfiguring existing computer systems and / or software applications. There is a further need to maintain data integrity of data stored within data silos without requiring multiple copies of the data to be stored within the data silos.
[0011] For example, as described above, existing systems lack mechanisms for accessing siloed data across disparate locations without creating new computational components. Because existing systems rely on the creation of new data silos, databases, computing systems, software applications, and the like to access siloed data, such new computational components require significant resources to effectively access the data. Furthermore, because these existing systems rely on the creation of new data silos, the time and energy expended can lead to extended computing system downtime. Moreover, because existing systems are prone to corrupting data during the process of creating such computational components, existing systems rely on creating various copies of the data silos themselves, which can further exacerbate the problem of wasting valuable computer memory resources.
[0012] To overcome these and other deficiencies of existing systems, the inventors have developed a system and method for reducing computational resource usage when accessing siloed data across heterogeneous locations via a unified metadata graph. For example, the system may receive a user-specified query in a graphical user interface (GUI) indicating a request to access a set of data objects, each data object of the set of data objects being stored in a respective data silo of a set of data silos among the heterogeneous locations. For example, the system may receive a user query to access data stored across various data silos. The system may then perform natural language processing on the user-specified query to determine a set of phrases corresponding to the user-specified query. For example, to enable non-technical users to access the data they desire, the system may determine a contextually accurate set of phrases (e.g., based on the user query) to provide the non-technical user with the data they are attempting to access.
[0013] The system then accesses the metadata graph to determine nodes corresponding to the set of phrases. The metadata graph may comprise (i) a set of nodes comprising (a) metadata indicating internal data objects stored in the data silo and (b) location identifiers of the data silo, and (ii) edges indicating data lineage between the set of nodes. For example, by using the metadata graph, the system can traverse the metadata graph, which indicates where data (e.g., data objects) are stored and which data is available in different data silos. In this way, data scientists do not need to create new data silos and / or reconfigure existing computing systems / software applications because the metadata graph may provide an abstraction layer about which data is stored where, thereby reducing computational resource utilization. Moreover, because the metadata graph includes data lineage between the set of nodes (e.g., a representation of the data stored within the data silo itself), the system can further provide information about where copies of data a user intends to access may reside, which the system can leverage to efficiently find where the copied data is hosted. The system then determines a data silo that stores at least one data object of the set of data objects using a location identifier corresponding to the determined node to retrieve at least one data object of the set of data objects via the data silo. The system then generates a visual representation of the at least one data object for display in a GUI. For example, the system can then provide data intended for access by a non-technical user.Thus, by leveraging the power of metadata graphs to access siloed data, the system can reduce the computational resource utilization caused by creating new data silos, computing systems, or software applications to access data stored across different data silos in disparate locations.
[0014] While the use of metadata graphs reduces data retrieval time when accessing siloed data across heterogeneous locations (e.g., data silos hosted in various locations), there is a further need to optimize the generation of such metadata graphs. For example, traditional approaches to locating data can involve manually generating tables containing metadata for siloed data, but generating these tables is inefficient and wastes significant computational resources (e.g., computer memory and processing power) because computer scientists must first find the metadata, normalize the data (e.g., based solely on opinion), and then create the tables. Not only is creating such tables inefficient, but these tables are also error-prone given the tremendous amount of data to consider and the various copies of data inherent in the many copies of data stored in different data silos. To reduce errors and overcome the inherent inefficiencies of traditional approaches, inventors have developed optimized data structures (e.g., metadata graphs) that reduce data retrieval time compared to parsing error-prone metadata tables. The inventors have further developed an optimized method for generating metadata that is less error-prone by leveraging large-scale language models, the metadata itself, and domain-specific languages to reduce the time it takes to generate such data structures while increasing metadata normalization and accuracy to ensure correct labeling of the metadata.
[0015] For example, the system can select a first LLM prompt from a set of large language model (LLM) prompts that corresponds to a first metadata identifier among the set of metadata identifiers. The LLM prompt can correspond to the first metadata identifier based on a data profile (e.g., a data schema, a data format, etc.) of the metadata identifier. The system can then augment the first LLM prompt with the first metadata identifier to be provided to the LLM, and the LLM is configured to generate a first intermediate output that indicates a second set of metadata identifiers that correspond to the first metadata identifier. For example, the system can provide the first metadata identifier to the first LLM prompt to cause the LLM to generate a set of semantically similar metadata identifiers. The set of semantically similar metadata identifiers can represent variations of the first metadata identifier (e.g., to “ask” the LLM what the LLM thinks the first metadata identifier represents).
[0016] The system can then augment the first LLM prompt with a first intermediate output (e.g., a second set of metadata identifiers) to be provided to an LLM, which is configured to generate a second intermediate output indicating filtered domain-specific metadata identifiers by accessing a set of domain-specific ontologies. For example, by providing the augmented LLM prompt to an LLM (e.g., communicatively coupled to the set of domain-specific ontologies), the LLM can leverage contextual knowledge provided by the domain-specific ontologies to generate normalized domain-specific metadata identifiers. Domain-specific ontologies include relationships between phrases, words, or descriptions of data present in an entity's computing system, which can provide a level of contextual knowledge to the entity. The LLM can leverage such contextual knowledge to generate filtered domain-specific metadata identifiers. Moreover, by using an LLM communicatively coupled to domain-specific ontologies, the system can reduce the amount of computational resources required to generate the metadata graph by reducing the dataset of metadata identifiers to be considered (e.g., via access to the domain-specific ontologies).
[0017] The system can then generate a domain-specific unified metadata graph via the LLM using (i) the first metadata identifier and (ii) the second intermediate output indicating the filtered domain-specific identifier. For example, the filtered domain-specific metadata identifier may be a traversable identifier, and the first metadata identifier may be a non-traversable identifier in the domain-specific unified metadata graph. By generating the domain-specific unified metadata graph with the traversable and non-traversable identifiers, the system reduces data search time by reducing the amount of information to be traversed when identifying where data is located (e.g., within a data silo via the metadata graph) while maintaining the verifiability and accuracy of the metadata graph (e.g., by storing the unfiltered, non-domain-specific first metadata identifier in association with the filtered domain-specific metadata identifier). In this way, the system maintains data integrity of the metadata of heterogeneous data silos by transforming the metadata into a verifiable metadata graph for efficiently locating and determining available underlying data stored across the data silos. Finally, to ensure data retrieval time efficiency, the system determines a performance metric of the generated domain-specific integrated metadata graph relative to a previous performance metric of another version of the domain-specific integrated metadata graph. If the performance metric of the generated domain-specific integrated metadata graph cannot satisfy the performance measure relative to the previous performance metric of the other version of the domain-specific integrated metadata graph, the system implements an update process on the domain-specific integrated metadata graph. In this way, the system can ensure that data retrieval time is minimal and accurate when generating, updating, or modifying the domain-specific integrated metadata graph.
[0018] In various implementations, the methods and systems described herein can reduce computational resource utilization when accessing siloed data across heterogeneous locations via a unified metadata graph. For example, the system can receive a query (e.g., via a GUI) indicating a request to access a set of data objects, each data object of the set of data objects being stored in a respective data silo of a set of data silos among the heterogeneous locations. The system can perform natural language processing on the query to determine a corresponding set of phrases. The system can then access a metadata graph to determine nodes corresponding to the set of phrases, the metadata graph comprising (i) a set of nodes comprising (a) metadata indicating internal data objects stored in the data silo and (b) a location identifier of the data silo, and (ii) edges indicating the data lineage of the set of nodes, the metadata graph being generated using a metadata data structure based on file-level and container-level metadata identifiers. The system can then determine a data silo that stored at least one data object of the set of data objects using the location identifier corresponding to the determined node to retrieve at least one data object of the set of data objects via the data silo. The system can then generate a visual representation of the at least one data object for display in a GUI.
[0019] In various implementations, the methods and systems described herein can reduce data search time when accessing siloed data across disparate locations by generating a unified metadata graph via a search expansion generation (RAG) framework. For example, the system selects a first LLM prompt from a set of LLM prompts that corresponds to a first metadata identifier among the set of metadata identifiers. The system then expands the first LLM prompt with the first metadata identifier to be provided to the LLM, and the LLM is configured to generate a first intermediate output indicating a second set of metadata identifiers corresponding to the first metadata identifier. The system then expands the first LLM prompt with the second set of metadata identifiers that correspond to the first metadata identifier to be provided to the LLM, and the LLM is configured to generate a second intermediate output indicating filtered domain-specific metadata identifiers by accessing a set of domain-specific ontologies. The system can then generate a domain-specific unified metadata graph via the LLM using (i) the first metadata identifier and (ii) the second intermediate output indicating the filtered domain-specific metadata identifiers. The filtered domain-specific metadata identifier may be a traversable identifier and the first metadata identifier may be a non-traversable identifier in the domain-specific integrated metadata graph. In response to determining that the first performance metric of the domain-specific integrated metadata graph fails to meet a performance measure for a second performance metric of another version of the domain-specific integrated metadata graph, the system performs an update process on the domain-specific integrated metadata graph.
[0020] This domain-specific unified metadata graph also enables the system to automatically build or apply artificial intelligence (AI) models. An AI sandbox according to implementations herein provides a low-code or no-code environment in which data from heterogeneous locations, as represented by the metadata graph, is used to automatically generate AI models for data analysis or to automatically apply existing AI models to the data.
[0021] In some implementations, a computer system generates a dataset for training or applying to an AI model. The computer system can receive a first natural language input from a user, the first natural language input including a set of phrases and instructions for analyzing data associated with the set of phrases using an artificial intelligence (AI) model. In response to the first natural language input, the computer system accesses a metadata graph to determine nodes corresponding to the set of phrases, the metadata graph including: (i) a set of nodes including (a) metadata indicating internal data objects stored in a data silo and (b) a location identifier of the data silo, and (ii) edges indicating the data lineage of the set of nodes. The system processes the internal data objects indicated by the determined nodes to generate a first set of application data and applies the AI model to the first set of application data to generate one or more first outputs. For example, the first output can include a classification of data items in the first set of application data or a prediction made based on the first application data. A representation of the one or more outputs is sent for display to the user. A second natural language input can then be received from the user, the second natural language input including instructions for modifying the first set of application data. Based on the second natural language input, the computer system generates a second set of application data. The AI model can then be applied to the second set of application data.
[0022] In some implementations, a computer system automates the deployment of AI models. The computer system may receive a first request for a first artificial intelligence (AI) model used by an entity to deploy the first AI model and make the first AI model available for use in a production environment to process input data and generate corresponding outputs. Based on a model deployment engine, the computer system selects a first model deployment location for the first AI model, where the first model deployment location may be selected from a set of one or more cloud provider environments or an on-premises environment operated by the entity. The computer system generates a script for deploying the first AI model at the first model deployment location, and after deploying the model, monitors operating parameters associated with the deployment of the first AI model at the selected model deployment location as the first AI model processes the input data and generates corresponding outputs. The computer system may update the model deployment engine based on the monitored operating parameters, and in response to a second request to deploy a second AI model, select a second model deployment location for the second AI model based on the updated model deployment engine.
[0023] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the implementation of the present technology. Nevertheless, it will be apparent to one skilled in the art that implementations of the present technology may be practiced without some of these specific details.
[0024] The phrases "in some implementations," "in several implementations," "according to some implementations," "in a depicted implementation," "in another implementation," and the like generally mean that a particular feature, structure, or characteristic that follows the phrase is included in at least one implementation of the technology and may be included in one or more implementations. Additionally, such phrases do not necessarily refer to the same or different implementations.
[0025] System Overview 1 illustrates a representation of a graphical user interface (GUI) for reducing computational resource usage when accessing siloed data across disparate locations via a unified metadata graph, according to some implementations of the present technology. For example, the user interface 100 may include a user-specified query input 102, a result output 104, a visual representation of at least one data object 106, and data lineage information 108 (e.g., 108a-108b) for the at least one data object. For example, the user-specified query input 102 may be a data field configured to receive a user-specified query as input. A user may provide a query to the user-specified query input 102 to access data that may be stored across disparate data silos of a computing system. The result output 104 may include one or more visual representations of the at least one data object 106 and the data lineage information 108 corresponding to the at least one data object 106. By way of example, in the context of a non-technical user attempting to find or otherwise access data that may be stored across a set of data silos on one or more computing systems, user interface 100 provides a mechanism for enabling such user to find the data that such user wants or needs.
[0026] Often, a user does not know which data silos (e.g., databases) host the data the user intends to retrieve, nor does the user know exactly what data the user may need for a given application. For example, a non-technical user, such as a business user, may want a list of all of the first names of users who were active last month. Thus, the user can provide a query indicating "I want all of the first names of users who were active last month" to the user-specified query input 102, and the system can generate a result output 104. As described later, the system can perform natural language processing on the user-specified query to obtain a set of phrases (e.g., keywords, semantically similar phrases, etc.) for searching the metadata graph. The metadata graph may be a graph that indicates where data is stored and what data is available. For example, when the user-specified query may be in a question format, the system can determine a set of phrases for accessing the metadata graph by removing unnecessary terms in the user-specified query. The set of phrases may not only be a "cleaned" version of the user-specified query, but also help target what data the user intends to retrieve. By leveraging access to the metadata graph, the system can display a results output 104, which can include a visual representation of at least one data object 106 (e.g., the data the user is attempting to access, the location of the data the user is attempting to access, the format of how the data the user is attempting to access is stored, etc.), and can also include a visual representation of data lineage information 108 (e.g., where copies of the data or similar data may be stored, the format of how the data is stored, etc.).In this way, non-technical users may be provided with a unified, easy-to-use user interface that provides a central access point for accessing data stored across different data silos in different locations while improving the user experience.
[0027] In some implementations, the visual representation of the at least one data object 106 may be interactive. For example, the visual representation of the at least one data object 106 may be a bidirectional link (e.g., a hyperlink) that, upon user selection of the visual representation of the at least one data object 106, may enable the user to access data associated with the at least one data object (e.g., by generating a visual representation of a table storing the at least one data object, generating a window showing the at least one data object, etc.). In this way, the user is enabled to quickly and efficiently browse the data that the user intends to access.
[0028] The right computing environment 2 is a block diagram illustrating some of the components typically incorporated in at least some of the computer systems and other devices on which the disclosed system operates. In various implementations, these computer systems and other devices 200 may include server computer systems, desktop computer systems, laptop computer systems, netbooks, mobile phones, personal digital assistants, televisions, cameras, car computers, electronic media players, web services, mobile devices, watches, wearables, glasses, smartphones, tablets, smart displays, virtual reality devices, augmented reality devices, etc.In various implementations, computer systems and devices include input components 204, including a keyboard, microphone, image sensor, touch screen, buttons, a touch screen, a trackpad, a mouse, a CD drive, a DVD drive, a 3.5 mm input jack, an HDMI input connection, a VGA input connection, a USB input connection, or other computing input components; output components 206, including a display screen (e.g., LCD, OLED, CRT, etc.), a speaker, a 3.5 mm output jack, lights, LEDs, haptic motors, or other output-related components; a processor 208, including a central processing unit (CPU) for executing computer programs, a graphical processing unit (GPU) for executing computer graphics programs and handling computing graphical elements; and data, including programs (e.g., applications 212a-212N, models 214a-214N, and other programs), and associated data, while the programs are being used. the computer system includes zero or more of: storage 210 including at least one computer memory for storing data, an operating system including a kernel, and device drivers; network connectivity components 216 for the computer system to communicate with other computer systems and send and / or receive data, such as via the Internet or another network and its networking hardware, such as switches, routers, repeaters, electrical and optical cables, light emitters and receivers, wireless transmitters and receivers, and the like; persistent storage device 218, such as a hard drive or flash drive, for persistently storing programs and data; and computer-readable medium drive 220 (e.g., at least one non-transitory computer-readable medium), which is a tangible storage means that does not involve a transitory, propagating signal, such as a floppy, CD-ROM, or DVD drive, for reading programs and data stored on a computer-readable medium.Although a computer system configured as described above is typically used to support the operation of the facility, those skilled in the art will understand that the facility may be implemented using devices of various types and configurations and having various components.
[0029] 3 is a system diagram illustrating an example computing environment in which the disclosed system operates in some implementations. In some implementations, the environment 300 includes one or more client computing devices 302a-d, examples of which may host a metadata graph 500 (FIG. 5) (or other system components). For example, the computing devices 302a-d may comprise distributed entities a-d, respectively. The client computing device 302 operates in a networked environment using logical connections through a network 304 to one or more remote computers, such as a server computing device. In some implementations, the client computing device 302 may correspond to device 200 (FIG. 2).
[0030] In some implementations, server computing device 306 is an edge server that receives client requests and coordinates the execution of those requests through other servers, such as servers 310a-c. In some implementations, server computing devices 306 and 310 comprise computing systems. While each server computing device 306 and 310 is logically represented as a single server, each server computing device may be a distributed computing environment encompassing multiple computing devices in the same or geographically disparate physical locations. In some implementations, each server computing device 310 corresponds to a group of servers. In some implementations, server computing devices 306 and 310 host large language models, sets of domain-specific ontologies, artificial intelligence models, user interfaces, web servers, or other computing components.
[0031] The client computing device 302 and the server computing devices 306 and 310 can function as servers or clients to other servers or client devices, respectively. In some implementations, the server computing devices (306, 310a-c) connect to corresponding databases (308, 312a-c). As mentioned above, each server computing device 310 can correspond to a group of servers, each of which can share a database or have its own database (e.g., a data silo). The databases 308 and 312 store (e.g., store) information such as predefined ranges, predefined thresholds, error thresholds, graphical representations, machine learning models, artificial intelligence models, natural language processing models, LLMs, LLM prompts, keywords, metadata graphs, location identifiers, lineage information, semantically similar phrases, file-level metadata identifiers, container-level metadata identifiers, system-level metadata identifiers, governance policies, usage measures, machine learning model training data, artificial intelligence model training data, performance metrics, data schemas, data profiles, or other information. In some implementations, databases 308 and 312 may be data silos.
[0032] Although databases 308 and 312 are logically represented as a single unit, databases 308 and 312 may each be a distributed computing environment encompassing multiple computing devices, may reside within their respective servers, or may reside in the same or geographically disparate physical locations.
[0033] The network 304 can be a local area network (LAN) or a wide area network (WAN), but can also be another wired or wireless network. In some implementations, the network 304 is the Internet or some other public or private network. The client computing devices 302 are connected to the network 304 through a network interface, such as by wired or wireless communication. Although the connections between the server computing devices 306 and 310 are shown as separate connections, these connections can be any type of local, wide area, wired, or wireless network, including the network 304 or a separate public or private network.
[0034] Accessing siloed data across disparate locations FIG. 4 is a flow diagram illustrating a process 400 for reducing computational resource usage when accessing siloed data across disparate locations via a unified metadata graph, according to some implementations of the present technology.
[0035] In act 402, process 400 receives a user-specified query indicating a request to access a set of data objects. For example, the system receives a user-specified query in a GUI indicating a request to access a set of data objects, each data object of the set of data objects stored in a respective data silo of a set of data silos in heterogeneous locations. A data object can be any object, piece of data, or information that can be stored in a data silo, such as a file, information contained within a file (e.g., first name, last name, email address, home address, business address, financial information, account identifier, account number, value, percentage, ratio, alphanumeric string, sentence, etc.), table, data structure, or other data object.
[0036] Data objects (e.g., that a user is attempting to access) may be stored across various data silos (e.g., databases) within a computing environment (e.g., environment 300 (FIG. 3)). For example, a user may wish to access account-related data for one or more user accounts. Yet, the account-related data may be stored in one or more data silos within the computing environment. For example, one data silo may indicate how many accounts are currently open / active (e.g., a first data object), while another data silo may indicate the names of users who have opened accounts (e.g., a second data object). The user may not be aware of where such data, if available, is located. Thus, the user may provide a user-specified query indicating a request to access a set of data objects, and the system may return the data (e.g., a set of data objects) to the user, as described below. In this manner, the system improves the user experience because the user can access such data without requiring prior knowledge of where the data may or may not reside.
[0037] In act 404, process 400 may perform natural language processing to determine a set of phrases. For example, the system may perform natural language processing on the user-specified query to determine a set of phrases corresponding to the user-specified query. Because data stored across data silos may include the same data (e.g., copies of data) or similar data, the system may determine a set of phrases corresponding to the user-specified query for efficiently searching the data stored across the data silos. As an example, one data silo that stores user account information, such as the user's last name, may store the user's last name as a variable called "last_name." Yet, another data silo that stores user account information may store the user's last name as a variable called "given_name." The stored data may be the same (e.g., each silo stores the user's last name), but the variable names may be different. Thus, when searching the data, the system may determine a set of phrases corresponding to the user-provided query for accessing the data.
[0038] In some implementations, the system determines a set of semantically similar phrases that correspond to a user-specified query. For example, the system parses the user-specified query for a set of keywords. The set of keywords may correspond to a set of data objects stored in a data silo. For example, a user provides a query (e.g., "I want the first names of all users who were active last month") The system parses the user-provided query for a set of keywords (e.g., first name, active, etc.). For each keyword, the system can determine a set of semantically similar phrases.
[0039] For example, data may be stored in different silos for different computer applications across an entity's computing systems, so the same or similar data may be stored in a variety of formats. For example, a database storing a table of users' account information may store the user's first name as a variable "name_first," "account_ID," "first_name," "name," or other. In this manner, the system determines a set of semantically similar phrases corresponding to each respective keyword in the set of keywords for searching the metadata graph to retrieve the data the user intends to receive.
[0040] The system can then use the set of semantically similar phrases corresponding to each keyword in the set of keywords to determine a set of phrases corresponding to the user-specified query. For example, continuing the above example, if the user-specified query is "I want the first names of all users who were active last month," the system can determine a first set of semantically similar phrases for "first name" (e.g., "name_first," "account_ID," "first_name," "name") to be used when accessing the metadata graph to determine nodes (e.g., that point to metadata for data objects stored in silos and lineage data for such data objects). In this way, the system can use the metadata graph to reduce the usage of computational resources for accessing siloed data because the system can more efficiently determine the location of needed data based on a set of semantically similar phrases (e.g., when traversing the metadata graph), as opposed to being limited to a single phrase, keyword, or variable name.
[0041] In some implementations, the system can determine semantically similar phrases by accessing a database. For example, the database can indicate a mapping between a first keyword and a second set of keywords. In some implementations, the database can store a predetermined set of keywords generated by a subject matter expert (SME). In this manner, the SME can create such a database to accurately determine which keywords are semantically similar to other keywords, thereby improving the accuracy with which semantically similar phrases are determined.
[0042] In some implementations, the database can be based on an artificial intelligence model. For example, due to the large volume of user-specified queries, the amount of semantically similar phrases, and the unique data that can be searched within data silos, the system can use an artificial intelligence model to determine a set of semantically similar phrases or generate a database to determine semantically similar phrases. The artificial intelligence model may be a machine learning model configured to receive keywords (e.g., phrases) as input and output a set of semantically similar keywords (e.g., semantically similar phrases). Due to the nature of machine learning models (or other artificial intelligence models) that are capable of learning associations between training data (e.g., labeled instances of keywords and semantically similar phrases), the model is not limited to a predefined set of keywords and phrases. For example, the machine learning model can generate new, undiscovered instances of semantically similar phrases corresponding to a given keyword that may be implausible to the human mind. Thus, by using the machine learning model, the system can determine a set of semantically similar phrases corresponding to each respective keyword. In this way, because the machine learning model is not limited to a predetermined set of keywords, the system can determine more robust semantically similar phrases, thereby expanding the range of possible semantically similar phrases that can be generated.
[0043] In response to accessing the database, the system can use each keyword to determine a set of semantically similar phrases corresponding to each keyword. For example, the system can parse the database using each keyword to determine a match between (i) each keyword and (ii) keywords in the database. Upon identifying a match, the system can obtain a set of semantically similar phrases corresponding to the keyword. In this manner, the system can reduce the use of computational resources when determining semantically similar phrases by using the matches, as opposed to performing natural language processing on each keyword to determine the set of semantically similar phrases.
[0044] In act 406, process 400 may access the metadata graph to determine nodes corresponding to the set of phrases. For example, the system may access the metadata graph to determine nodes corresponding to the set of phrases. The metadata graph may include (i) a set of nodes and (ii) edges that indicate the data lineage of the set of nodes. The set of nodes may include (a) metadata that indicates internal data objects stored in the data silos, and (b) location identifiers for the data silos. As an example, the metadata graph may be a graph data structure that indicates metadata for information stored in the set of data silos in environment 300 (FIG. 3).
[0045] As described above, when accessing data that may be stored in data silos in heterogeneous locations, each data silo may be associated with a unique configuration for accessing the data stored within the data silo. When designing a computing system / software application, data scientists and computer scientists may carefully design the data silos, computing systems, and software applications to communicate effectively with each other via one or more communication protocols. Yet, this creates scalability issues when scaling the computing system because the data required for a given computing system / software application may be inaccessible due to the configuration of either the computing system / software application or the data silo itself. Furthermore, searching for the required data may be difficult because the information stored in one data silo may be the same underlying data as in another data silo, albeit with different variable names (e.g., variable identifiers, metadata identifiers, etc.). When searching for such required data for a given computing system / software application, existing systems can parse each and every available data silo for a given match between the data stored in the data silo and the data intended to be accessed (e.g., the required data). Yet parsing each and every data silo in an environment wastes valuable computer processing and memory resources caused by determining whether a match exists between each and every data silo and the information stored in the data silo.
[0046] To address these technical deficiencies, accessing the metadata graph to determine nodes corresponding to a set of phrases (e.g., phrases, keywords, alphanumeric strings corresponding to a user-specified query) can be leveraged to quickly and efficiently identify and access data, thereby reducing computational resource usage.
[0047] Referring to FIG. 5 , which illustrates an exemplary representation of a metadata graph, the metadata graph 500 may include a set of nodes 502a-502m and edges 504a-504q. Each node 502 may be linked or connected to one or more other nodes via one or more edges 504. Each node 502 may indicate metadata for one or more data silos, such as metadata for internal data objects stored within a given data silo (e.g., file-level metadata), metadata for the data silo itself (e.g., container-level metadata), and a location identifier for the given data silo (e.g., where the data silo is located, such as a compute component node identifier, a server identifier, etc.). Each edge 504 may indicate the data lineage of a set of nodes. For example, each edge may represent a lineage relationship between a first node and a second node. That is, each edge may indicate whether a node is a data source for another node or a derivative of another node.
[0048] For example, FIG. 6 shows an expanded view of a metadata graph. In some implementations, the expanded view of the metadata graph 600 may correspond to a portion of the metadata graph 500 in accordance with some implementations of the present technology. By way of example, a first node 602a may indicate metadata for one or more data objects stored within a data silo. For example, the first node 602a may include a file-level metadata identifier 606a, a container-level metadata identifier 608a, and a data silo location identifier 610a. The file-level metadata identifier 606a may be a variable name, a file name, or other identifier that indicates a piece of data stored within a given silo. For example, the file-level metadata identifier may be any identifier (e.g., a variable identifier, a file format, a file size, an access timestamp, or other file-level metadata) that describes data stored within a file stored in the data silo. The container-level metadata identifier 608a may be an identifier that identifies a possible format (e.g., tabular format, tabular format, graphical format, dictionary, etc.) of the data silo, one or more configurations of the data silo (e.g., communication protocols, accessibility parameters, etc.), or other container-level metadata associated with the given data silo. The location identifier 610a may be an identifier that indicates the location of the given data silo. For example, the location identifier may indicate the computer node with which the data silo is associated (e.g., stored, hosted, connected, etc.), the computer system with which the data silo is associated, the location of a server with which the data silo is hosted or otherwise associated, or other location identifier.
[0049] Each node in the set of nodes (e.g., nodes 602a-602d) can have its own file-level metadata identifier 606, container-level metadata identifier 608, or location identifier 610. Because each node in the set of nodes can represent an abstract view of how data is derived from one another, where data is located, and what data is available, the system can leverage the metadata graph to efficiently find where data is located, along with the lineage information for the data itself. That is, the nodes can represent an abstract view of how data is stored across data silos, the relationships between data stored in data silos, and where data is stored between data silos. For example, a first node 602a can be linked to a second node 602b via a first edge 604a. In some implementations, the first edge 604a can indicate lineage information for the node, such as when the second node 602b is a data source for the first node 602a. Nevertheless, in other implementations, the first edge 604a may indicate lineage information, such as when the first node 602a is a data source for the second node 602b, according to some implementations of the present technology. It should be recognized by those skilled in the art that each node 602 may be linked to other nodes via edges 604, with each edge indicating lineage information between one or more nodes in the set of nodes.By representing data objects via a metadata graph that indicates (i) where the data objects (e.g., data stored in a data silo) are located, (ii) the metadata of the data object itself, (iii) the metadata of the data silo that stored the data object, and (iv) the location of such data silo, the system can more efficiently traverse the metadata graph to access data stored in data silos in heterogeneous locations, as opposed to existing systems that rely on users manually parsing each and every data silo for a match between the data they are trying to access and the data stored in the silo itself, thereby reducing the use of computational resources when accessing siloed data across heterogeneous locations.
[0050] In some implementations, the system can determine nodes corresponding to the set of phrases by traversing the metadata graph. For example, the system can traverse each node in the set of nodes of the metadata graph. The system can compare a metadata identifier of a given node to each phrase in the set of phrases. For example, the metadata identifier may be a file-level, container-level, or other identifier that indicates that a given data silo contains data related to the phrase. For example, the metadata identifier may be "first_name" (e.g., a file-level metadata identifier) that indicates that the data silo contains a user's first name. In response to determining that the metadata identifier matches at least one phrase in the set of phrases, the system can determine nodes corresponding to the set of phrases.
[0051] For example, as opposed to traversing a metadata graph using a single phrase, the system traverses the metadata graph and compares each phrase in a set of phrases with the metadata identifier of a given node. That is, as opposed to existing techniques that traverse a graph (e.g., a metadata graph or other graph) using a given keyword, the system traverses the graph using a set of phrases. In this way, the system can more efficiently determine nodes that correspond to a set of phrases because the system does not need to perform multiple traversals of the graph using a different phrase each time, thereby reducing computational resource usage.
[0052] In some implementations, the system can determine another data silo that stored the second data object. For example, the system can traverse each node of the set of nodes (e.g., of the metadata graph) to identify a metadata identifier that matches at least one phrase of the set of phrases. In response to determining that the metadata identifier matches at least one phrase of the set of phrases, the system determines a first node that corresponds to the set of phrases. Although the system may have determined a first node that corresponds to the set of phrases (e.g., thereby determining a data silo that stored a data object associated with the set of phrases), the system can still continue traversing the metadata graph to determine other locations (e.g., of data silos) that host the given data object.
[0053] For example, if a user-specified query indicates "I want all locations where the user's first name resides," the system can continue to traverse the set of nodes using edges connected to the given node. For example, in response to determining that a first node corresponds to a set of phrases, the system can perform a second traversal of nodes in the set of nodes to determine a second node using an edge that indicates a first data lineage of the first node. The first data lineage of the first node can indicate a second node that contains information that is a source of the information associated with the first node. For example, each edge in the metadata graph can indicate a lineage of a data object. Because each node in the set of nodes indicates metadata (e.g., data of data), edges between nodes can indicate that one node is a source of another node (or alternatively, a derived data source of another node).
[0054] To illustrate, with reference to FIG. 6 , a system may determine a first node corresponding to a set of phrases, such as a first node 602a. The system may traverse using a first edge 604a to a second node 602b, using a second edge 604b to a third node 602c, or using a third edge 604c to a fourth node 602d. In some implementations, after performing a first traverse (e.g., from the first node 602a to the second node 602b), the system may perform a second traverse (e.g., from the first node 602a to the third node 602c). The system may iteratively repeat such traversals until each node has been traversed or until, after performing a traversal, no nodes remain that correspond to the set of phrases. In this way, the system may more efficiently access siloed data by traversing the metadata graph, as opposed to parsing each and every data silo in a computing environment for matches.
[0055] Thus, the system can determine a second data silo that stores a second data object (e.g., the same data object or a similar data object related to at least one of the data objects) by retrieving a second data object of the set of data objects through the second data silo using a location identifier corresponding to the second node. That is, the system can determine alternative locations (e.g., data silos) where a given data object may be stored by traversing the metadata graph using edges connected to the determined node. In this manner, the system can determine all locations where the same or similar data may be stored. In some implementations, the system can then generate a visual representation of the second data object on a GUI. In this manner, the user can be provided with additional data of interest to the user.
[0056] In some implementations, in response to determining each data silo in which a given data object is stored, the system can perform one or more data aggregation techniques. For example, the system can remove unnecessary instances of the data itself. For example, because the metadata graph is an abstraction that indicates where data is located and what data a given silo may contain, the system can remove all but one instance of the data (e.g., the data object) to reduce the amount of computer memory utilized.
[0057] In some implementations, process 400 can generate a metadata graph using the generated metadata data structure. For example, the system can retrieve (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers from each data silo in a given environment (e.g., environment 300). Each file-level metadata identifier in the set of file-level metadata identifiers indicates metadata for a given data object stored in the respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates metadata for a respective data silo in the set of data silos in the given environment. The system can generate sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, respectively. For example, the system can perform natural language processing on the file-level and container-level metadata identifiers to determine sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, respectively. For example, for a file-level metadata identifier of “first_name,” the system can generate sets of semantically similar metadata identifiers such as “name_first,” “account_ID,” “user_id,” “name,” or others.
[0058] The system may then generate a metadata data structure for mapping each semantically similar metadata identifier in the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier. For example, to enable the system to efficiently search data across the metadata graph, the system may generate a normalized metadata identifier corresponding to each semantically similar phrase (e.g., by using natural language processing, a machine learning model, an artificial intelligence model, etc.). For example, a normalized metadata identifier for a set of semantically similar metadata identifiers of "name_first," "account_ID," "user_id," and "name" may be "first_name_ID," and the metadata data structure maps "first_name_ID" to each of the semantically similar metadata identifiers. In some implementations, the system may generate a metadata graph using the generated metadata data structure (e.g., the normalized metadata identifiers, the set of semantically similar metadata identifiers, etc.). Additionally or alternatively, the system may generate the metadata graph based on an artificial intelligence model. In this manner, the system can optimize the metadata graph by using normalized container-level and file-level metadata identifiers associated with nodes in the metadata graph to enable more efficient data searches.
[0059] Referring again to FIG. 4, in act 408, process 400 may determine a data silo that stored at least one data object. For example, the system may use a location identifier corresponding to the determined node (e.g., of act 406) to determine a data silo of the set of data silos that stored at least one data object of the set of data objects, in order to retrieve at least one data object of the set of data objects via the data silo. Because each node of the set of nodes of the metadata graph includes a location identifier corresponding to a data silo (e.g., indicating which data silo stores a given data object), the system may access the data silo using the location identifier to retrieve the data object. For example, the system may use the location identifier of the determined node to determine which data silo hosts at least one data object of the set of data objects. In some implementations, the system may use location identifiers corresponding to other determined nodes to determine each data silo that stored each data object of the set of data objects, in order to retrieve the set of data objects. Using the location identifier, the system may determine a communication protocol associated with the determined data silo to retrieve the at least one data object. For example, since each data silo may be associated with a unique communication protocol, the system can identify which communication protocol the determined data silo is associated with, select the communication protocol (e.g., query language, access protocol, configuration, etc.) to communicate with the data silo, and provide a query to the data silo. Thus, the system can retrieve at least one data object via the query. In this manner, the system can reduce the amount of wasted computational resources when accessing siloed data across heterogeneous locations via a metadata graph.
[0060] In act 410, process 400 may generate a visual representation of at least one data object for display. For example, the system may generate a visual representation of the at least one data object for display in a GUI. In some implementations, the visual representation of the at least one data object includes lineage information for the at least one data object. For example, with reference to FIG. 1 , a visual representation of at least one data object 106 may be presented for display within user interface 100 along with lineage information 108 corresponding to the at least one data object. It should be recognized by those skilled in the art that user interface 100 may include one or more visual representations (e.g., of data objects and / or lineage information) that may correspond to a set of data objects according to one or more implementations of the present technology.
[0061] In some implementations, the system can use an artificial intelligence model to generate an intended result. For example, the system can receive a second user-specified query via a second GUI indicating a request to generate an intended result. For example, a user can provide a query indicating a request to generate an intended result (e.g., a prediction) using an artificial intelligence model. The intended result can be any user-specified prediction that the user wants to receive. In the context of a non-technical user, such a user may be ignorant about which artificial intelligence / machine learning model to select to generate a given prediction, what data to use to train a given artificial intelligence / machine learning model, or other components / data to use to generate a given prediction. Nevertheless, the user may know what the user wants to discover (e.g., how many accounts will be opened in the next three months, in which week a company is likely to receive an influx of opened accounts, what the expected cost of monitoring a set of accounts for fraudulent activity over a given period of time is, how many users / accounts are active, how many users / accounts are inactive, etc.). To enable such non-technical users to obtain the intended results, the system may provide a GUI (which may be the same GUI or similar to the GUI described in FIG. 1) that enables the user to provide a query to generate the intended results, and may provide recommendations as to which artificial intelligence / machine learning model should be used to generate the intended results, and which training data should be used to train the artificial intelligence / machine learning model to generate the intended results.
[0062] The system can provide a second user-specified query to the artificial intelligence model to generate a recommendation, the recommendation including (i) a second artificial intelligence model to be used to generate the intended result and (ii) a second set of data objects to be used in training the second artificial intelligence model. For example, the system can provide the user-specified query to an artificial intelligence model (e.g., a machine learning model, model 702 (FIG. 7)) trained to generate the recommendation. The artificial intelligence model can generate a recommendation indicating which artificial intelligence model to use to generate the intended result and with which training data the given artificial intelligence model should be trained to generate the intended result. For example, the recommended artificial intelligence model can be an artificial intelligence model or machine learning model that can be configured to generate the intended result. Such a recommended artificial intelligence model / machine learning model can be a deep learning model, a neural network, a convolutional neural network, a recurrent neural network, a support vector machine, a natural language processing model, a KNN model, a linear regression model, a logistic regression model, a random forest model, a Bayesian model, or other artificial intelligence / machine learning model. In this way, the system provides recommendations on which artificial intelligence model should be used to generate the intended result and which training data should be used to train the artificial intelligence model, thereby reducing the utilization of computational resources that would otherwise be wasted by a non-technical user performing numerous incorrect iterations of training a machine learning model to generate the intended result.
[0063] In response to receiving a user selection indicating acceptance of the recommendation, the system may (i) access a database to obtain a second artificial intelligence model and (ii) use the metadata graph to obtain a second set of data objects, according to some implementations of the present technology. For example, the system may generate a message (e.g., a notification, a user-selectable object, etc.) to enable the user to accept the recommendation (e.g., via a button, a text-based command, a checkbox, etc.). In some implementations, the system may automatically accept the recommendation without a user selection to accept the recommendation. In this manner, the system may automatically select a recommended artificial intelligence model and training data for producing the intended result, thereby improving the user experience. The system may then access a database (e.g., an artificial intelligence model database) that stores untrained or pre-trained artificial intelligence / machine learning models and obtain the recommended artificial intelligence model (e.g., via an artificial intelligence model identifier, a machine learning model identifier, etc.). The system may also access the metadata graph to obtain a second set of data objects (e.g., to be used as training data for the recommended artificial intelligence / machine learning model). For example, the second set of data objects may be training data stored in one or more data silos of environment 300 that will be used as training data for the artificial intelligence model. In response to obtaining the recommended artificial intelligence model and the second set of data objects, the system may train the recommended artificial intelligence model using the second set of data objects (e.g., training data) and apply the recommended artificial intelligence model (e.g., to the input data) to generate the intended result. For example, the system may provide new input data (e.g., new data obtained via the metadata graph) as input to the recommended artificial intelligence model for generating the intended result (e.g., based at least in part on a user-specified query).In this manner, non-technical users may be enabled to use artificial intelligence models to generate one or more intended results, thereby improving the user experience.
[0064] In some implementations, the system can determine whether the output of an artificial intelligence model is approved to be provided to one or more computing systems. For example, because artificial intelligence and machine learning models are used in various domains related to entities (e.g., companies, businesses, etc.), the use of such artificial intelligence / machine learning models may be required to conform to one or more governance standards when using such models for one or more functions. Because non-technical users may use such models to generate predictions, discover new relationships between existing data, or for other functions, the system can ensure that the use of such models, the data provided to the models, and the outputs generated by the models comply with one or more industry, government, or internal standards. In this way, the system can reduce the likelihood of data breaches, thereby improving data security.
[0065] For example, the system may access a governance database to obtain a set of policies that dictate usage measures corresponding to a set of data objects. The governance database may store policies (e.g., governance policies, industry standards, internal company policies, etc.) that dictate usage measures (e.g., definitions or other measures regarding how data may be used, generated, provided to other computing systems, provided to external computing environments, made public, etc.). The system may access the governance database to obtain a set of policies that dictate usage measures for a second set of data objects (e.g., data used to train a recommended artificial intelligence model) and may use the set of policies to determine whether the second set of data objects is approved for use to train the recommended artificial intelligence model. For example, in some implementations, the system may provide (i) the second set of data objects and (ii) the obtained set of policies (e.g., corresponding to the second set of data objects) to another artificial intelligence / machine learning model (e.g., model 702 (FIG. 7)) configured to generate a prediction about whether the second set of data objects is approved for use to train a second artificial intelligence model. The system may also determine whether the output of the second artificial intelligence model (e.g., the recommended artificial intelligence model) is approved for provision to one or more computing systems using a second set of policies that dictate usage measures corresponding to the artificial intelligence model predictions. For example, the second set of policies may include information regarding which types of artificial intelligence model predictions may be transmitted, provided, published, or sent to internal or external computing systems.In response to (i) approval of the second set of data objects to be used to train a second artificial intelligence model and (ii) approval of the output of the second artificial intelligence model to be provided to one or more computing systems, the system may apply the second artificial intelligence model (e.g., a recommended artificial intelligence model) to generate an intended result. In this manner, the system may scrutinize the training data and any output that may be generated by the artificial intelligence model before generating the intended result, thereby reducing the likelihood of a data breach caused by providing such output to one or more computing systems.
[0066] Referring to FIG. 7, FIG. 7 shows a diagram 700 of an artificial intelligence model according to some implementations of the present technology. The model 702 can take in input 704 and provide output 706. The input can include multiple datasets, such as a training dataset and a test dataset. Each of the multiple datasets (e.g., the input 704) can include a data subset related to user data, predicted predictions and / or errors, and / or actual predictions and / or errors. In some embodiments, the output 706 can be fed back to the model 702 as an input for training the model 702 (e.g., alone or together with a user indication of the accuracy of the output 706, a label associated with the input, or other reference feedback information). For example, the system can receive a first labeled feature input, where the first labeled feature input is labeled with a known prediction for the first labeled feature input. The system can then train a first machine learning model to classify the first labeled feature input (e.g., a response to a user-provided query) with the known prediction.
[0067] In various implementations, the model 702 can update its configuration (e.g., weights, biases, or other parameters) based on an evaluation of its predictions (e.g., output 706) and reference feedback information (e.g., user instructions about accuracy, reference labels, or other information). In various implementations, if the model 702 is a neural network, the connection weights may be adjusted to reconcile the difference between the neural network's predictions and the reference feedback. In a further use case, one or more neurons (or nodes) of the neural network may require their respective errors to be sent backward through the neural network (e.g., backpropagating errors) to facilitate the update process. The updates to the connection weights may, for example, reflect the magnitude of the error propagated backward after a forward pass is completed. In this way, for example, the model 702 may be trained to generate better predictions.
[0068] In some implementations, model 702 may include an artificial neural network. In such implementations, model 702 may include an input layer and one or more hidden layers. Each neuronal unit of model 702 may be connected to many other neuronal units of model 702. Such connections may be constraining or inhibitory in their effect on the activation states of the connected neuronal units. In some implementations, each individual neuronal unit may have a summary function that combines the values of all of its inputs. In some implementations, each connection (or the neuronal unit itself) may have a threshold function that a signal must exceed before propagating to other neuronal units. Model 702 may be self-learning and trainable rather than explicitly programmed, and may perform significantly better in certain domains of problem solving than traditional computer programs. During training, the output layer of model 702 may correspond to a classification of model 702, and inputs known to correspond to this classification may be input to the input layer of model 702 during training. During testing, inputs with no known classification may be input to the input layer, and a determined classification may be output.
[0069] In some implementations, model 702 may include multiple layers (e.g., signal paths traverse from front layers to back layers). In some implementations, backpropagation techniques may be utilized by model 702, with forward stimuli being used to reset weights on "front" neural units. In some implementations, stimuli and inhibition for model 702 may be more fluid, with connections interacting in a more chaotic and complex manner. During testing, the output layer of model 702 may indicate whether a given input corresponds to the model's 702 classification (e.g., a response to a user-provided query).
[0070] In some implementations, the model (e.g., model 702) can automatically perform an action based on the output 706. In some implementations, the model (e.g., model 702) may not perform any action. The output of the model (e.g., model 702) may be used to direct or otherwise generate a metadata graph, determine a set of phrases, determine semantically similar phrases, provide artificial intelligence / machine learning model recommendations, determine whether a data object is approved to be used to train an artificial intelligence / machine learning model, determine whether an artificial intelligence / machine learning model output is approved to be provided to one or more computing systems, generate a response, or generate other information, in accordance with one or more implementations of the present technology.
[0071] In some implementations, a model (e.g., model 702) can be trained based on training information stored in database 308 or database 312 to generate recommendations. For example, the recommendations may be recommendations for a given artificial intelligence / machine learning model to generate an intended result and recommendations for which training data should be used when training the given artificial intelligence / machine learning model. Model 702 can take a first set of training information as input 704 and generate an output (e.g., a recommendation, multiple recommendations) as output 706. The first set of training information can include a user-specified query indicating a request to generate an intended result (e.g., a prediction), an artificial intelligence / machine learning model identifier used to generate the intended result, training data used to train the artificial intelligence / machine learning model used to generate the intended result, or other information. For example, model 702 can learn associations between the first set of training information to generate recommendations as output 706. The output 706 may be a recommendation as to which artificial intelligence model should be selected to produce the intended result and which training data should be used to train the artificial intelligence model to produce the intended result. In some embodiments, the output 706 may be fed back to the model 702 to update one or more configurations (e.g., weights, offsets, or other parameters) based on the model's 702 evaluation of its predictions (e.g., output 706) and reference feedback information (e.g., user instructions for accuracy, reference labels, ground truth information, known recommendations, etc.). The first set of training information may be historical training information that has been used to train a prior artificial intelligence / machine learning model to produce a given intended result.In this manner, model 702 may be trained to generate one or more recommendations regarding which artificial intelligence / machine learning models can produce the intended results, as well as the training data necessary to train such artificial intelligence / machine learning models, thereby enabling non-technical users to take advantage of artificial intelligence / machine learning models.
[0072] In some implementations, a model (e.g., model 702) can be trained based on training information stored in database 308 or database 312 to make approval decisions. For example, model 702 can be trained to determine whether training data for a given artificial intelligence / machine learning model is approved for use in training the artificial intelligence / machine learning model and whether the output of the artificial intelligence / machine learning model is approved for publication, transmission, or provision to one or more computing systems. For example, as described above, with the increasing number of artificial intelligence and machine learning models used in business contexts, such models are under scrutiny and must be vetted before being applied to sensitive user data. To vet such a model, model 702 can take a second set of training information as input 704 and generate an output (e.g., an approval, multiple approvals) as output 706. The second set of training information may include predictions generated by the artificial intelligence / machine learning model, an artificial intelligence / machine learning model identifier used to generate the predictions, training data used to train the artificial intelligence / machine learning model used to generate the predictions, a set of policies that indicate usage measures corresponding to data objects (e.g., training data) used to train the artificial intelligence / machine learning model used to generate the predictions, a second set of policies that indicate usage measures corresponding to the artificial intelligence model predictions, or other information. For example, model 702 may learn associations between the second set of training information to generate an approval as output 706. Output 706 may be an approval that indicates whether the second set of data objects (e.g., training data) is approved for use to be used to train the artificial intelligence / machine learning model and whether the output (e.g., predictions) of the artificial intelligence / machine learning model is approved to be provided to one or more computing systems.In some embodiments, the output 706 may be fed back to the model 702 to update one or more configurations (e.g., weights, offsets, or other parameters) based on the model's 702 evaluation of its predictions (e.g., output 706) and reference feedback information (e.g., user instructions for accuracy, reference labels, ground truth information, known recommendations, etc.). The second set of training information may be historical information that has been used to provide recommendations for different data objects (e.g., training data) and machine learning models. In this manner, the model 702 may be trained to scrutinize the artificial intelligence model / machine learning model, its input data, its training data, and its output data before being used, according to one or more implementations of the present technology.
[0073] Generating a unified metadata graph FIG. 8 illustrates a process for generating a unified metadata graph via a search extension generation (RAG) framework, according to some implementations of the present technology.
[0074] In act 802, process 800 selects a first LLM prompt. For example, the system selects a first LLM prompt from a set of LLM prompts that corresponds to a first metadata identifier in the set of metadata identifiers. Each LLM prompt in the set of LLM prompts may be associated with a data schema, data format, data type, or other characteristic of the metadata identifier. For example, an LLM prompt associated with a data type of the metadata identifier may reference a file-level metadata identifier, a container-level metadata identifier, a system-level metadata identifier, or another metadata identifier. For example, the type of metadata identifier may dictate the structure of the LLM prompt to be selected for use.
[0075] An LLM prompt can be structured with respect to the data type of the metadata identifier. For example, a structured LLM prompt can refer to input configured to be interpreted by an LLM in a structured format. A structured LLM prompt is a prompt for a text-to-text language model (e.g., an LLM) that is structured so that the text included in the structured LLM prompt is interpreted and understood by the LLM. Each LLM prompt can be structured for the data schema, data format, data type, data profile, or other characteristics of the metadata identifier, so that the LLM prompt can provide additional information to the LLM when generating output. For example, a structured LLM prompt structured for the data type of the metadata identifier can include one or more attributes, tags, labels, or other information that indicate that the metadata identifier included in the LLM prompt is of a particular type. In this way, the LLM can generate a more accurate first intermediate output that indicates a set of metadata identifiers that correspond to a first metadata identifier.
[0076] For example, referring to FIGS. 9A-9B , which show illustrative diagrams of LLM prompts according to some implementations of the present technology, an exemplary LLM prompt 900 may include prompt 1 902, prompt 2 910, prompt 3 914, prompt 4 916, and prompt 5 922. Additionally, for illustrative purposes, an LLM 906 is shown. Prompt 1 902 may include a level 903, a first metadata identifier 904, and first prompt text 905. For example, the level 903 may reference a data schema, data format, data type, or other characteristic of the first metadata identifier 904 to provide the LLM 906 with additional information about which kind or type of metadata identifier the LLM should consider. The first prompt text 905 may be structured text associated with the level 903. For example, the first prompt text 905 may be unique to the level 903. For example, the first prompt text 905 may be shown to indicate "provide a set of similar identifiers of," which may be text corresponding to the level 903, where the level 903 indicates a file-level metadata identifier type and the first metadata identifier 904 indicates a file-level metadata identifier. In some implementations, the first prompt text 905 may differ based on the metadata identifier of the first metadata identifier 904. For example, if the first metadata identifier 904 is a container-level metadata identifier, the first prompt text 905 may alternatively recite "provide a set of similar container-level identifiers of," where the level 903 indicates the first prompt text 905. That is, upon determining the metadata identifier type of the first metadata identifier 904, prompt 1 902 may be selected, where prompt 1 902 is associated with the level 903 indicating a metadata identifier type, and the first metadata identifier 904 further includes the correct first prompt text 905.In this manner, the LLM prompt may be structured based on the data schema, data format, data type, or other characteristics of the first metadata identifier to obtain more accurate results from the LLM, as opposed to generic LLM prompts in existing systems that do not rely on specifically generated LLM prompts.
[0077] To generate the metadata graph, the system can utilize a RAG framework. For example, a RAG framework, or alternatively, RAG, can refer to a framework that enables an artificial intelligence model (e.g., a large-scale language model) to access data sources that may contain information that undergo updates without requiring retraining of the entire LLM. Traditionally, LLMs are trained on large corpora of data to provide output based on input prompts. Yet, LLMs are often limited to the training data on which the LLM is trained, and the training process for LLMs is exceptionally computationally intensive. To overcome these shortcomings of LLMs, a RAG method can be employed to ensure that the LLM is provided with the latest data without requiring full retraining of the LLM.
[0078] Moreover, using RAG, the LLM can query for additional information on which the LLM was not previously trained. For example, while an LLM can often produce outputs that may appear factually correct on the surface, the LLM lacks a mechanism for deciphering between what is true and what is not. Rather, the LLM stipulates that the output it interprets is the most correct output given the input (e.g., the prompt). The LLM can be communicatively coupled to one or more data sources to provide a mechanism that allows the LLM to access information it was not previously trained on as well as provide a source of truth (e.g., verifiable information on which the LLM may generate output). For example, in the context of generating a metadata graph via the RAG framework, the LLM can be communicatively coupled to the raw data component 1010 (e.g., to “extract” metadata identifiers) and a set of domain-specific ontologies in the entity domain ontology component 1008 to return the generated filtered domain-specific metadata identifiers (based on the augmented LLM prompts that include the extracted metadata identifiers).
[0079] For example, referring to FIG. 10A , FIG. 10A shows a subsystem diagram 1000 illustrating an example RAG framework environment for generating a unified metadata graph. The subsystem diagram 1000 can provide an example of a RAG framework environment according to some implementations of the present technology, where the RAG framework environment can include a user interface 1002, a metadata graph 1004, an LLM 1006, a domain ontology component 1008, a raw data component 1010, a feedback component 1012, and communication links 1014a-1014p. For example, the user interface 1002 is a user interface for receiving a user-specified query indicating a request to access a set of data objects (e.g., as described in act 402 of FIG. 4). The metadata graph 1004 can be a metadata graph (e.g., as described in act 406 of FIG. 4). The LLM 1006 can be any LLM configured to provide an output in response to an input (e.g., BERT, Claude, Cohere, Ernie, Falcon40B, Galactica, etc.) For example, the LLM 1006 can be configured to receive an LLM prompt and provide a response (e.g., text, graphical, etc.) in response to the LLM prompt.
[0080] The communications links 1014a-1014p may enable communications between the user interface 1002, the metadata graph 1004, the LLM 1006, the domain ontology component 1008, the raw data component 1010, the feedback component 1012, or other components (shown or not shown). For example, the communications links 1014a-1014p may include the Internet, a mobile phone network, a mobile voice or data network (e.g., a 5G or LTE network), a cable network, a public switched telephone network, or other types of communications networks or combinations of communications networks. The communications links 1014a-1014p may include one or more communications paths, separately or together, such as satellite paths, fiber optic paths, cable paths, paths supporting Internet communications (e.g., IPTV), free-space connections (e.g., for broadcast or other wireless signals), or any other suitable wired or wireless communications paths or combinations of such paths.
[0081] The domain ontology component 1008 may be a database, server, or other computational component configured to store a set of domain ontologies related to entities. For example, the domain ontology component 1008 may store a domain ontology that indicates a set of concepts and categories in a given subject area (e.g., a domain), providing information about the nature of the concepts / categories and the relationships between the concepts / categories. According to one or more implementations of the present technology, the domain ontology component 1008 may store a set of domain ontologies specific to an entity of the system. For example, if the entity is a company, the domain ontology may reflect domain-specific knowledge (e.g., nomenclature, taxonomy, lexicography) about the terms used in the entity's domain. For example, if the entity is a bank, the domain ontology component 1008 may include an ontology that relates a given financial term to other financial terms to infer the context in which the given financial term is used.
[0082] Such domain-specific (e.g., entity-specific) contextual knowledge is advantageous to leverage in connection with generating a metadata graph, because such knowledge can be used to generate normalized filtered domain-specific metadata identifiers that match the terminology used by the entities in their daily work. By communicatively coupling the LLM to a set of entity-specific domain ontologies, the system can generate normalized filtered domain-specific metadata identifiers to be used in generating the metadata graph. In doing so, the system can extract, identify, and locate data that users of the system intend to locate from the metadata graph based on a common terminology, context, or domain. Moreover, by leveraging domain ontologies associated with entities, the system reduces errors when generating normalized filtered domain-specific metadata identifiers because the entity's terminology and contextual knowledge is maintained. For example, different entities may have different meanings for a given term. By leveraging domain-specific ontologies, the system can reduce errors when generating normalized filtered domain-specific metadata identifiers to be used in the metadata graph because the LLM can "look up" a set of domain-specific ontologies to validate the output of the LLM (e.g., normalized filtered domain-specific metadata identifiers). In addition to validating the output of the LLM (which may, e.g., include the metadata graph 1004 itself), the feedback component 1012 can be used to update, confirm, or verify (e.g., additions to) the metadata graph 1004 during the metadata graph generation process. For example, the feedback component 1012 can include one or more user or automated inputs to verify the accuracy of the metadata graph 1004 (described in more detail below).
[0083] The raw data component 1010 can be a data source that provides raw data to the LLM. For example, the raw data component 1010 provides metadata identifiers, data profiles (e.g., of metadata identifiers, data silos, data objects, systems), or other raw data to the LLM 1006 when generating the metadata graph 1004. For example, the raw data component 1010 can obtain raw data from data silos of a system. For example, the system can receive raw data from a set of data silos that includes a set of metadata identifiers that indicate (i) a file-level metadata identifier, (ii) a container-level metadata identifier, or (iii) a system-level metadata identifier. The file-level metadata identifier can indicate metadata for a data object stored within one of the set of data silos, the container-level metadata identifier can indicate metadata for one of the data silos, and the system-level metadata identifier can indicate metadata for a computing system that hosts one of the data silos. For example, a file-level metadata identifier may indicate a label for a data object stored within a data silo, a container-level metadata identifier may indicate a label for a data format in which the data silo stores data, and a system-level metadata identifier may indicate a label for an operating system or system identifier that hosts the data silo.
[0084] In some implementations, the system can perform a crawling process across a set of data silos. For example, to obtain raw metadata, the system can perform a crawling process across a set of data silos associated with an entity (e.g., a company, a merchant, a corporation, a business, a computing environment, etc.) to obtain raw data comprising a set of metadata identifiers. For example, referring to FIG. 10B showing a subsystem diagram of raw data component 1010, the crawling process may be performed by crawler 1016, which can be any database crawling service configured to extract metadata from data silos 1015. Data silos 1015 may be the same or similar data silos as those described in acts 402-410 of process 400 (FIG. 4). For example, crawler 1016 can generate a set of crawl queries to obtain file-level, container-level, or system-level metadata (e.g., metadata values, metadata identifiers, etc.). Parser 1018 can parse the set of crawl queries to extract file-level, container-level, or system-level metadata identifiers. In this way, the system can retrieve all available metadata from data silos for use in generating a more robust and accurate metadata graph, as opposed to existing methods that rely on manual labeling techniques. In some implementations, metadata is retrieved via a combination of a crawler 1016 and a parser 1018, and via manually labeled data entities.
[0085] In some implementations, the system can generate a data profile for each data silo in the set of data silos. When generating a domain-specific integrated metadata graph, the data silos themselves can store various data of different schemas, types, and formats and can also have different contexts. Profiling such data silos is advantageous because these data profiles can indicate valuable contextual information that can impact the given structure of the LLM prompt, thereby impacting the ultimate results received by the LLM. For example, each LLM prompt can be key to achieving a particular intended result (e.g., to obtaining a normalized metadata identifier with a particular context or domain), so when providing a prompt to the LLM, the structure of the prompt can include various data elements to achieve more efficient and accurate results. As an example, an LLM prompt augmented with a metadata identifier and a data type corresponding to this metadata identifier can cause more accurate results to be generated, as opposed to an LLM prompt that only includes a metadata identifier (e.g., because it may lack additional contextual information). Thus, the system can generate a data profile for each data silo in a set of data silos for expanding or selecting structured LLM prompts for processing.
[0086] For example, the parser 1018 can extract a first value from each data silo of the set of data silos. Because each data silo can store a unique set of data, the system only needs to extract at least one value from each of the set of data silos. Yet, in other implementations, the system can extract one or more values from each data silo of the set of data silos. The profiler 1020 can then determine a data type corresponding to each first value extracted from each of the set of data silos. For example, the profiler 1020 may be a logical component that can determine the data type corresponding to the first value. The data type may relate to the data schema of the first value, the format of the first value, whether the first value is an integer, a character, a floating point, a double-precision floating point, or some other data type. Using the data type of the first value, the profiler 1020 can generate a data profile for each data silo of the set of data silos that indicates the data type of the values stored in the data silo. For example, the system may generate a data profile (e.g., file, text file, tag, etc.) associated with each data silo (e.g., container) in the entity's computing system that indicates the data type of the values stored in each of the data silos. Such data profiles may be stored in a database for later retrieval and associated with their respective data silos. In this manner, the system may index the data types associated with each data silo in order to accurately select structured LLM prompts with contextual information (e.g., data profile).
[0087] In some implementations, to select a first LLM prompt from the set of LLM prompts, the system can filter the set of LLM prompts. For example, as described above, the data profile (e.g., data types) of each data silo can add advantageous contextual information to use when selecting structured LLM prompts to generate a domain-specific integrated metadata graph. For example, by augmenting a specifically designed LLM prompt with contextual information (e.g., the data profile of the data silo) from which the metadata identifiers originate, the system can achieve more accurate results, in contrast to existing systems that cannot add such contextual information but rather rely on learned knowledge of the LLM itself.
[0088] Thus, the system can determine the data silos that store data corresponding to the first metadata identifier. For example, the system can compare the first metadata identifier with each metadata identifier stored in each of the data silos for a match. In other implementations, the system can still reference a database that stores a mapping between the metadata identifiers and the data silos that stored the data associated with the metadata identifiers. The system can then retrieve a data profile that corresponds to the data silos that stored the data corresponding to the first metadata identifier. For example, as described above, the system can retrieve a generated data profile for the data silo.
[0089] The system can then use the retrieved data profile to filter the set of structured LLM prompts to generate a set of filtered LLM prompts. For example, each LLM prompt in the set of LLM prompts may be tagged with one or more tags indicating (i) a metadata identifier, (ii) a data profile (e.g., data type), (iii) the architecture of the LLM prompt, and / or (iv) other tags (e.g., data schema, data format, or other characteristics). The system can filter the set of structured LLM prompts into a subset of LLM prompts (e.g., a filtered set of LLM prompts) to reduce the amount of computational resources utilized when comparing LLM prompts. The filtering not only reduces utilization of computational resources (e.g., computer memory and processing power), but can also provide a reduced set of LLM prompts to select from based on the data profile of the data silo associated with the metadata identifier, thereby improving LLM prompt selection accuracy. The system can then select a first structured LLM prompt from the set of filtered LLM prompts that corresponds to a first metadata identifier in the set of metadata identifiers. For example, the system selects a first structured LLM prompt based on a match between a tag of the LLM prompt that indicates the data format of the LLM prompt and the data format of the metadata identifier.
[0090] Referring again to FIG. 8 , in act 804, process 800 augments a first LLM prompt with a first metadata identifier. For example, the system augments a first LLM prompt (e.g., a structured LLM prompt) with the first metadata identifier to be provided to an LLM. The LLM may be configured to generate a first intermediate output indicating a second set of metadata identifiers corresponding to the first metadata identifier. For example, the LLM may be communicatively coupled to (i) the raw data component and (ii) the domain ontology component, and the LLM is configured to generate the first intermediate output indicating the second set of metadata identifiers corresponding to the first metadata identifier without accessing the domain ontology component.
[0091] 10A , the LLM 1006 is communicatively coupled to both the raw data component 1010 and the domain ontology component 1008. While the LLM is communicatively coupled to each of the raw data component 1010 and the domain ontology component 1008, the LLM can communicate with the raw data component 1010 to generate a first intermediate output. For example, the first intermediate output may be an intermediate output such that the first intermediate output is not a final output of the LLM. For example, consistent with the RAG framework, the system can augment a first LLM prompt with a first metadata identifier, which will be provided to the LLM 1006 to generate a second set of metadata identifiers corresponding to the first metadata identifiers.
[0092] 9 , for example, prompt 1 902 may reflect a first LLM prompt. The system may extend (e.g., add, update, place, etc.) prompt 1 902 with a first metadata identifier (e.g., first metadata identifier 904). The extended version of prompt 1 902 may then be provided by the system as input to LLM 906 for generating first intermediate output 908. For example, LLM 906 may be the same as or similar to LLM 1006 ( FIG. 10A ) according to some implementations of the present technology. LLM 906 processes prompt 1 902 to generate first intermediate output 908. The first intermediate output 908 may be a set of metadata identifiers corresponding to first metadata identifier 904. For example, to determine what a metadata identifier (e.g., first metadata identifier 904) means, what it could be, or what it resembles, the system can provide the first metadata identifier to an LLM, which will receive a generated set of metadata identifiers, descriptions, descriptions, or other information corresponding to or otherwise associated with the first metadata identifier. In some implementations, the first intermediate output may be the same as or similar to semantically similar phrases, as described in act 404 of process 400 (FIG. 4). By doing so, the system can generate a set of semantically similar phrases that correspond to the first metadata identifier, thereby expanding the range of contextual information that the LLM will consider to later generate more accurate normalized domain-specific metadata identifiers that are key to unique entities.
[0093] Referring again to FIG. 8 , in act 806, process 800 augments the first LLM prompt with a set of metadata identifiers corresponding to the first metadata identifier. For example, the system augments the first LLM prompt with a second set of metadata identifiers corresponding to the first metadata identifier to be provided to the LLM. The LLM can be configured to generate a second intermediate output indicating the filtered domain-specific metadata identifiers by accessing the set of domain-specific ontologies. For example, referring to FIG. 9A , the system augments prompt 2 910 with the first intermediate output 908 corresponding to the first metadata identifier 904 to be provided as input to the LLM 906. In some implementations, prompt 2 910 may be the same as or similar to prompt 1 902, while in other embodiments, prompt 2 910 may be different from prompt 1 902. For example, prompt 3 914 may represent a single prompt that combines information from prompt 1 902 and prompt 2 910 into a single, updatable prompt. That is, as opposed to having two separate prompts to achieve a given goal, the system can expand the prompt multiple times upon receiving each output from the LLM 906.
[0094] For example, referring again to prompt 2 910, prompt 2 910 may include second prompt text 907 and first intermediate output 908. The second prompt text 907 may be structured text associated with level 903 or first intermediate output 908. For example, second prompt text 907 may be unique to level 903. For example, second prompt text 907, shown to indicate "return a domain-specific identifier for," may be text corresponding to level 903, where level 903 indicates a file-level metadata identifier type and first metadata identifier 904 indicates a file-level metadata identifier. In some implementations, second prompt text 907 may differ based on the metadata identifier of first metadata identifier 904. For example, if the first metadata identifier 904 is a container-level metadata identifier, the second prompt text 907 may alternatively recite "return domain-specific for the container-level identifier of," where the level 903 indicates the second prompt text 907. That is, upon determining the metadata identifier type of the first metadata identifier 904, prompt 2 910 may be selected, where prompt 2 910 is associated with the level 903 indicating the metadata identifier type, and where the first metadata identifier 904 further includes the correct second prompt text 907. In this manner, the LLM prompt may be structured based on the data schema, data format, data type, or other characteristics of the first metadata identifier to obtain more accurate results from the LLM, as opposed to the generic LLM prompts of existing systems that do not rely on a specifically generated LLM prompt.
[0095] Additionally or alternatively, second prompt text 907 may be associated with first intermediate output 908. For example, second prompt text 907 may be expanded into prompt 2 910 when the system receives first intermediate output 908. For example, the system may change, update, expand, or otherwise alter prompt 1 902 to reflect prompt 2 910 (e.g., including second prompt text 907 and first intermediate output 908). Prompt 3 914 is shown to provide an illustrative example. Prompt 3 914 can be a combined prompt (e.g., of prompt 1 902 and prompt 2 910). In some implementations, prompt 3 914 may be a resulting prompt. For example, when the system originally selected prompt 1 902 to be provided to the LLM to generate the first intermediate output 908, the system may only provide the information of prompt 1 902 to the LLM to generate the first intermediate output 908. When the system receives the first intermediate output from the LLM, the system can expand the original prompt (e.g., Prompt 1 902) to generate Prompt 3 914, which includes second prompt text 907 and the first intermediate output 908. In some implementations, the system can provide Prompt 3 914 in its entirety to the LLM 906 to generate the second intermediate output 912, which indicates the filtered domain-specific metadata identifiers, by accessing a set of domain-specific ontologies. In yet other implementations, the system can provide only the new information in Prompt 3 914 to the LLM 906 to generate the second intermediate output 912. For example, the system can only provide the second prompt text 907 and the first intermediate output 908 to the LLM 906 to generate the second intermediate output 912, thereby reducing the amount of computational resources required by the LLM to process input data (e.g., prompt information).
[0096] The second intermediate output 912 may be a filtered domain-specific metadata identifier. For example, to reduce data search time when accessing data stored in various disparate data silos through non-technical users, it is necessary to maintain domain-specific information in the context of the entity's system, allowing users to quickly search for the data they need without the burden of knowing the correct nomenclature of the data. For example, a non-technical user may try to find the name of an account. Yet, because computer engineers, data scientists, and other more technically savvy users are the ones who set up, create, or otherwise maintain the data silos, values, identifiers, phrases, or other data markers may differ from those familiar to the non-technical user. Non-technical users may be business-minded and understand the domain-specific language in which entities officially operate, but computer engineers and data scientists often do not, leading them to label data without considering the business's (e.g., the entity's) domain-specific context. To overcome this, the system can provide prompt 2 910 (or alternatively prompt 3 914) to an LLM communicatively coupled to the domain ontology component 1008 to generate a second intermediate output 912 (e.g., a filtered domain-specific metadata identifier).
[0097] For example, referring to FIG. 10 , the LLM 1006 can be provided as input with an expanded LLM prompt indicating a second set of metadata identifiers (e.g., the first intermediate output) to generate a second intermediate output indicating filtered domain-specific metadata identifiers by accessing the domain ontology component 1008. The LLM 1006 can extract one or more domain ontologies, including a set of domain ontologies specific to entities of the system. For example, as described above, if the entity is a company, the domain ontology can reflect domain-specific knowledge (e.g., nomenclature, taxonomy, lexicography) about terms used in the entity's domain. For example, if the entity is a bank, the domain ontology component 1008 can include an ontology that relates financial terms to other financial terms to infer the context in which a given financial term is used. The LLM 1006 can verify the first intermediate output (e.g., the second set of metadata identifiers corresponding to the first metadata identifiers) using one or more domain-specific ontologies included in the domain ontology component 1008. For example, the LLM 1006 may compare each metadata identifier in the second set of metadata identifiers to keywords, phrases, strings, or other domain-specific values in the domain-specific ontology to determine (i) the meaning of each metadata identifier in the second set of metadata identifiers, or (ii) the filtered domain-specific metadata identifiers.
[0098] Because the LLM 1006 can be an unsupervised artificial intelligence model, the LLM 1006 can be trained to determine the meaning of each metadata identifier in the second set of metadata identifiers by accessing a domain-specific ontology. The domain-specific ontology may be a pre-defined ontology created by one or more subject matter experts for a given entity. The LLM 1006 can determine the filtered domain-specific metadata identifiers by accessing the domain-specific ontology. For example, during the comparison process (e.g., the LLM compares or otherwise processes the first intermediate output), the LLM 1006 can determine that the first intermediate output (e.g., the second set of metadata identifiers corresponding to the first metadata identifiers) corresponds to (e.g., is associated with, matches, etc.) common filtered domain-specific metadata identifiers present in the domain-specific ontology. For example, the domain-specific metadata identifiers are considered “filtered” because they are filtered to a single representative domain-specific metadata identifier that corresponds to a potential match to the first intermediate output as generated via the LLM 1006. In this way, the system can reduce the amount of computational resources involved in generating the metadata graph, as the filtered domain-specific metadata identifiers are used to generate the metadata graph.
[0099] Referring to FIG. 10C , which illustrates a subsystem diagram of the domain ontology component 1008, the domain ontology component 1008 may be communicatively coupled to a domain thesaurus 1024 and concepts 1022. For example, the domain ontology component 1008 may host a set of domain-specific ontologies generated at least in part based on subject matter experts, while the domain thesaurus 1024 and concepts 1022 may contribute to the generation of the domain ontology component 1008 of the domain-specific ontologies. The concepts 1022 may include an entity-specific set of terms, phrases, or other concepts commonly used throughout a system of entities (e.g., FIG. 3 ). The thesaurus 1024 may include a data structure that maps the entity-specific set of terms, phrases, or concepts to other terms, phrases, or other concepts used throughout the system of entities. For example, the thesaurus 1024 may represent a digital thesaurus of terms, phrases, or concepts. The ontology component 1008, in some implementations, can aggregate information stored in thesaurus 1024 and concepts 1022 to automatically generate one or more domain-specific ontologies (e.g., via one or more ontology creation models). Additionally or alternatively, the ontology component 1008 can leverage an SME to create a set of domain-specific ontologies. In this way, the system can maintain the accuracy of the domain-specific context on which metadata identifiers rely when generating the metadata graph, thereby maintaining the system's domain-specific language of entities.
[0100] In Act 808, process 800 generates a metadata graph. For example, the system may generate a domain-specific unified metadata graph via an LLM using (i) the first metadata identifier and (ii) the second intermediate output indicating the filtered domain-specific metadata identifier. The LLM may be configured to generate a graph (e.g., an undirected graph, a directed graph, a directed acyclic graph, etc.) using the first metadata identifier and the filtered domain-specific metadata identifier. In some implementations, the LLM may be provided with a prompt instructing it to generate a graph (e.g., a metadata graph), the prompt including the first metadata identifier, the filtered domain-specific metadata identifier, and prompt text instructing it to generate the metadata graph. In some implementations, Acts 802-808 may be repeated iteratively until all metadata in data silo 1015 (FIG. 10B) has been processed by the system.
[0101] 9B , prompt 4 916 may include third prompt text 918 that instructs generating a graph 920, a first metadata identifier 904, and a second intermediate output 912. The third prompt text 918 may be structured text associated with the level 903 or the second intermediate output 912. For example, the third prompt text 918 may be unique to the level 903, the first metadata identifier 904, and the second intermediate output 912, and the LLM 906 is for generating a domain-specific integrated metadata graph based on the level 903, the first metadata identifier 904, or the second intermediate output 912. The LLM may generate the metadata graph 920 using at least a portion of the information included in prompt 4 916. For example, the LLM may be trained to generate a graph data structure including first metadata identifiers, filtered domain-specific metadata identifiers, or other information (e.g., file-level metadata identifiers, container-level metadata identifiers, system-level metadata identifiers, location identifiers, data lineage, etc.) as described in act 406 of process 400 (FIG. 4) and in FIGS. 5-6.
[0102] In some implementations, the LLM can be provided with prompt 5 922, which can represent a single prompt that combines the information from prompt 1 902, prompt 2 910, and prompt 4 916 into a single, updatable prompt. That is, as opposed to having three separate prompts to achieve a given goal, the system can expand the prompt multiple times in terms of receiving respective outputs from the LLM 906. When the LLM 906 is provided with a prompt (e.g., prompt 4 916 or prompt 5 922) as input, the LLM 906 can generate a metadata graph 920.
[0103] Referring to FIG. 11 , which shows an illustrative representation of a generated metadata graph, the LLM 906 can generate a metadata graph 1100. The metadata graph 1100 may be the same as or similar to the metadata graph 920 ( FIG. 9B ), the metadata graph 1004 ( FIG. 10A ), the metadata graph 500 ( FIG. 5 ), or the metadata graph 600 ( FIG. 6 ). The metadata graph 1100 can represent a domain-specific unified metadata graph in accordance with some implementations of the present technology. The metadata graph 1100 can include nodes 1102a-d, such as a fifth node 1102a, a sixth node 1102b, a seventh node 1102c, and an eighth node 1102d. Each node 1102a-d can indicate metadata for one or more data objects stored within a given data silo. For example, the fifth node 1102a may include a file-level metadata identifier 1106a, a container-level metadata identifier 1108a, a location identifier 1110a, and a domain-specific metadata identifier 1112a. Additionally or alternatively, the fifth node 1102a may include a system-level metadata identifier or other information, not shown. The domain-specific metadata identifier 1112a may be the same as or similar to the second intermediate output 912 indicating the filtered domain-specific metadata identifier.
[0104] Each of the nodes 1102a-1102d may be linked to one or more other nodes. For example, the fifth node 1102a may be linked to the sixth node 1102b via a second edge 1104a. In some implementations, the second edge 1104a may indicate lineage information of the node, such as when the sixth node 1102b is a data source for the fifth node 1102a. Yet, in other implementations, the second edge 1104a may indicate lineage information, such as when the fifth node 1102a is a data source for the sixth node 1102b, according to some implementations of the present technology. It should be recognized by those skilled in the art that each node 1102 may be linked to other nodes via edges 1104, each edge indicating lineage information between one or more nodes in the set of nodes.
[0105] In some implementations, one or more of the identifiers included in nodes 1102a-1102d are traversable. To efficiently traverse the metadata graph 1100, the system can traverse the metadata graph based on a single traversable identifier while ignoring other identifiers included in the node. For example, the traversable identifier can be the filtered domain-specific metadata identifier 1112a. As referred to herein, a traversable identifier is an identifier that the system looks for when a user provides a query to locate data, while a non-traversable identifier is an identifier that the system associates with and stores with node 1102a and does not look for when traversing the metadata graph 1100. In this way, the system reduces the amount of computational resources traditionally utilized when searching large tables for strings because the metadata graph (i) is a graph that provides direction (e.g., a directed graph) and (ii) uses entity-specific, domain-specific, and contextually accurate metadata identifiers to locate the same instance of a data object stored throughout the entity system. For example, as the system traverses the metadata graph, the system may compare a set of phrases (e.g., as described in act 406 of process 400 (FIG. 4)) with the traversable metadata identifiers of the metadata graph 1100. While the system may traverse the metadata graph 1100 based on the filtered domain-specific metadata identifiers, the metadata graph 1100 may still store other metadata identifiers (which may be, for example, file-level metadata identifiers 1106a, container-level metadata identifiers 1108a, or system-level metadata identifiers) in association with the nodes 1102a-1102d to maintain the information for future use.For example, when the metadata graph locates a given data object using the filtered domain-specific metadata identifier 1112a, the system can then retrieve file-level, container-level, or system-level location information or other information about the given data object.
[0106] 8, in some implementations, the system may perform a validation process on the generated metadata graph 1100 (FIG. 11). For example, in some implementations, the system calculates query-to-results performance metrics and accuracy performance metrics. For example, the validation process may include providing automatically generated or user-provided test queries to the domain-specific integrated metadata graph to measure the performance of the domain-specific integrated metadata graph.
[0107] For example, referring to FIG. 10D illustrating a subsystem diagram of the feedback component 1012, the feedback component 1012 can include a versioning component 1026, a generator 1028, a result 1030, and an update component 1034. The versioning component 1026 can store a previous version of the metadata graph 1100 ( FIG. 11 ). For example, the versioning component can store the most recent version of the metadata graph 1100 before generating an updated version of the metadata graph 1100 ( FIG. 11 ). The generator 1028 can generate a test query for providing the metadata graph 1100. For example, the test query may be the same as or similar to a user-specified query, as discussed in act 402 of process 400 ( FIG. 4 ). Additionally or alternatively, the test query may be a historical user-specified query, as discussed in act 402 of process 400 ( FIG. 4 ). The test query can be utilized to generate one or more results. For example, results 1030 may determine (e.g., generate) one or more performance metrics based on test queries as generated by generator 1028. For example, results 1030 may store historical performance metrics for other versions of metadata graph 1100 and performance metrics for the current version of metadata graph 1100 (FIG. 11). The performance metric may be a query-to-results performance metric, an accuracy metric, or other performance metric.
[0108] Once performance metrics for the metadata graph 1100 are generated, a decision can be made to update the metadata graph 1100 ( FIG. 11 ). The update component 1034 can automatically trigger the update process to the metadata graph 1100 ( FIG. 11 ). In some implementations, the update component 1034 can still trigger the update process to the metadata graph 1100 in conjunction with third-party input 1032. For example, the third-party input 1032 can be a third-party source of information (e.g., a website, a computing device, etc.). As another example, the third-party input 1032 can be a subject matter expert input. In this manner, by leveraging subject matter expert input, a human can verify the accuracy of an LLM-generated metadata graph before publishing such a metadata graph for use across systems. Adding the opinion of a subject matter expert to the generation process of the metadata graph 1100 can enhance the accuracy with which the metadata graph is generated to avoid any unintentional LLM-related errors.
[0109] To perform the verification process, the system can provide a first query (e.g., a test query) requesting the location of a first data item to each of (i) the domain-specific unified metadata graph and (ii) other versions of the domain-specific unified metadata graph. As an example, the system can test the most recent iteration of the domain-specific unified metadata graph to find the location of a given data item (e.g., stored in a data silo). Nevertheless, to ensure that the most recent modification to the domain-specific unified metadata graph results in a better metadata graph, the system compares performance metrics of the domain-specific metadata graph with previous (or other) versions of the metadata graph. For example, the system can calculate a query-to-results performance metric indicating the time period between when a query is provided to each metadata graph and when results are received from each metadata graph. The time period may be in epoch time, seconds, milliseconds, determinations, UNIX time measured in microseconds, etc. Such query-to-results performance metrics can be generated for each of the domain-specific unified metadata graph and other versions of the domain-specific unified metadata graph (e.g., previous versions of the unified metadata graph).
[0110] In some implementations, the system can calculate accuracy metrics (e.g., performance metrics) of the domain-specific unified metadata graph and other versions of the domain-specific unified metadata graph. For example, the accuracy metric may be a degree of accuracy (e.g., a percentage, a decimal value, a ratio, an integer, a binary value, a numeric value, an alphanumeric value, etc.) of the results generated from the domain-specific unified metadata graph and other versions of the domain-specific metadata graph. The accuracy metric may be generated based on human evaluation of the results (e.g., results returned from submitting a query to each metadata graph). For example, a subject matter expert (e.g., a data scientist, software developer, computer engineer) can verify the accuracy of the results for each of the domain-specific metadata graph and other versions of the domain-specific metadata graph. Because each of the metadata graphs integrates the “domain,” “context,” and “terminology” of a given entity's system, the subject matter expert can verify the accuracy of the generated results returned by each metadata graph in locating a given data item. In this way, experts can verify the accuracy of the results, which can lead to more accurate generation of domain-specific integrated metadata graphs. Nevertheless, in other implementations, the accuracy metrics may be generated automatically without human intervention. For example, the accuracy metrics can be based on a comparison of results generated from each metadata graph by one or more implementations of the present technology to historical results.
[0111] The system can calculate the accuracy metric by sampling and examining one or more portions of the results, or all of the results. For example, the system can select a sample set of results to determine the accuracy metric, and can vary the size of the sample set until the desired accuracy metric threshold is met.
[0112] In some implementations, the system can determine whether to perform an update process on the metadata graph. For example, the system can determine whether a performance metric of the metadata graph (e.g., metadata graph 1100) satisfies a performance measure related to a second performance metric of another version (e.g., a previous version) of the metadata graph. In some implementations, determining whether a performance metric of the metadata graph satisfies the performance measure can be based on whether (i) the query-to-result performance metric of the metadata graph fails to exceed the query-to-result performance metric of the other version of the metadata graph, or (ii) the result accuracy metric of the metadata graph meets or exceeds the result accuracy metric of the other version of the metadata graph. In this manner, when the performance measure fails to be satisfied, the system can perform an update process on the metadata graph. In short, if the metadata graph (i) returns results faster than the previous version of the metadata graph and (ii) returns results more accurately than the previous version of the metadata graph, the metadata graph should not be updated. Nevertheless, if the metadata graph (i) returns results slower than the previous version of the metadata graph, or (ii) returns results that are less accurate than the previous version of the metadata graph, the metadata graph will be updated.
[0113] 8, in act 810, process 800 performs an update process on the metadata graph. For example, in response to determining that a first performance metric of the domain-specific unified metadata graph fails to meet a performance measure for a second performance metric of another version of the domain-specific unified metadata graph, the system may perform an update process on the domain-specific unified metadata graph (e.g., metadata graph 1100 (FIG. 11)). As described above, a metadata graph will be updated if it (i) returns results slower than a previous version of the metadata graph or (ii) returns results that are less accurate than a previous version of the metadata graph.
[0114] In some implementations, the update process can be performed by updating the nodes and edges of the metadata graph to the nodes and edges of the previous version of the metadata graph. For example, the system can determine a set of inconsistencies between the domain-specific unified metadata graph and the other version of the domain-specific unified metadata graph. The set of inconsistencies can reflect inconsistencies between (i) the nodes of the domain-specific unified metadata graph and the other version of the domain-specific unified metadata graph, and (ii) the edges connected to at least one node of the domain-specific unified metadata graph and the other version of the domain-specific unified metadata graph. For example, the system can traverse each graph to determine newly added nodes, edges, metadata identifiers, or other information. For example, the system can first traverse the current version of the domain-specific unified metadata graph and store a tabular representation of the current version of the domain-specific unified metadata graph in a database. The system can then retrieve tabular versions of other versions (e.g., previous versions) of the domain-specific unified metadata graph, if available. In some implementations, the system can traverse the other version of the domain-specific unified metadata graph and store a tabular representation of the other version of the domain-specific unified metadata graph in a database. The system can then compare the two tabular versions of the metadata graph to each other to identify inconsistencies between the two versions. For example, the system can identify one newly added node (e.g., and the metadata or location identifier that the node contains) and two newly added edges that connect this node to other nodes in the current version of the metadata graph (as compared to the previous version of the metadata graph). The system can then update the domain-specific unified metadata graph with updated nodes and edges of the other version of the domain-specific unified metadata graph that correspond to the set of inconsistencies.In this way, the system can revert to a previous version of the unified metadata graph when performance metrics fail to be met.
[0115] In some implementations, the system can effect an update process based on detecting the addition of a data silo. For example, in some implementations, the system can perform an update process on the domain-specific unified metadata graph when a data silo is added to a computing environment associated with an entity. The system can monitor a computing environment associated with the entity (e.g., FIG. 3) using one or more network discovery tools, such as SNMP, LLDP, CDP, or others, to identify when a new device is added to the entity's computing environment. The system can then communicate with the new device to determine the type of the new device using the IP address determined via the network discovery tool to verify the addition of the new device (e.g., using a “ping” command). The system can further compare the determined IP address of the new device with a database that stores device information for the computing environment. For example, a table can store device IP addresses and device types for the entity's computing system. The system can use the table to determine whether the newly added device is a data silo (e.g., a data source, a database, etc.). In response to detecting the addition of a data silo, the system can cause an update process to be performed on the domain-specific unified metadata graph. For example, the system can extract metadata from newly added data silos to update or regenerate the domain-specific unified metadata graph 1100 (FIG. 11). For example, the system can iteratively repeat one or more of the processes described above until all metadata from the data silos (and newly added data silos) has been processed by the system, thereby generating a system-wide domain-specific unified metadata graph.
[0116] Artificial Intelligence Sandbox The above-described domain-specific integrated metadata graph enables computer systems to efficiently extract siloed data across disparate locations. Artificial intelligence (AI) sandboxes enable the use of data from these disparate locations to generate AI models or automatically apply existing AI models for data analysis. Sandboxes provide a low-code or no-code environment, allowing users to build and apply models to extract insights or predictions from data even when they lack the expertise or time to build AI models and data pipelines for the data to be processed through the models, or when they are unaware that a model would help solve a problem. As described herein, AI sandboxes include automated tools for deploying and managing models that ensure the models operate efficiently without requiring human intervention. AI sandboxes can be used to automate any or all steps in a model's lifecycle, such as training new models, fine-tuning or improving existing models, deploying models, or generating a pipeline of data suitable for analysis by the model once deployed. By integrating these capabilities, Sandbox provides a practical solution for transforming user interactions with data and AI models, enhancing accessibility and usability while maintaining operational efficiency.
[0117] 12 is a block diagram illustrating components in an AI sandbox 1200, according to some implementations. As shown in FIG. 12, AI sandbox 1200 can include a model review assistant 1210, a data processor 1220, a model generator 1230, a model governor 1240, and a model automator 1250. Other implementations of AI sandbox 1200 can include additional, fewer, or different components, or can divide functionality differently among the components.
[0118] The model review assistant 1210 interacts with users and orchestrates other components of the AI sandbox 1200 to generate and apply AI models. The model review assistant 1210 can leverage large-scale language models (LLMs) to both interact with users and perform tasks related to generating, applying, or improving AI models as the user interacts with the assistant 1210. In some implementations, the model review assistant 1210 generates a chat-style interface where user input is received and information is output to the user.
[0119] The data processor 1220 identifies relevant data objects applicable to the model the user requests to train or apply. The data processor 1220 processes the data to ensure cleanliness and integrity, making the data suitable for analysis by the AI model. The data can be prepared for various stages of AI development, including generating, training, and testing datasets to train, fine-tune, or evaluate models. Additionally, the data processor 1220 can provide archiving capabilities to securely store clean data and generated training samples. Using this archived data, the data processor 1220 can generate documentation that facilitates inspection of the model and its use, ensuring transparency and compliance with regulatory standards. The data processor 1220 is further described with respect to FIG. 13.
[0120] The model generator 1230 builds an AI model based on user input and data output by the data processor 1220. The model generator 1230 is configured to train a machine learning model using the cleaned and processed dataset from the data processor 1220. The model generator 1230 can analyze the dataset output by the processor 1220, for example, by creating a set of features from the dataset and identifying a target variable for the model. The model generator 1230 can also recommend a type of machine learning model to build from a list of available models or based on model performance metrics. Using the processed dataset and the selected model type, the model generator 1230 builds a model, which may include training a new model, fine-tuning an existing model, or preparing a model for deployment. The model generator 1230 is further described with respect to FIG. 14.
[0121] The model governor 1240 performs higher-level analysis and governance assessment tasks to ensure that the model complies with a set of governance measures. These governance measures can relate to policies or procedures within the organization in which the model review assistant 1210 operates, or policies from external organizations that the organization must follow (such as policies or regulations implemented by a government or standard-setting body, or social commitments signed by the organization). If the model does not comply with the governance measures, the model governor 1240 can cause the model to be modified until it is in compliance. The model governor 1240 can further generate documentation for the model, which can be archived and stored for subsequent analysis and inspection. The model governor 1240 is further described with respect to FIG. 15.
[0122] Model Automator 1250 manages the deployment of AI models. Models can have different requirements for their deployment. For example, some models are large and require significant amounts of computing resources. Some models are used for applications where results are needed quickly. Entities that use AI models can have access to a variety of different locations for deploying the models, such as one or more cloud providers or on-premises machines or elastic resources. Model Automator 1250 can orchestrate the deployment of models across these various different locations. Model Automator 1250 can provide a central system that can access any applicable APIs for deploying and invoking models, the data used to train these models, and the application data processed by the models. Thus, Model Automator 1250 can determine how to efficiently deploy models to enable them to successfully implement and manage computing resources. Model Automator is further described with respect to FIG. 18 .
[0123] 13 is a block diagram illustrating components within data processor 1220, according to some implementations. Data processor 1220 can generally prepare data for various purposes associated with an AI model, including generating training data for training or fine-tuning a model, generating test data for testing a model, or preparing data for application to a model.
[0124] When a request for data (e.g., to train, fine-tune, or test a model, or to apply to a model) is received, the data processor 1220 accesses a dataset 1310 that has been granted permissions for the entity requesting access. The dataset 1310 can be accessed using any of the techniques described above, including using a metadata graph to identify data objects within data silos in a data repository.
[0125] The data processor 1220 may include a data preprocessor 1320 that performs various operations on the accessed dataset 1310 to prepare the data for desired use. The data preprocessor may perform operations to generate training or test datasets for training, fine-tuning, and / or testing a model. Such operations may include, for example, normalizing the data, converting the data from one type to another, converting the data to an appropriate format or structure, or aggregating or reducing the data. In some implementations, the data preprocessor 1320 may apply anonymization operations that remove, encrypt, or obscure personally identifiable or private information in the accessed dataset 1310. Some implementations of the data preprocessor 1320 also apply operations to make the data compliant with policies or regulations. For example, the data preprocessor 1320 may remove data from the dataset 1310 that is determined to be inaccurate. An example process for detecting and removing noise from a dataset is described in U.S. Patent Application No. 18 / 736,407, filed June 6, 2024, which is incorporated herein by reference in its entirety.
[0126] The processed data output by the data preprocessor 1320 can be passed to other elements of the data processor 1220 to generate a set of data for training or testing an AI model. The data preprocessor 1320 can additionally or alternatively generate a set of application data 1322 for application to an existing AI model. The data preprocessor 1320 can perform cleansing, normalization, conversion, or other processing operations on the accessed dataset 1310 to generate the application data. These operations can include anonymizing the data as described above, or, in another example, applying a preprocessing model to the dataset 1310 that counteracts distortions in the model to which the application dataset 1322 will be applied. The preprocessing model can include one or more data modification operators that modify the raw dataset 1310 by, for example, adding features, removing features, or changing values in the dataset 1310 to create a set of modified data. When applied to an AI model, the modified data items from the modified data set cause the model to produce outputs that do not exhibit distortions in the model. An example process for detecting and counteracting distortion in a model via a pre-processing model is described in U.S. Patent Application No. 18 / 783,409, filed July 25, 2024, which is incorporated herein by reference in its entirety.
[0127] The data analyzer 1330 analyzes the data to understand the data types and how the data will be used and generates a usable set of data (e.g., training data, test data, or inference data set). The data analyzer 1330 receives the raw data and / or the preprocessed data set output by the preprocessor 1320 and evaluates properties such as the data type, the range of the data, linkages or dependencies between the data, or other information that helps the system understand what the data is, how the data is used, and how the data can be used in a model. The data analyzer 1330 can further generate synthetic data to supplement the existing data based on the data analyzer's 1330 analysis of the properties of the data. In some implementations, the data analyzer 1330 includes a data profile generator 1340 and / or a sampler 1350, which are described below.
[0128] The data profile generator 1340 generates a data profile 1342 for the dataset output by the data preprocessor 1320. The data profile generator 1340 can analyze the structure of the dataset and the interrelationships between data within the set to ensure the dataset is appropriate for the intended artificial intelligence application. Such analysis can include, for example, calculating statistics of the dataset (e.g., mean or standard deviation), checking and, if necessary, correcting the data types within the dataset, checking for data anomalies (e.g., missing values, duplicates, or outliers), detecting patterns or correlations between data items within the dataset, or detecting asymmetries or imbalances in the distribution of data within the dataset that may affect model performance. The data profile 1342 output by the data profile generator 1340 can further include metadata describing the dataset, which can be used to generate model documentation or perform inspections and compliance checks.
[0129] In some implementations, the data processor 1220 generates a set of training data that can be used to train a machine learning model. Accordingly, the data processor 1220 may further include a sampler 1350 that samples the datasets output by the data preprocessor 1320 and the data analyzer 1330 to generate training samples 1352 or test samples 1354. For example, the sampler 1350 may generate a representative random sample of items of data from the processed datasets for each of the training dataset and the test dataset. The training samples 1352 can be further input to a synthesis fabricator to generate additional training or test cases.
[0130] 14 is a block diagram illustrating components within the model generator 1230, according to some implementations. As described above, the model generator 1230 trains a machine learning model using the cleaned and / or processed dataset output by the data processor 1220. The model training procedure can be performed based in part on input 1405 received from a user (e.g., via the model review assistant 1210).
[0131] 14, the model generator 1230 may include a feature set generator 1410 that collects the data profile 1342 and training samples 1352 generated by the data processor 1220. The feature set generator 1410 creates a set of features from the dataset that can be used to train a machine learning model. When creating the features, the feature set generator 1410 may convert the data (e.g., in the training samples 1352) into an appropriate machine-readable format, for example, by converting non-numeric values to numeric values, converting numeric formats to the same type, vectorizing the data, or the like.
[0132] The features output by feature set generator 1410 can be defined based in part on user input. For example, when a user is creating a model using AI sandbox 1200, the user can directly specify features of interest for the model or can provide information that feature set generator 1410 uses to identify features in a dataset.
[0133] In some implementations, the feature set generator 1410 further uses feedback to learn over time how to select relevant features for a given application. For example, as a user interacts with the AI sandbox 1200 to generate models, the feature set generator 1410 can evaluate features used in each model, such as the data types input to the model, the model's target variable, or the model's performance, to build a robust mapping between the model's features and attributes. The feature set generator 1410 can then use this mapping to recommend features for other models or train a feature selection model based on the mapping. In another example, the feature set generator 1410 uses model performance feedback to recommend or select features for a given model. The feature set generator 1410 can, for example, select a first set of features and receive feedback indicating the performance (e.g., accuracy) of a model trained using the first set of features. The generator 1410 can then select a second set of features, receive feedback indicating the performance of the model trained using the second set of features, and compare the performance measures to determine whether the first or second set of features produced better results.
[0134] The goal selector 1420 selects a target variable for the machine learning model based on the features output by the feature set generator 1410 and / or based on user instructions received via the model review assistant 1210. The target variable specifies the output that the machine learning model will predict or classify. To identify the target variable, the goal selector 1420 can receive variable identification from the model review assistant 1210 generated based on the user input 1405. For example, the model review assistant 1210 can use an LLM to identify a user-specified target variable in a natural language user input. The goal selector 1420 can compare the variable identified in the user's input with the features output by the feature set generator 1410 to determine whether the user-specified variable is present in the set of features or can be derived from the set of features. In some cases, the goal selector 1420 can use an LLM to evaluate the set of features and identify the closest match to the user-specified target variable. The goal selector 1420 can additionally or alternatively use patterns of user behavior to identify the target variable. For example, if a user has recently used AI sandbox 1200 to generate a model based on a particular target variable in a corresponding dataset, goal selector 1420 can determine that the user may be interested in generating a model based on the same target variable in a different dataset. Similarly, goal selector 1420 detects similarities between the user's actions within the enterprise (such as other models the user has built, models the user has used, or data the user has created or accessed) and the actions of other users who have built models using AI sandbox 1200. Based on these similarities, goal selector 1420 can determine that the user is likely to build a model for a specified target variable because other users have built models for the same specified target variable.
[0135] The goal selector 1420 can additionally or alternatively use feedback from a user or other system when selecting a goal variable. This feedback can be used in place of the process for selecting a goal variable from a user's natural language input or based on the user's past activity described above, or can be used to improve the selector 1420's ability to identify the correct goal variable based on natural language input or user activity. In one example, after a model is initially trained for one goal variable, feedback from a user or external system can be used to determine that the model should be trained for a different goal variable (e.g., because the user provides direct input indicating that the trained model is not producing the desired output, or because another system identifies an error in the trained model's output). In other cases, the goal selector 1420 generates a mapping between goal variables used in other models and attributes of the models (such as features input into the other models, data types used in the other models, users who created the other models, or performance of the other models) to predict goal variables that are likely to be relevant to the developed model.
[0136] The model selector 1430 evaluates whether a model should be generated or deployed and, if so, recommends the type of machine learning model the model generator 1230 should build. The model type can be one of a set of different model architectures, such as a neural network, a random forest, or a support vector machine. Additionally or alternatively, the model type can be selected from a set of commercially available or existing models that can be used as is or fine-tuned for a specific purpose. Similarly, the model selector 1430 can select from pre-trained models, which can be models previously developed within the organization where the model review assistant 1210 operates or received from an external source. For each application, the model selector 1430 can recommend an individual model or a set of multiple models to achieve the user's desired objective. For example, the model selector 1430 can recommend generating a series of models that can be used together in an ensemble method. When recommending multiple models, the model selector 1430 can recommend generating all new models using a selected set of pre-trained or fine-tuned models, or combining new models with pre-trained or fine-tuned models. The model selector 1430 can also recommend specific ensemble learning techniques that allow these models to be used together. For example, the model selector 1430 can recommend generating a series of three models in a voting ensemble, where the predictions for each model are combined by majority voting or averaging. In another example, the model selector 1430 recommends a stacking ensemble, in which the outputs of several base models are used as inputs to a metalarner model that makes the final prediction, or a boosting ensemble, in which models are trained sequentially, each focusing on correcting the errors of its predecessors to improve overall performance.
[0137] When selecting a model, the model selector 1430 can receive an explicit model selection from the user. For example, the model review assistant 1210 can provide the user with a list of available types of models, and the user can select a model from the list. Alternatively, the model selector 1430 can recommend a model type for a given application. To recommend a model, the model selector 1430 can apply a set of input data to each of multiple types of models and calculate metrics for each model type. The metrics can include, for example, a measure of how quickly the model generates output for a set of input data (e.g., latency, model output speed, or model response time variance), a measure of how accurate the model's output is, the amount of memory used by the model, or a measure of the cost of using the model. Based on the metrics, the model selector 1430 can recommend one or more model types that achieve a particular goal, such as the fastest execution or the most accurate results. In other cases, rather than applying the input data to multiple models and calculating metrics used for model selection, the model selector 1430 can recommend a model type based on historical model performance. For example, the model selector 1430 may evaluate historical data that indicates that one type of model is typically faster but less accurate than another type of model, allowing the model selector 1430 to make a recommendation of a model type based on whether speed or accuracy is more important for a given application. Some implementations of the model selector 1430 output an identifier of the recommended model to the model builder 1440 to enable the recommended model to be built. In other implementations, the model selector 1430 outputs the recommended model type to the user, so that the user can choose between the recommended model types or select a different type of model.
[0138] The model builder 1440 builds models 1445, which may include training a new model, fine-tuning an existing model, or packaging or refining a model for deployment without any further training. The model builder 1440 receives the selected model type from the model selector 1430 and may determine whether to train, fine-tune, package, or perform other tasks based on the selected model. When training a new model or fine-tuning an existing model, the model builder 1440 trains the model type selected by the user or the model selector 1430 based on the target variable identified by the goal selector 1420 and the set of training samples 1352 generated by the data processor 1220. The model builder 1440 may then test the model using test samples 1354 generated by the data processor 1220 and retrain as necessary, for example, until the model converges or reaches a specified accuracy threshold on the test data set.
[0139] FIG. 15 is a block diagram illustrating the functionality of the model governor 1240, according to some implementations. As shown in FIG. 15, the model governor 1240 can leverage a model review assistant 1210 to evaluate a model 1445, including using the model review assistant 1210 to interface with a large language model (LLM) 1510 to evaluate the model 1445 for higher-level analysis and governance tasks. In some implementations, the LLM 1510 uses a RAG-based process 1520 to retrieve policies, procedures, know-how, or other applicable governance information from a data repository, such as an industry knowledge repository 1522 or an enterprise knowledge repository 1524. The industry knowledge repository 1522 can store, for example, scientific models, national regulations, or data standards. The enterprise knowledge repository 1524 can store information such as organizational policies, organization-specific classifications, or the organization's business mission. The model review assistant 1210 can generate queries to the LLM 1510 that cause the LLM 1510 to evaluate the model 1445 against applicable governance information.
[0140] Based on the evaluation, the model governor 1240 can update the model 1445, create additional models or data that bring the model 1445 into compliance with the governance measures, or generate documentation describing how the model does or does not comply with the governance information. For example, an organization may be subject to regulations that dictate that a particular type of model (e.g., a model for approving loan applications) must produce outputs that are not biased toward particular backgrounds (e.g., race, gender, age, sexual orientation, or geographic location or loan applicant). If the model governor 1240 detects that such a loan approval model is incorrectly biasing its conclusions based on one or more of these backgrounds, the model governor 1240 can cause the model to be retrained or fine-tuned to reduce the biased conclusions. Alternatively, the model governor 1240 can cause a second model to be generated, configured to pre-process application data before the application data is input into the model 1445 to correct the model bias. In another example, if the model governor 1240 determines that the model complies with each of a set of governance measures, the model governor 1240 can generate documentation (optionally using the LLM 1510) that describes the governance measures against which the model 1445 was evaluated and how the model was determined to comply with each of the measures.
[0141] The model governor 1240 can further generate an explainability layer 1530 for the model 1445. The explainability layer 1530 includes data associated with the model 1445 that explains what the model is doing, how the model makes its decisions, what data was used to train the model, what data is input or output from the model, any modifications applied to the input data before the data is processed through the model, or any other characteristics specified by corporate, regulatory, or other standards for the organization where the model was generated. The model governor 1240 can generate the explainability layer using the LLM 1510 by analyzing the model 1445 itself and / or governance documents retrieved from a repository such as the industry knowledge repository 1522 or the corporate knowledge repository 1524.
[0142] Automated generation of AI models 16A is a flowchart illustrating a process 1600 for automatically generating an artificial intelligence (AI) model, according to some implementations. The process 1600 can be performed by a computer system, such as one or more systems implementing components of the artificial intelligence sandbox 1200.
[0143] As shown in FIG. 16A , the computer system receives natural language input from a user at 1602. The natural language input may include explicit instructions for the computer system to train an AI model or a general query that the computer system can process as instructions to generate a model. The input may include a set of phrases that implicitly or explicitly indicate desired properties of the model, such as the data to be processed by the model and / or the intended output from the model. In one example, a user may provide the input, “I need help analyzing my investment portfolio.” The computer system processes this input to determine whether a model would be useful in answering the query or whether the query can be answered without the model. For example, the computer system may determine that a model would be useful in analyzing an investment portfolio, while a more direct query (e.g., “Do I own any shares of XYZ Corp?”) would not be required or would be complicated by the model. When submitting a natural language input, a user may not be aware that a model should be useful, much less the type of model that should produce the best results, the data that should be used to train or input the model, or how to go about training and deploying the model.
[0144] Input can be received, for example, through model review assistant 1210, which can provide a chat-like interface through which input can be received from a user and information can be output to the user. User input to the chat interface can be received as natural language input and / or as other types of input, such as selecting an item from a list. Similarly, model review assistant 1210 can output information to the user in natural language format (e.g., using an LLM to generate output), in a graphical format such as a data plot, or in other formats. An example of a chat interface 1700 through which user input can be received is depicted in FIG. 17A. In the example, the user provided natural language input 1705, "I need help analyzing my investment portfolio." Other users can provide natural language input that more directly instructs the computer system to generate a model (e.g., "Create a model to analyze my investment portfolio"). The phrase "my investment portfolio" can be processed by the computer system as an indication of data to be processed by the model.
[0145] At 1604, the computer system accesses a metadata graph based on user input. The metadata graph may include (i) a set of nodes comprising (a) metadata indicating internal data objects stored in data silos and (b) location identifiers of the data silos, and (ii) edges indicating data lineage between the set of nodes. As described above, the computer system traverses the metadata graph, which indicates where data is stored, what data is available in different data silos, and the data lineage between the nodes of the graph, thereby enabling the computer system to efficiently find data within the set of data silos. By traversing the metadata graph, the system can determine nodes corresponding to a set of phrases in the natural language input received from the user. In the example interface illustrated in FIG. 17A , the computer system selects four candidate data sources associated with one or more nodes of the metadata graph identified based on the phrase “my investment portfolio,” each data source including one or more data items. For example, a data source for the implementation of investment A may include a set of measurements of the value of investment A at different times. In some implementations, such as in the example shown in Figure 17A, a user can provide additional input to select from among the candidate data sources shown in Figure 17B. Alternatively, the computer system can proceed to train a new model with the data objects identified based on the user's natural language input.
[0146] At 1606, the computer system processes the retrieved data objects using the metadata graph to generate a training dataset. The computer system may pre-process the data objects, such as by removing biased, incorrect, or irrelevant values from the dataset. Once the data objects are cleaned, the system may sample a training dataset from the data objects and / or input the data objects to a synthetic data generator to generate synthetic training data. Similarly, the computer system may generate a set of test data for testing the model once it has been trained.
[0147] At 1608, the computer system selects a model type for the model to be trained. The model type can be selected from among different model architectures and / or from commercial packages or existing models using model metrics associated with each available model type. In some cases, the model type can be output to the user via a chat interface that allows the user to select the type of model to train. The computer system then trains the selected type of model using the generated set of training data at 1610.
[0148] After the model is trained, whether through the process of Figure 16A or another process, the computer system can also interact with a user to automatically generate a pipeline of application data to be processed through the model. Figure 16B is a flowchart illustrating a process 1620 for automatically generating a data pipeline for an AI model according to some implementations. Process 1620 can be performed at least in part by the same computing system that performs process 1600 of Figure 16A, or can be performed by one or more different computing systems.
[0149] As shown in FIG. 16B , the computer system receives 1612 a first natural language input from a user, which may include a set of phrases and instructions for using an AI model to analyze data associated with the set of phrases. Like the input for training a model, the input for deploying a model can be received via a chat interface generated by the model review assistant 1210. For example, FIG. 17C illustrates user input 1710 received via the chat interface instructing the computer system to “take a look at last year’s investment mix.” The phrase “last year’s investment mix” can be used to identify the set of data to which the model will be applied, while “take a look” is interpreted in the context of the chat session as an instruction to deploy the model generated during the chat session.
[0150] At 1624, the computer system uses the natural language input to access a metadata graph that indicates internal data objects stored in the data silo. The metadata graph can be the same graph used to identify data objects for generating the training dataset described with respect to Figure 16A. Using the metadata graph, the system can determine nodes that correspond to a set of phrases in the first natural language input.
[0151] At 1626, the computer system processes the internal data object indicated by the determined node to generate a first set of data. For example, the computer system may remove personally identifiable or private information from the internal data object. In another example, the system applies a set of data correction operators to the internal data object to generate a set of corrected data based on the internal data object, the data correction operators being configured to remove bias from the internal data object or to remove or compensate for inaccurate or irrelevant data within the object.
[0152] At 1628, the computer system applies the AI model to the first set of application data to generate one or more outputs based on the application data. For example, the computer system uses the model to classify items of data in the first set of application data or make a prediction based on one or more of the application data items.
[0153] At 1630, the computer system sends a representation of one or more outputs to the user for display. For example, Figure 17D illustrates the computer system generating plot 1715 to illustrate the results produced when the model is applied to a set of data specified by the user (e.g., "Last Year's Investment Mix").
[0154] Based on the displayed output, the user can determine that modifications to the model or the data processed by the model are needed to obtain the desired results. For example, the displayed output can indicate that the data input to the model was incomplete or incorrect. Thus, the user can iteratively interact with the model review assistant 1210 to modify the model's inputs until the expected output is received. These iterative interactions can include further natural language input from the user via a chat interface provided by the model review assistant 1210.
[0155] For example, at 1612, the computer system receives a second natural language input including instructions to modify the first set of application data (e.g., by adding data to the first set, removing data from the first set, or modifying values within the first set). FIG. 17D further illustrates an example input 1720 in which the user instructs the computer system to “add data source 4.” The computer system then generates a second set of application data based on the second natural language input at 1614. In response to the instructions in the second input, generating the second set of application data may involve retrieving additional data objects using the metadata graph and processing the data objects to obtain application data, modifying data values within the first set of application data (e.g., by modifying the way the data objects were processed to produce the first set of application data), or removing data from the first set of application data. The AI model can then be applied to the second set of application data at 1616, and output produced from the application of the model can be displayed to the user. This iterative process can be repeated until the user is satisfied with the output of the model.
[0156] Deploying AI models Once an AI model is trained and determined to comply with the model's governance parameters, the model can be deployed to make predictions based on application data input into the model. Model automator 1250 determines how the model should be deployed for use in a production environment and orchestrates the deployment.
[0157] 18 is a schematic diagram illustrating the operation of a model automator 1250, according to some implementations. As shown in FIG. 18, the model automator 1250 can include a model deployment engine 1810, a model deployment engine updater 1812, an orchestrator 1814, and a data handler 1816.
[0158] Model automator 1250 selects a model deployment location for an AI model using model deployment engine 1810, which can be selected from among several available computing environments. For example, as illustrated in FIG. 18 , model automator 1250 can access multiple public cloud environments (e.g., from first public cloud provider 1820A and second public cloud provider 1820B) as well as one or more on-premises environments (such as on-premises machines 1830A or on-premises elastic resources 1830B).
[0159] Various deployment environments for AI models can offer various advantages and disadvantages, particularly regarding cost efficiency, computational capacity, and privacy. On-premises environments often have limited capabilities because they are constrained by the available hardware at an enterprise. Yet, on-premises environments can offer privacy benefits that are not available in some cloud environments. On the other hand, cloud environments offer scalable resources, but costs can vary significantly depending on the cloud provider and the amount of computational power used. For example, cloud providers typically charge based on usage, sometimes using a tiered pricing system in which the price of a margin amount of computing resources varies depending on the total amount of resources used in a given time frame. Cloud and on-premises environments can also differ in their ease of integration with existing systems, flexibility in resource allocation, and their likelihood of downtime or service interruptions.
[0160] Model automator 1250 can monitor operational parameters associated with deployment environments available to the automator. Operational parameters can include relatively dynamic data, such as the amount of computational capacity available for use by a given model, the computing resources used during execution of a model deployed to an environment, the response time from a model deployed to an environment, or the accuracy of results produced by a model when deployed to an environment. Other example operational parameters include data that is static or changes over longer periods of time, such as the privacy policy of the operator of the deployment environment, the average amount of downtime for the environment, or the number of service interruptions to the environment. Model automator 1250 can measure some of the operational parameters as the automator deploys models to various available environments. Alternatively, model automator 1250 can obtain operational parameters from other sources, such as the operator of the environment or other systems performing computing tasks within the environment.
[0161] Additionally, model automator 1250 can facilitate the publication of models for use by other users. Within some organizations, it may be desirable for a model generated by one user to be made available to other users. For example, many users within an organization may perform similar tasks and would benefit from a model created by another user to assist with these tasks. Once a model is published, model automator 1250 can manage access rights or entitlements to the model. For example, model automator 1250 can link a model to access rights that specify that only users within a particular department of the organization can use the model, or that the model can only be used by users who have permission to access certain data (e.g., the data on which the model was trained).
[0162] The model deployment engine 1810 includes rules, models, or other logic tools that enable the model automator 1250 to select a model deployment location. For example, the model deployment engine 1810 can include one or more trained decision models, such as decision trees or random forests, a knowledge graph, or a rules engine. Upon determining a deployment location for a given model, the model deployment engine 1810 can input information such as parameters of the model itself (e.g., the size of the model, privacy considerations associated with the model, or information indicating whether the model will be used in real time or in a batch processing flow), parameters of the data that will be processed through the model when deployed (e.g., the amount of data processed per model iteration, the location of the data, or privacy considerations associated with the data), or operational parameters associated with available deployment locations. Based on one or more of these inputs, the model deployment engine 1810 selects a model deployment location for the model. In some implementations, the model deployment engine 1810 includes an explainability layer that enables the engine to output an explanation for its selected model deployment location. For example, if a user is using a chat-style interface from the model review assistant 1210 to create and deploy a model, the model deployment engine 1810 can generate explanations of its decisions that can be output to the user via the model review assistant 1210.
[0163] In one example, the model deployment engine 1810 evaluates the cost-efficiency of available environments and deploys the model to environments whose cost-efficiency is greater than a specified threshold. Cost-efficiency can be measured based on the amount of computing resources the model is expected to use and the expected cost of using those resources for each available environment. The model deployment engine 1810 can determine cost-efficiency as a standalone metric associated with each available model deployment environment or as a differential metric comparing the cost of deploying the model in one environment to the cost of deploying the model in another environment. In other cases, the model deployment engine 1810 selects a deployment location based at least in part on a privacy policy associated with the model, the input data to be processed by the model, or the output produced when the model processes the input. For example, a model is deployed to an on-premises environment if the model has a privacy policy that limits its use to an on-premises system, but is deployed to a cloud environment if there is no such restriction. In yet other cases, the model deployment engine 1810 can select a model deployment location based in part on whether the model will be used to process data in real time or in a batch process. For example, if the model is to be used in a real-time processing flow, the engine 1810 may select a cloud environment to deploy the model based on a determination that the cloud environment can more easily scale resources than an on-premises environment to ensure availability of the model, whereas if the model is to be used in a batch processing flow, the engine 1810 may cause the model to be deployed in an on-premises environment based on a determination that execution of the batch process can be delayed, if necessary, until computing resources are available.
[0164] The model deployment engine updater 1812 can update the model deployment engine 1810 based on its continued observation of the operating parameters. For example, when the model deployment engine 1810 includes a trained decision model, the model deployment engine updater 1812 can retrain the decision model based on the updated operating parameters. When the model deployment engine 1810 includes explicit rules, the model deployment engine updater 1812 can update the rules by instructing a large language model (LLM) to modify existing rules or create new rules based on the observed operating parameters. For example, if a cloud service provider updates its privacy policy, the model deployment engine updater 1812 can prompt the LLM to evaluate the updated privacy policy to determine whether it complies with the privacy policy of the organization operating the model automator 1250. Based on the evaluation, the LLM can then update rules about acceptable deployment locations for a particular type of AI model if, for example, the first cloud service provider modifies its privacy policy such that it no longer complies with the privacy policy associated with the particular AI model or the data that the particular AI model processes.
[0165] Model automator 1250 can use model deployment engine 1810 to both select a model deployment location for a new model, or new instance of a model, and to move a model from one deployment location to another. For example, as engine 1810 is updated in response to observed operating parameters, model automator 1250 can periodically use engine 1810 to reevaluate whether a model is deployed to a location that meets the engine's metrics (e.g., whether the environment's cost-effectiveness is still greater than a corresponding threshold or whether the environment still complies with a privacy policy associated with the data). In another example, model automator 1250 can determine that a new instance of a model should be deployed to a second environment when the original instance of the model uses computing resources in a first environment that exceed a given threshold (e.g., a specified pricing tier from the cloud computing provider operating the first environment).
[0166] Orchestrator 1814 facilitates deployment of the AI model to a location selected by model automator 1250 based on model deployment engine 1810, which makes the model available for use in a production environment. When deploying a model, orchestrator 1814 can configure the model for deployment to the selected infrastructure (including transporting the model to the selected infrastructure) and configure the infrastructure for the model (e.g., by spinning up the resources required for the model). Orchestrator 1814 is configured to automatically and seamlessly deploy the model in any of the available environments using platform-specific APIs. Orchestrator 1814 can also leverage containerization technologies such as Docker and orchestration frameworks such as Kubernetes to package the model with all necessary dependencies and scale and manage the containerized application.
[0167] The orchestrator 1814 can determine deployment attributes required for deployment of each model or that will improve the performance of the deployed model. These deployment attributes can include, for example, the model's language (such as Python or R), the model's infrastructure needs (such as available memory, processing speed, available parallelism, or hardware type), or configuration parameters for Docker file or Kubernetes deployment. In some implementations, the orchestrator 1814 maintains a set of patterns or templates, each of which specifies a corresponding type of deployment attribute for a model. Some of these patterns or templates can be initially provided by a user. Other patterns or templates can be automatically generated by the orchestrator 1814, for example, by identifying similar types of deployment attributes for a model. The orchestrator 1814 can automatically update the patterns or templates over time as the orchestrator 1814 deploys models using attributes in the patterns or templates.
[0168] After deploying a model, orchestrator 1814 can monitor the deployed model to verify that the deployment was successful. Success can be measured, for example, by an indicator specifying whether the model is producing results, by measurements of model operating parameters (e.g., latency, memory utilization, or CPU or GPU utilization), by measurements of model performance (e.g., accuracy or precision), or by a combination of factors. Orchestrator 1814 can determine that a model was not successfully deployed, for example, if the model is not producing results, if model operating parameters fall outside specified ranges, or if model performance differs from expected model performance by at least a threshold amount. When orchestrator 1814 determines that a model was not successfully deployed, orchestrator 1814 can decide to roll back the configuration to a previous configuration, deploy the model on a different infrastructure, stop model deployment entirely until errors are corrected, or take other corrective action to improve model deployment. Orchestrator 1814 can also update deployment templates based on successful or unsuccessful deployments. For example, if it is determined that configuration parameters in a Docker file caused model operating parameters for a particular type of model to fall outside of expected ranges, orchestrator 1814 can update the deployment template for the type of model to ensure that the correct configuration parameters are used in future Docker files for the same type of model.
[0169] The data handler 1816 makes data available to the deployed AI model at the location where the AI model is deployed. In some implementations, the data handler 1816 identifies a pipeline of application data for the deployed model, which is processed by the data processor 1220 described above. Using an API associated with the environment in which the model is deployed, the data handler 1816 can generate a script that makes the application data available to the environment for processing by the model. The data handler 1816 can also handle data privacy policies, check data for compliance with ethical or fairness standards, and / or obtain or generate approvals for data that can be used in a given situation.
[0170] 19 is a flowchart illustrating a process 1900 for automating the deployment of an AI model, according to some implementations. The process 1900 can be performed by a computing system, such as a system implementing aspects of the model automator 1250.
[0171] 19, a computer system receives a first request to deploy a first AI model to make the first AI model available for use in a production environment at 1902. In some implementations, the request can be received as natural language input to a chat interface, similar to the input described with respect to FIGS.
[0172] At 1904, the computer system selects a first model deployment location for the first AI model based on the model deployment engine. The deployment location can be selected from among a cloud provider environment or an environment operated at least in part by the entity controlling the first AI model (an “on-premises environment”). For example, an entity may contract with multiple cloud providers to use computing environments maintained by the cloud providers, but may also maintain some of its own computing infrastructure. The computer system can select whether to deploy the first model to an on-premises infrastructure or a cloud environment, and / or select a specific location within the on-premises infrastructure or a specific cloud environment where the model will be deployed. To select the model deployment location, the model deployment engine can take into account the nature of the first AI model itself, the nature of the deployment location, or considerations regarding model governance or entity-specific policies of the entity controlling the first AI model. The model deployment engine can include one or more trained models, such as decision trees or random forests, one or more rule-based systems, such as knowledge graphs or rule engines, or a combination of logic or tools that enable the computer system to make a determination about the model deployment location.
[0173] At 1906, the computer system generates a script to deploy the first AI model to the first model deployment location. For example, the computer system can generate a script that calls a platform-specific API for the selected model deployment location, leverage cloud technologies (such as containerization or orchestration) for a cloud deployment location or file system or server management technologies for an on-premises deployment location, and ensure that the deployed model has access to a data pipeline with the data the model is configured to process.
[0174] After deploying the first AI model, the computer system may monitor operating parameters associated with the deployment of the model at 1908. The operating parameters may include, for example, the computational cost used by the first AI model, or the response time from the first AI model when deployed at a selected location, the privacy policy of the environment that includes the deployment location of the model, or a measurement of downtime or service interruption of the environment that includes the deployment location.
[0175] Based on the operating parameters, the computer system may update the model deployment engine at 1910, for example, by retraining a trained model in the model deployment engine or by updating rules in the engine. For example, the computer system may update the model when a cloud provider modifies its privacy policy or when actual operating parameters monitored by the computer system differ from the operating parameters used to train the model deployment engine.
[0176] The updated model deployment engine can then be used to deploy a second AI model or redeploy the first AI model to another location, at 1912. In some cases, the second AI model can be deployed to the same environment as the first AI model or to a different environment than the first model based on observed operating parameters from the deployment of the first model. For example, if the first model is deployed to an on-premises system and the on-premises system is approaching its computational capacity, the second model can be deployed to a cloud environment. In other cases, the second AI model is a second instance of the first model deployed to another location. For example, if an entity deploys a first AI model to a first cloud environment where the entity is approaching a certain computational cost threshold, a second instance of the first model can be automatically deployed to the second cloud environment to reduce computational costs in the first environment. In still other cases, the computer system can determine that the first AI model should be moved from one deployment location to another. For example, if the privacy policy of the first cloud provider changes after the first model is deployed in the first provider's environment, the computer system can migrate the first model from the first provider's environment to another cloud provider's environment.
[0177] Process 1900 can be repeated as additional models are deployed by the computer system. The computer system can thus iteratively improve its knowledge of how well models perform in different environments, the cost of deploying the models to those environments, and how well the environments comply with governance or policy considerations. Leveraging this iteratively improved knowledge, the computer system can improve its ability to automatically and efficiently deploy AI models.
[0178] section Clause 1. A system for reducing computational resource usage when accessing siloed data across heterogeneous locations via a unified metadata graph, the system comprising: at least one hardware processor; receiving, at a graphical user interface (GUI), a user-specified query that, when executed by the at least one hardware processor, indicates a request to access a set of data objects, wherein each data object in the set of data objects is stored in a respective data silo in a set of data silos among the heterogeneous locations; parsing the user-specified query for a set of keywords, wherein each keyword in the set of keywords is associated with a set of data objects; performing natural language processing on the set of keywords to determine a set of semantically similar phrases corresponding to each keyword in the set of keywords; and accessing the metadata graph to determine nodes corresponding to the set of semantically similar phrases. and (ii) an edge indicating data lineage between a first node and a second node in the set of nodes, wherein the metadata graph comprises: (i) a set of nodes indicating (a) metadata of internal data objects stored in the data silo and (b) a location identifier of the data silo; and (ii) an edge indicating data lineage between a first node and a second node in the set of nodes, wherein the metadata graph is generated using a metadata data structure based on the file-level and container-level metadata identifiers.
[0179] 2. The system of claim 1, wherein the metadata graph is generated by: retrieving from the second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier from the set of file-level metadata identifiers indicates metadata for a given data object stored in a respective data silo, and each container-level metadata identifier from the set of container-level metadata identifiers indicates metadata for a respective data silo in the second set of data silos; generating sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, respectively; generating a metadata data structure for mapping each semantically similar metadata identifier from the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; and generating the metadata graph using the generated metadata data structure.
[0180] Clause 3. The system of clause 1, further comprising instructions for receiving, via a second GUI, a second user-specified query indicating a request to generate an intended result; and providing the second user-specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising (i) a second artificial intelligence model to be used to generate the intended result; and (ii) a second set of data objects to be used in training the second artificial intelligence model; and in response to receiving a user selection indicating acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model; and (ii) obtaining the second set of data objects using the metadata graph, training the second artificial intelligence model using the set of data objects, and applying the second artificial intelligence model to generate the intended result.
[0181] Clause 4. The system of clause 3, further comprising instructions for: accessing a governance database to obtain a set of policies that prescribe usage measures corresponding to the second set of data objects; using the set of policies that prescribe usage measures corresponding to the second set of data objects to determine whether the second set of data objects is approved for use in training a second artificial intelligence model; using the second set of policies that prescribe usage measures corresponding to the artificial intelligence model predictions to determine whether output of the second artificial intelligence model is approved for provision to one or more computing systems; and applying the second artificial intelligence model to generate the intended results in response to (i) the second set of data objects being approved for use in training the second artificial intelligence model, and (ii) the output of the second artificial intelligence model being approved for provision to one or more computing systems.
[0182] Clause 5. A method for reducing computational resource usage when accessing siloed data across disparate locations via a unified metadata graph, the method comprising: receiving, at a graphical user interface (GUI), a user-specified query indicating a request to access a set of data objects, wherein each data object in the set of data objects is stored in a respective data silo in a set of data silos among the disparate locations; performing natural language processing on the user-specified query to determine a set of phrases corresponding to the user-specified query; and accessing the metadata graph to determine nodes corresponding to the set of phrases, the metadata graph comprising: (i)(a) a metadata graph including: and (ii) edges indicating data lineage of the set of nodes, wherein the metadata graph is generated using a metadata data structure based on the file-level and container-level metadata identifiers; determining a data silo that stores at least one data object of the set of data objects using the location identifier corresponding to the determined node to retrieve at least one data object of the set of data objects through the data silo; and generating a visual representation of the at least one data object for display on a GUI.
[0183] Clause 6. The method of clause 5, wherein the metadata graph is generated by retrieving from the second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier from the set of file-level metadata identifiers indicates metadata for a given data object stored in a respective data silo, and each container-level metadata identifier from the set of container-level metadata identifiers indicates metadata for a respective data silo in the second set of data silos; generating sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, respectively; generating a metadata data structure for mapping each semantically similar metadata identifier from the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; and generating the metadata graph using the generated metadata data structure.
[0184] Clause 7. The method of clause 5, further including receiving, via a second GUI, a second user-specified query indicating a request to generate an intended result; and providing the second user-specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising (i) a second artificial intelligence model to be used to generate the intended result; and (ii) a second set of data objects to be used in training the second artificial intelligence model; and in response to receiving a user selection indicating acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model, and (ii) obtaining the second set of data objects using the metadata graph, training the second artificial intelligence model using the set of data objects, and applying the second artificial intelligence model to generate the intended result.
[0185] Clause 8. The method of clause 7, further including: accessing a governance database to obtain a set of policies that prescribe usage measures corresponding to the second set of data objects; using the set of policies that prescribe usage measures corresponding to the second set of data objects to determine whether the second set of data objects is approved for use in training a second artificial intelligence model; using the second set of policies that prescribe usage measures corresponding to the artificial intelligence model predictions to determine whether an output of the second artificial intelligence model is approved for provision to one or more computing systems; and applying the second artificial intelligence model to generate an intended result in response to (i) the second set of data objects being approved for use in training the second artificial intelligence model, and (ii) the output of the second artificial intelligence model being approved for provision to one or more computing systems.
[0186] Clause 9. The method of clause 5, wherein determining the set of phrases corresponding to the user-specified query further includes parsing the user-specified query for a set of keywords, each keyword in the set of keywords being associated with a set of data objects; determining, for each keyword in the set of keywords associated with the set of data objects, a set of semantically similar phrases corresponding to each keyword in the set of keywords; and determining the set of phrases corresponding to the user-specified query using the set of semantically similar phrases corresponding to each keyword in the set of keywords.
[0187] Clause 10. The method of clause 9, wherein determining semantically similar phrases corresponding to each keyword in the set of keywords further includes accessing a database indicating a mapping between the first keyword and the set of second keywords, and in response to accessing the database, using each keyword to determine a set of semantically similar phrases corresponding to each keyword.
[0188] Clause 11. The method of clause 5, wherein accessing the metadata graph further includes traversing each node of the set of nodes to identify a metadata identifier that matches at least one phrase of the set of phrases, and determining a node that corresponds to the set of phrases in response to determining that the metadata identifier matches at least one phrase of the set of phrases.
[0189] Clause 12. The method of clause 5, wherein accessing the metadata graph further includes: traversing each node of the set of nodes to identify a metadata identifier that matches at least one phrase of the set of phrases; determining a first node corresponding to the set of phrases in response to determining that the metadata identifier matches at least one phrase of the set of phrases; performing a second traversal of nodes of the set of nodes using an edge that indicates a first data lineage of the first node, where the first data lineage of the first node indicates a second node comprising information that is a source of the information associated with the first node; determining a second data silo that stores a second data object of the set of data objects using a location identifier corresponding to the second node to retrieve a second data object of the set of data objects via the second data silo; and generating a second visual representation of the second data object for display to the GUI.
[0190] Clause 13. The method of clause 5, wherein the visual representation of the at least one data object comprises lineage information of the at least one data object.
[0191] Clause 14. A method for implementing a method of accessing a set of data objects, the method comprising: receiving, at a graphical user interface (GUI), a user-specified query that, when executed by one or more processors, indicates a request to access a set of data objects, wherein each data object in the set of data objects is stored in a respective data silo in a set of data silos among heterogeneous locations; performing natural language processing on the user-specified query to determine a set of phrases corresponding to the user-specified query; and accessing a metadata graph to determine nodes corresponding to the set of phrases, the metadata graph comprising: (i) (a) metadata indicating internal data objects stored in the data silos; and (b) data silos. and (ii) edges indicating data lineage of the set of nodes, wherein the metadata graph is generated using a metadata data structure based on the file-level and container-level metadata identifiers; and one or more non-transitory computer-readable media having stored thereon instructions to perform operations including: accessing a data silo that stores at least one data object from the set of data objects using the location identifier corresponding to the determined node to retrieve at least one data object from the set of data objects through the data silo; and generating a visual representation of the at least one data object for display on a GUI.
[0192] Clause 15. The medium of Clause 14, wherein the metadata graph is generated by retrieving from the second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier from the set of file-level metadata identifiers indicates metadata for a given data object stored in a respective data silo, and each container-level metadata identifier from the set of container-level metadata identifiers indicates metadata for a respective data silo in the second set of data silos; generating sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, respectively; generating a metadata data structure for mapping each semantically similar metadata identifier from the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; and generating the metadata graph using the generated metadata data structure.
[0193] Clause 16. The medium of clause 14, wherein the instructions, when executed by one or more processors, further cause the medium to perform operations including receiving, via a second GUI, a second user-specified query indicating a request to generate an intended result; and providing the second user-specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising (i) the second artificial intelligence model to be used to generate the intended result; and (ii) a second set of data objects to be used in training the second artificial intelligence model; and, in response to receiving a user selection indicating acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model, and (ii) obtaining the second set of data objects using the metadata graph, training the second artificial intelligence model using the set of data objects, and applying the second artificial intelligence model to generate the intended result.
[0194] Clause 17. The medium of clause 16, wherein the instructions, when executed by one or more processors, further cause the medium to perform operations including: accessing a governance database to obtain a set of policies that prescribe usage measures corresponding to the second set of data objects; using the set of policies that prescribe usage measures corresponding to the second set of data objects to determine whether the second set of data objects is approved for use in training a second artificial intelligence model; using the second set of policies that prescribe usage measures corresponding to the artificial intelligence model predictions to determine whether output of the second artificial intelligence model is approved for provision to one or more computing systems; and, in response to (i) the second set of data objects being approved for use in training the second artificial intelligence model, and (ii) the output of the second artificial intelligence model being approved for provision to one or more computing systems, applying the second artificial intelligence model to generate an intended result.
[0195] Clause 18. The medium of clause 14, wherein determining the set of phrases corresponding to the user-specified query further includes parsing the user-specified query for a set of keywords, each keyword in the set of keywords being associated with a set of data objects; determining, for each keyword in the set of keywords associated with the set of data objects, a set of semantically similar phrases corresponding to each keyword in the set of keywords; and determining the set of phrases corresponding to the user-specified query using the set of semantically similar phrases corresponding to each keyword in the set of keywords.
[0196] Clause 19. The medium of clause 18, wherein determining semantically similar phrases corresponding to each keyword in the set of keywords further includes accessing a database indicating a mapping between the first keyword and the set of second keywords, and in response to accessing the database, using each keyword to determine a set of semantically similar phrases corresponding to each keyword.
[0197] Clause 20. The medium of clause 14, wherein accessing the metadata graph further includes traversing each node of the set of nodes to identify a metadata identifier that matches at least one phrase of the set of phrases, and determining a node that corresponds to the set of phrases in response to determining that the metadata identifier matches at least one phrase of the set of phrases. US01 Allowed
[0198] Item 21. A system for reducing data search time when accessing siloed data across disparate locations by generating a unified metadata graph via a search expansion generation (RAG) framework, the system comprising: at least one hardware processor; and when executed by the at least one hardware processor, receiving raw data from a set of data silos, the raw data comprising a set of metadata identifiers indicating (i) file-level metadata identifiers, (ii) container-level metadata identifiers, and (iii) system-level metadata identifiers; selecting a first structured LLM prompt from a set of structured large language model (LLM) prompts, the first structured LLM prompt corresponding to a first metadata identifier in the set of metadata identifiers; augmenting the first structured LLM prompt with the first metadata identifier, the first structured LLM prompt being provided to an LLM communicatively coupled to a set of domain-specific ontologies, the LLM performing a first intermediate output process without accessing the set of domain-specific ontologies, the first intermediate output process indicating a second set of metadata identifiers corresponding to the first metadata identifier. augmenting the first structured LLM prompt with a second set of metadata identifiers corresponding to the first metadata identifiers to be provided to the LLM, wherein the LLM is configured to generate a second intermediate output indicating filtered domain-specific metadata identifiers by accessing a set of domain-specific ontologies; generating a domain-specific integrated metadata graph via the LLM using (i) the first metadata identifiers and (ii) the second intermediate output indicating the filtered domain-specific metadata identifiers, wherein the filtered domain-specific metadata identifiers are traversable identifiers and the first metadata identifiers are non-traversable identifiers in the domain-specific integrated metadata graph; performing a validation process on the domain-specific integrated metadata graph by comparing a first performance metric of the domain-specific integrated metadata graph with a second performance metric of another version of the domain-specific integrated metadata graph, and wherein the first performance metric isand at least one non-transitory memory storing instructions that cause the system to perform an update process on the domain-specific unified metadata graph in response to determining that the second performance metric of the other version of the domain-specific unified metadata graph cannot be met or exceeded.
[0199] Clause 22. The system of clause 21, wherein the set of raw data is received by performing a crawling process across a set of data silos associated with entities to obtain raw data comprising a set of metadata identifiers.
[0200] Clause 23. The system of clause 21, wherein the instructions, when executed by the at least one hardware processor, further cause the system to: extract a first value from each of the set of data silos; determine a data type corresponding to each first value extracted from each of the set of data silos; and generate a data profile for each data silo in the set of data silos, the data profile indicating a data type of a value stored in the data silo.
[0201] Clause 24. The system of clause 23, wherein selecting a first structured LLM prompt from the set of structured LLM prompts that corresponds to a first metadata identifier among the set of metadata identifiers further includes determining a data silo that stores data corresponding to the first metadata identifier, retrieving a first data profile that corresponds to the data silo that stores the data corresponding to the first metadata identifier, filtering the set of structured LLM prompts using the first data profile to generate a set of filtered LLM prompts, and selecting a first structured LLM prompt from the set of filtered LLM prompts that corresponds to the first metadata identifier among the set of metadata identifiers.
[0202] Clause 25. The validation process provides a first query requesting a location of a first data item to each of (i) the domain-specific integrated metadata graph and (ii) another version of the domain-specific integrated metadata graph, the providing of the first query causing generation of a first result indicating the location of the first data item from the domain-specific integrated metadata graph and a second result indicating the location of the first data item for the other version of the domain-specific integrated metadata graph; and providing a first sub-performance metric for the domain-specific integrated metadata graph and a second sub-performance metric for the other version of the domain-specific integrated metadata graph. 22. The system of claim 21, further comprising: calculating a second sub-performance metric for the domain-specific unified metadata graph, where the first sub-performance metric and the second sub-performance metric are query-to-result performance metrics; and calculating a third sub-performance metric for the domain-specific unified metadata graph and a fourth sub-performance metric for the other version of the domain-specific unified metadata graph, where the third sub-performance metric and the fourth sub-performance metric are result accuracy metrics, where the first performance metric comprises the first sub-performance metric and the third sub-performance metric, and the second performance metric comprises the second sub-performance metric and the fourth sub-performance metric.
[0203] Clause 26. The system of clause 25, wherein determining that the first performance metric fails to meet or exceed the second performance metric is based on determining that (i) the first sub-performance metric fails to meet or exceed the second sub-performance metric, or (ii) the third sub-performance metric fails to meet or exceed the fourth sub-performance metric.
[0204] Clause 27. The system of clause 21, wherein performing an update process on the domain-specific unified metadata graph includes determining a set of mismatches between the domain-specific unified metadata graph and another version of the domain-specific unified metadata graph, the set of mismatches reflecting mismatches between (i) nodes in the domain-specific unified metadata graph and the other version of the domain-specific unified metadata graph and (ii) edges connected to at least one node in the domain-specific unified metadata graph and the other version of the domain-specific unified metadata graph; and updating the domain-specific unified metadata graph with updated nodes and edges in the other version of the domain-specific unified metadata graph that correspond to the set of mismatches.
[0205] Clause 28. The system of clause 21, wherein the instructions, when executed by the at least one hardware processor, further cause the system to detect the addition of a data silo to a computing environment associated with the first entity, and in response to detecting the addition of the data silo, cause a second update process to be performed on the domain-specific unified metadata graph.
[0206] Clause 29. The system of clause 21, wherein the file-level metadata identifier indicates metadata for a given data object stored within a respective data silo of the set of data silos, the container-level metadata identifier indicates metadata for a respective data silo of the set of data silos, and the system-level metadata identifier indicates metadata for a computing system hosting a respective data silo of the set of data silos.
[0207] Clause 30. The system of clause 21, wherein the domain-specific unified metadata graph comprises: (i) a first node indicating (a) the filtered metadata identifier, (b) the first metadata identifier, and (c) a data silo location identifier associated with the first metadata identifier, and (ii) at least one edge indicating data lineage between the first node and a second node.
[0208] Clause 31. A method for reducing data search time when accessing siloed data across disparate locations by generating a unified metadata graph via a search expansion generation (RAG) framework, comprising: selecting a first LLM prompt from a set of large language model (LLM) prompts, the first LLM prompt corresponding to a first metadata identifier among the set of metadata identifiers; augmenting the first LLM prompt with the first metadata identifier to be provided to the LLM, the LLM being configured to generate a first intermediate output indicating a second set of metadata identifiers corresponding to the first metadata identifier; and augmenting the first LLM prompt with the second set of metadata identifiers corresponding to the first metadata identifier to be provided to the LLM, the LLM being configured to access a set of domain-specific ontologies. generating a domain-specific unified metadata graph via the LLM using (i) the first metadata identifiers and (ii) the second intermediate output indicating the filtered domain-specific metadata identifiers, wherein the filtered domain-specific metadata identifiers are traversable identifiers and the first metadata identifiers are non-traversable identifiers in the domain-specific unified metadata graph; and performing an update process on the domain-specific unified metadata graph in response to determining that a first performance metric of the domain-specific unified metadata graph fails to satisfy a performance measure for a second performance metric of another version of the domain-specific unified metadata graph.
[0209] Clause 32. The method of clause 31, further comprising: extracting a first value from each of the set of data silos; determining a data type corresponding to each first value extracted from each of the set of data silos; and generating a data profile for each data silo of the set of data silos, the data profile indicating the data types of values stored in the data silo.
[0210] Clause 33. The method of clause 32, wherein selecting a first LLM prompt from the set of LLM prompts that corresponds to a first metadata identifier among the set of metadata identifiers further includes determining a data silo that stores data corresponding to the first metadata identifier, retrieving a first data profile that corresponds to the data silo that stores the data corresponding to the first metadata identifier, filtering the set of LLM prompts using the first data profile to generate a set of filtered LLM prompts, and selecting the first LLM prompt from the set of filtered LLM prompts that corresponds to the first metadata identifier among the set of metadata identifiers.
[0211] Clause 34. Performing a validation process on a domain-specific integrated metadata graph, the validation process providing a first query requesting a location of a first data item to each of (i) the domain-specific integrated metadata graph and (ii) another version of the domain-specific integrated metadata graph, the first query causing generation of a first result indicating the location of the first data item from the domain-specific integrated metadata graph and a second result indicating the location of the first data item for the other version of the domain-specific integrated metadata graph. 32. The method of clause 31, further comprising: calculating a second sub-performance metric for the other version of the graph, wherein the first sub-performance metric and the second sub-performance metric are query-to-result performance metrics; and calculating a third sub-performance metric for the domain-specific integrated metadata graph and a fourth sub-performance metric for the other version of the domain-specific integrated metadata graph, wherein the third sub-performance metric and the fourth sub-performance metric are result accuracy metrics, wherein the first performance metric comprises the first sub-performance metric and the third sub-performance metric, and the second performance metric comprises the second sub-performance metric and the fourth sub-performance metric.
[0212] Clause 35. The method of clause 34, wherein determining that a first performance metric satisfies a performance measure for a second performance metric is based on determining that (i) the first sub-performance metric meets or exceeds the second sub-performance metric, or (ii) the third sub-performance metric fails to meet or exceed the fourth sub-performance metric.
[0213] Clause 36. The method, when executed by one or more processors, includes selecting a first LLM prompt from a set of large language model (LLM) prompts that corresponds to a first metadata identifier of the set of metadata identifiers; augmenting the first LLM prompt with the first metadata identifier to be provided to the LLM, wherein the LLM is configured to generate a first intermediate output indicating a second set of metadata identifiers that correspond to the first metadata identifier; and augmenting the first LLM prompt with the second set of metadata identifiers that correspond to the first metadata identifier to be provided to the LLM, wherein the LLM is configured to generate a second intermediate output indicating the filtered domain-specific metadata identifiers by accessing a set of domain-specific ontologies. and generating a domain-specific integrated metadata graph via the LLM using (i) the first metadata identifiers and (ii) a second intermediate output indicating filtered domain-specific metadata identifiers, wherein the filtered domain-specific metadata identifiers are traversable identifiers and the first metadata identifiers are non-traversable identifiers in the domain-specific integrated metadata graph; and performing an update process on the domain-specific integrated metadata graph in response to determining that a first performance metric of the domain-specific integrated metadata graph fails to meet a performance measure for a second performance metric of another version of the domain-specific integrated metadata graph.
[0214] Clause 37. The medium of clause 36, wherein the instructions, when executed by one or more processors, further cause the medium to perform operations including: extracting a first value from each of the set of data silos; determining a data type corresponding to each first value extracted from each of the set of data silos; and generating a data profile for each data silo of the set of data silos, the data profile indicating the data type of the values stored in the data silo.
[0215] Clause 38. The medium of clause 37, wherein selecting a first LLM prompt from the set of LLM prompts that corresponds to a first metadata identifier among the set of metadata identifiers further includes determining a data silo that stored data corresponding to the first metadata identifier, retrieving a first data profile that corresponds to the data silo that stored the data corresponding to the first metadata identifier, filtering the set of LLM prompts using the first data profile to generate a set of filtered LLM prompts, and selecting the first LLM prompt from the set of filtered LLM prompts that corresponds to the first metadata identifier among the set of metadata identifiers.
[0216] Clause 39. The instructions, when executed by one or more processors, perform a validation process on a domain-specific unified metadata graph, the validation process providing a first query requesting a location of a first data item to each of (i) the domain-specific unified metadata graph and (ii) another version of the domain-specific unified metadata graph, the first query causing generation of a first result indicating a location of the first data item from the domain-specific unified metadata graph and a second result indicating a location of the first data item for the other version of the domain-specific unified metadata graph; providing a first sub-performance metric for the domain-specific unified metadata graph, and a second sub-performance metric for the domain-specific unified metadata graph. and computing a second sub-performance metric for the other version of the domain-specific unified metadata graph, the first sub-performance metric and the second sub-performance metric being query-to-results performance metrics, and computing a third sub-performance metric for the domain-specific unified metadata graph and a fourth sub-performance metric for the other version of the domain-specific unified metadata graph, the third sub-performance metric and the fourth sub-performance metric being result accuracy metrics, the first performance metric comprising the first sub-performance metric and the third sub-performance metric, and the second performance metric comprising the second sub-performance metric and the fourth sub-performance metric.
[0217] Clause 40. The medium of clause 39, where the determination that the first performance metric fails to meet the performance measure for the second performance metric is based on determining that (i) the first sub-performance metric fails to meet or exceed the second sub-performance metric, or (ii) the third sub-performance metric fails to meet or exceed the fourth sub-performance metric. US02 Allowed.
[0218] Clause 41. Receiving, at a computer system, a first natural language input from a user including a set of phrases and instructions for analyzing data associated with the set of phrases using an artificial intelligence (AI) model; and, in response to the first natural language input, accessing a metadata graph to determine nodes corresponding to the set of phrases, the metadata graph comprising: (i) a set of nodes comprising: (a) metadata indicating internal data objects stored in a data silo, and (b) a location identifier of the data silo, and (ii) edges indicating data lineage of the set of nodes; and processing one or more internal data objects indicated by the determined nodes to generate a first set of application data; and generating one or more first outputs. 1. A computer-implemented method comprising: applying an AI model to a first set of application data, wherein one or more first outputs comprise classifications of data items in the first set of application data or one or more predictions made based on the first set of application data; transmitting a representation of the one or more first outputs for display to a user; receiving a second natural language input from the user, wherein the second natural language input includes instructions for modifying the first set of application data; generating, by a computer system, a second set of application data based on the received second natural language input; and applying the AI model to the second set of application data to generate the one or more second outputs.
[0219] Clause 42. The method of clause 41, wherein processing the internal data object to generate the first set of application data includes removing personally identifiable or private information from the internal data object.
[0220] Clause 43. The method of clause 41, wherein processing the internal data object to generate a first set of application data includes applying a set of data modification operators to the internal data object to generate a set of modified data based on the internal data object, the first set of application data including one or more modified data items from the set of modified data.
[0221] Clause 44. The method of clause 41, wherein processing the internal data object to generate the first set of application data includes removing inaccurate data items from the internal data object.
[0222] Clause 45. The method of clause 41, further comprising: before receiving the first natural language input, receiving a third natural language input including instructions for generating an AI model; accessing the metadata graph to determine a node corresponding to the third natural language input; processing internal data objects associated with the node corresponding to the third natural language input to generate a set of training data; and training the AI model using the set of training data to generate a trained AI model, wherein applying the AI model to the first set of application data comprises applying the trained AI model to the first set of application data.
[0223] Clause 46. The method of clause 45, wherein training the AI model includes accessing values of one or more model metrics associated with each of a plurality of model types; selecting, by the computer system, a model type for the trained AI model from among the plurality of model types based on the accessed values of the one or more model metrics; and training the selected model type.
[0224] Clause 47. The method of clause 45, wherein training the AI model includes receiving a user selection of a model type for the AI model and training the selected model type.
[0225] Clause 48. The method of clause 41, further comprising generating, by the computer system, a chat interface for display to the user, wherein the first natural language input is received via the chat interface, and wherein sending a representation of one or more outputs for display to the user comprises displaying the one or more outputs via the chat interface.
[0226] Clause 49. The method of clause 41, wherein generating the second set of application data based on the second natural language input includes adding a data item to the first set of application data, removing a data item from the first set of application data, or applying a data modification operator to a value to modify the value of the data item in the first set of application data.
[0227] Clause 50. One or more processors and one or more non-transitory computer-readable storage media having executable instructions stored thereon, the instructions, when executed by the one or more processors, comprising: receiving a first natural language input from a user including a set of phrases, and instructions for analyzing data associated with the set of phrases using an artificial intelligence (AI) model; in response to the first natural language input, accessing a metadata graph to determine nodes corresponding to the set of phrases, the metadata graph comprising: (i) a set of nodes comprising: (a) metadata indicating internal data objects stored in a data silo, and (b) a location identifier of the data silo, and (ii) an edge indicating a data lineage of the set of nodes; accessing the first natural language input from a user including a set of phrases and an edge indicating a data lineage of the set of nodes; and one or more non-transitory computer-readable storage media that cause the system to: process one or more internal data objects indicated by the determined node to generate a set of application data; apply an AI model to the first set of application data to generate one or more first outputs; transmit a representation of the one or more first outputs for display to a user; receive a second natural language input from the user, the second natural language input including instructions for modifying the first set of application data; generate the second set of application data based on the received second natural language input; and apply the AI model to the second set of application data to generate one or more second outputs.
[0228] Clause 51. The system of clause 50, wherein processing the internal data objects to generate a first set of application data includes removing personally identifiable or private information from the internal data objects.
[0229] Clause 52. The system of clause 50, wherein processing the internal data object to generate a first set of application data includes applying a set of data modification operators to the internal data object to generate a set of modified data based on the internal data object, the first set of application data including one or more modified data items from the set of modified data.
[0230] Clause 53. The system of clause 50, wherein processing the internal data object to generate the first set of application data includes removing inaccurate data items from the internal data object.
[0231] Clause 54. The system of clause 50, wherein the instructions, when executed by the one or more processors, further cause the system to: before receiving the first natural language input, receive a third natural language input including instructions for generating an AI model; access the metadata graph to determine a node corresponding to the third natural language input; process internal data objects associated with the node corresponding to the third natural language input to generate a set of training data; and train the AI model using the set of training data to generate a trained AI model, wherein applying the AI model to the first set of application data includes applying the trained AI model to the first set of application data.
[0232] Clause 55. The system of clause 50, wherein the instructions, when executed by the one or more processors, further cause the system to generate a chat interface for display to the user, wherein the first natural language input is received via the chat interface, and wherein sending a representation of one or more outputs for display to the user includes displaying the one or more outputs via the chat interface.
[0233] Clause 56. The system of clause 50, wherein generating the second set of application data based on the second natural language input includes adding a data item to the first set of application data, removing a data item from the first set of application data, or applying a data modification operator to a value to modify a value of the data item in the first set of application data.
[0234] Clause 57. A non-transitory computer-readable storage medium having stored thereon executable instructions, the instructions, when executed by one or more processors of the system, comprising: receiving a first natural language input from a user including a set of phrases, and instructions for analyzing data associated with the set of phrases using an artificial intelligence (AI) model; and, in response to the first natural language input, accessing a metadata graph to determine nodes corresponding to the set of phrases, the metadata graph comprising: (i) a set of nodes comprising: (a) metadata indicating internal data objects stored in a data silo, and (b) a location identifier of the data silo, and (ii) an edge indicating a data lineage of the set of nodes; and accessing the first natural language input from a user including a set of phrases and an edge indicating a data lineage of the set of nodes. a first natural language input from the user, the second natural language input including instructions for modifying the first set of application data; generating the second set of application data based on the received second natural language input; and applying the AI model to the second set of application data to generate one or more second outputs.
[0235] Clause 58. The non-transitory computer-readable storage medium of clause 57, wherein processing the internal data object to generate a first set of application data comprises applying a set of data modification operators to the internal data object to generate a set of modified data based on the internal data object, the first set of application data comprising one or more modified data items from the set of modified data.
[0236] Clause 59. The non-transitory computer-readable storage medium of clause 57, wherein the instructions, when executed by one or more processors, further cause the system to: before receiving the first natural language input, receive a third natural language input including instructions for generating an AI model; access the metadata graph to determine a node corresponding to the third natural language input; process internal data objects associated with the node corresponding to the third natural language input to generate a set of training data; and train the AI model using the set of training data to generate a trained AI model, wherein applying the AI model to the first set of application data includes applying the trained AI model to the first set of application data.
[0237] Clause 60. The non-transitory computer-readable storage medium of clause 57, wherein the instructions, when executed by one or more processors, further cause the system to generate a chat interface for display to the user, wherein the first natural language input is received via the chat interface, and wherein sending a representation of one or more outputs for display to the user includes displaying the one or more outputs via the chat interface.
[0238] Clause 61. A computer-implemented method comprising: receiving a first request for a first AI model used by an entity to deploy the first artificial intelligence (AI) model and make the first AI model available for use in a production environment for processing input data and generating corresponding outputs; selecting a first model deployment location for the first AI model based on a model deployment engine, the model deployment engine being configured to select the first model deployment location from among a set of one or more cloud provider environments or an on-premises environment operated by the entity; generating a script for deploying the first AI model at the first model deployment location; monitoring operating parameters associated with deployment of the first AI model at the selected model deployment location as the first AI model processes the input data and generates corresponding outputs; updating the model deployment engine based on values of the monitored operating parameters; and selecting a second model deployment location for the second AI model based on the updated model deployment engine in response to a second request to deploy a second AI model.
[0239] Clause 62. The computer-implemented method of clause 61, wherein the model deployment engine comprises a decision tree, a knowledge graph, or a rules engine, and wherein selecting the first model deployment location based on the model deployment engine includes selecting the first model deployment location based on one or more of: a computational cost of deploying the first AI model at the selected model deployment location, a computational capacity of the selected model deployment location, a response time from the selected model deployment location, or a measure of accuracy of the first AI model when deployed at the selected model deployment location.
[0240] Clause 63. The computer-implemented method of clause 61, wherein selecting a first model deployment location based on the model deployment engine includes selecting the first model deployment location based on a privacy policy associated with at least one of the first AI model, the input data, or the corresponding output.
[0241] Clause 64. The computer-implemented method of clause 61, wherein the operational parameters include one or more of: a computational cost used by the first AI model at the first model deployment location; a response time from the first AI model when deployed at the first model deployment location; available computational capacity in an environment that includes the first model deployment location; a privacy policy of the environment that includes the first model deployment location; or a measure of downtime or service interruption in the environment that includes the first model deployment location.
[0242] Clause 65. The computer-implemented method of clause 61, wherein the model deployment engine comprises a trained decision model, and wherein updating the model deployment engine includes retraining the trained decision model based on a difference between values of the monitored operational parameters and values of the set of operational parameters on which the model deployment engine was trained.
[0243] Clause 66. The computer-implemented method of clause 61, wherein the second AI model is a second instance of the first AI model, and the second model deployment location for the second AI model is a different location than the first model deployment location for the first AI model.
[0244] Clause 67. The computer-implemented method of clause 61, further including selecting a third model deployment location for the first AI model, different from the first model deployment location, based on the updated model deployment engine; and generating a script for deploying the first AI model to the third model deployment location.
[0245] Clause 68. The computer-implemented method of clause 61, further including selecting a third model deployment location for the first AI model, different from the first model deployment location, based on the model deployment engine and the operating parameters, and generating a script for deploying the first AI model to the third model deployment location.
[0246] Clause 69. The computer-implemented method of clause 68, wherein the operational parameters include a computational cost associated with deploying the first AI model at the first model deployment location, and wherein selecting the third model deployment location includes determining to move the first AI model to the third model deployment location when the computational cost associated with deploying the first AI model at the first model deployment location is greater than a predicted computational cost associated with deploying the first AI model at the third model deployment location.
[0247] Section 70. A method for deploying a first artificial intelligence (AI) model to an entity, comprising: one or more processors; and one or more non-transitory computer-readable storage media having executable instructions stored thereon, the instructions, when executed by the one or more processors, receiving a first request for a first AI model to be used by the entity to deploy the first AI model and make the first AI model available for use in a production environment for processing input data and generating corresponding output; selecting a first model deployment location for the first AI model based on a model deployment engine, the model deployment engine selecting a first model deployment location for the first AI model from among a set of one or more cloud provider environments or an on-premises environment operated by the entity; and one or more non-transitory computer-readable storage media configured to select a model deployment location; generate a script for deploying a first AI model at the first model deployment location; monitor operating parameters associated with deployment of the first AI model at the selected model deployment location as the first AI model processes input data and generates corresponding outputs; update a model deployment engine based on values of the monitored operating parameters; and, in response to a second request to deploy a second AI model, select a second model deployment location for the second AI model based on the updated model deployment engine.
[0248] Clause 71. The system of clause 70, wherein the model deployment engine comprises a decision tree, a knowledge graph, or a rules engine, and wherein selecting the first model deployment location based on the model deployment engine includes selecting the first model deployment location based on one or more of a computational cost of deploying the first AI model at the selected model deployment location, a computational capacity of the selected model deployment location, a response time from the selected model deployment location, or a measure of accuracy of the first AI model when deployed at the selected model deployment location.
[0249] Clause 72. The system of clause 70, wherein selecting the first model deployment location based on the model deployment engine includes selecting the first model deployment location based on a privacy policy associated with at least one of the first AI model, the input data, or the corresponding output.
[0250] Clause 73. The system of clause 70, wherein the operational parameters include one or more of: a computational cost used by the first AI model at the first model deployment location; a response time from the first AI model when deployed at the first model deployment location; available computational capacity in an environment that includes the first model deployment location; a privacy policy in an environment that includes the first model deployment location; or a measurement of downtime or service interruption in an environment that includes the first model deployment location.
[0251] Clause 74. The system of clause 70, wherein the model deployment engine comprises a trained decision model, and wherein updating the model deployment engine includes retraining the trained decision model based on a difference between values of the monitored operational parameters and values of the set of operational parameters on which the model deployment engine was trained.
[0252] Clause 75. The system of clause 70, wherein the second AI model is a second instance of the first AI model, and wherein the second model deployment location for the second AI model is a different location than the first model deployment location for the first AI model.
[0253] Clause 76. The system of clause 70, wherein the instructions, when executed by the one or more processors, further cause the system to: select a third model deployment location for the first AI model, different from the first model deployment location, based on the updated model deployment engine; and generate a script for deploying the first AI model to the third model deployment location.
[0254] Clause 77. The system of clause 70, wherein the instructions, when executed by the one or more processors, further cause the system to: select a third model deployment location for the first AI model, different from the first model deployment location, based on the model deployment engine and operating parameters; and generate a script for deploying the first AI model to the third model deployment location.
[0255] Clause 78. The system of clause 77, wherein the operational parameters include a computational cost associated with deploying the first AI model at the first model deployment location, and wherein selecting the third model deployment location includes determining to move the first AI model to the third model deployment location when the computational cost associated with deploying the first AI model at the first model deployment location is greater than a predicted computational cost associated with deploying the first AI model at the third model deployment location.
[0256] Clause 79. A non-transitory computer-readable storage medium having executable instructions stored thereon, the instructions, when executed by one or more processors of the system, including: receiving a first request for a first artificial intelligence (AI) model used by an entity to deploy the first AI model and make the first AI model available for use in a production environment for processing input data and generating corresponding output; and selecting a first model deployment location for the first AI model based on a model deployment engine, wherein the model deployment engine selects a first model deployment location for the first AI model from among a set of one or more cloud provider environments or an on-premises environment operated by the entity. a non-transitory computer-readable storage medium configured to select a model deployment location; generate a script for deploying a first AI model at the first model deployment location; monitor operating parameters associated with deployment of the first AI model at the selected model deployment location as the first AI model processes input data and generates corresponding outputs; update a model deployment engine based on values of the monitored operating parameters; and, in response to a second request to deploy a second AI model, select a second model deployment location for the second AI model based on the updated model deployment engine.
[0257] Clause 80. The non-transitory computer-readable storage medium of clause 79, wherein the model deployment engine comprises a decision tree, a knowledge graph, or a rule engine, and wherein selecting the first model deployment location based on the model deployment engine includes selecting the first model deployment location based on one or more of a computational cost of deploying the first AI model at the selected model deployment location, a computational capacity of the selected model deployment location, a response time from the selected model deployment location, or a measure of accuracy of the first AI model when deployed at the selected model deployment location.
[0258] Clause 81. A system for reducing computational resource usage when accessing siloed data across disparate locations via a unified metadata graph, the system comprising: at least one hardware processor; and when executed by the at least one hardware processor, identifying a set of keywords associated with a request to access a set of data objects; performing natural language processing on the set of keywords to determine a set of semantically similar phrases corresponding to each keyword in the set of keywords; accessing the metadata graph to determine nodes corresponding to the set of semantically similar phrases, the metadata graph comprising: (i) a set of nodes indicating (a) metadata of internal data objects stored in the data silo, and (b) a location identifier of the data silo; and (ii) a set of nodes indicating (a) metadata of internal data objects stored in the data silo, and (b) a location identifier of the data silo. and edges indicating data lineage between a first node and a second node in the set of data objects, wherein the metadata graph is generated using a metadata data structure based on file-level and container-level metadata identifiers; determining a data silo that stores at least one data object of the set of data objects using a location identifier corresponding to the determined node to retrieve at least one data object of the set of data objects through the data silo; and generating a visual representation of the at least one data object for display on a graphical user interface (GUI), wherein the visual representation of the at least one data object comprises lineage information of the at least one data object.
[0259] Clause 82. The system of clause 81, wherein the metadata graph is generated by retrieving from the second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier from the set of file-level metadata identifiers indicates metadata for a given data object stored in a respective data silo, and each container-level metadata identifier from the set of container-level metadata identifiers indicates metadata for a respective data silo in the second set of data silos; generating sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, respectively; generating a metadata data structure for mapping each semantically similar metadata identifier from the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; and generating the metadata graph using the generated metadata data structure.
[0260] Clause 83. The system of clause 81, further comprising instructions for receiving, via a second GUI, a second user-specified query indicating a request to generate an intended result; and providing the second user-specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising (i) a second artificial intelligence model to be used to generate the intended result, and (ii) a second set of data objects to be used in training the second artificial intelligence model; and in response to receiving a user selection indicating acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model, and (ii) obtaining the second set of data objects using the metadata graph, training the second artificial intelligence model using the set of data objects, and applying the second artificial intelligence model to generate the intended result.
[0261] Clause 84. The system of clause 83, further comprising instructions for: accessing a governance database to obtain a set of policies that prescribe usage measures corresponding to the second set of data objects; using the set of policies that prescribe usage measures corresponding to the second set of data objects to determine whether the second set of data objects is approved for use to train a second artificial intelligence model; using the second set of policies that prescribe usage measures corresponding to the artificial intelligence model predictions to determine whether output of the second artificial intelligence model is approved for provision to one or more computing systems; and applying the second artificial intelligence model to generate an intended result in response to (i) the second set of data objects being approved for use to train the second artificial intelligence model, and (ii) the output of the second artificial intelligence model being approved for provision to one or more computing systems.
[0262] Clause 85. A method for reducing computational resource usage when accessing siloed data across disparate locations via a unified metadata graph, the method comprising: identifying a set of keywords associated with a user-specified query for accessing a set of data objects; performing natural language processing on the user-specified query to determine a set of phrases corresponding to the user-specified query; accessing the metadata graph to determine nodes corresponding to the set of phrases, the metadata graph comprising: (i) a set of nodes comprising: (a) metadata indicating internal data objects stored in the data silo; and (b) a location identifier of the data silo; and (ii) edges indicating data lineage of the set of nodes, the metadata graph being generated using a metadata data structure based on file-level and container-level metadata identifiers; and, to retrieve at least one data object from the set of data objects via the data silo, determining a data silo that stored at least one data object from the set of data objects using the location identifier corresponding to the determined node; and generating a representation of the at least one data object.
[0263] Clause 86. The method of clause 85, wherein the metadata graph is generated by retrieving from the second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier from the set of file-level metadata identifiers indicates metadata for a given data object stored in a respective data silo, and each container-level metadata identifier from the set of container-level metadata identifiers indicates metadata for a respective data silo in the second set of data silos; generating sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, respectively; generating a metadata data structure for mapping each semantically similar metadata identifier from the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; and generating the metadata graph using the generated metadata data structure.
[0264] Clause 87. The method of clause 85, further including receiving, via a second GUI, a second user-specified query indicating a request to generate an intended result; and providing the second user-specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising (i) a second artificial intelligence model to be used to generate the intended result; and (ii) a second set of data objects to be used in training the second artificial intelligence model; and in response to receiving a user selection indicating acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model, and (ii) obtaining the second set of data objects using the metadata graph, training the second artificial intelligence model using the set of data objects, and applying the second artificial intelligence model to generate the intended result.
[0265] Clause 88. The method of clause 87, further including: accessing a governance database to obtain a set of policies that prescribe usage measures corresponding to the second set of data objects; using the set of policies that prescribe usage measures corresponding to the second set of data objects to determine whether the second set of data objects is approved for use to train a second artificial intelligence model; using the second set of policies that prescribe usage measures corresponding to the artificial intelligence model predictions to determine whether output of the second artificial intelligence model is approved for provision to one or more computing systems; and applying the second artificial intelligence model to generate the intended results in response to (i) the second set of data objects being approved for use to train the second artificial intelligence model, and (ii) the output of the second artificial intelligence model being approved for provision to one or more computing systems.
[0266] Clause 89. The method of clause 85, wherein determining the set of phrases corresponding to the user-specified query further comprises parsing the user-specified query for a set of keywords, each keyword in the set of keywords being associated with a set of data objects, and for each keyword in the set of keywords associated with the set of data objects, determining a set of semantically similar phrases corresponding to each keyword in the set of keywords, and determining the set of phrases corresponding to the user-specified query using the set of semantically similar phrases corresponding to each keyword in the set of keywords.
[0267] Clause 90. The method of clause 89, wherein determining semantically similar phrases corresponding to each keyword in the set of keywords further includes accessing a database indicating a mapping between the first keyword and the set of second keywords, and in response to accessing the database, using each keyword to determine a set of semantically similar phrases corresponding to each keyword.
[0268] Clause 91. The method of clause 85, wherein accessing the metadata graph further includes traversing each node of the set of nodes to identify a metadata identifier that matches at least one phrase of the set of phrases, and determining a node corresponding to the set of phrases in response to determining that the metadata identifier matches at least one phrase of the set of phrases.
[0269] Clause 92. The method of clause 85, wherein accessing the metadata graph further includes: traversing each node of the set of nodes to identify a metadata identifier that matches at least one phrase of the set of phrases; determining a first node corresponding to the set of phrases in response to determining that the metadata identifier matches at least one phrase of the set of phrases; and performing a second traversal of nodes of the set of nodes using an edge that indicates a first data lineage of the first node, where the first data lineage of the first node indicates a second node comprising information that is a source of the information associated with the first node; determining a second data silo that stores a second data object of the set of data objects using a location identifier corresponding to the second node to retrieve a second data object of the set of data objects via the second data silo; and generating a second representation of the second data object.
[0270] Clause 93. The method of clause 85, wherein the representation of the at least one data object comprises lineage information for the at least one data object.
[0271] Clause 94. One or more non-transitory computer-readable media having stored thereon instructions that, when executed by one or more processors, cause operations including: performing natural language processing on a user-specified query to determine a set of phrases corresponding to the user-specified query; accessing a metadata graph to determine nodes corresponding to the set of phrases, the metadata graph comprising: (i) a set of nodes comprising: (a) metadata indicating internal data objects stored in the data silo, and (b) a location identifier of the data silo; and (ii) edges indicating data lineage of the set of nodes, the metadata graph being generated using a metadata data structure based on the file-level and container-level metadata identifiers; determining a data silo that stored at least one data object of the set of data objects using the location identifier corresponding to the determined node to retrieve at least one data object of the set of data objects through the data silo; and generating a representation of the at least one data object.
[0272] Clause 95. The medium of clause 94, wherein the metadata graph is generated by retrieving from the second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier from the set of file-level metadata identifiers indicates metadata for a given data object stored in a respective data silo, and each container-level metadata identifier from the set of container-level metadata identifiers indicates metadata for a respective data silo in the second set of data silos; generating sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, respectively; generating a metadata data structure for mapping each semantically similar metadata identifier from the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; and generating the metadata graph using the generated metadata data structure.
[0273] Clause 96. The medium of clause 94, wherein the instructions, when executed by one or more processors, further cause the medium to perform operations including receiving, via a second GUI, a second user-specified query indicating a request to generate an intended result; and providing the second user-specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising (i) the second artificial intelligence model to be used to generate the intended result, and (ii) a second set of data objects to be used in training the second artificial intelligence model; and, in response to receiving a user selection indicating acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model, and (ii) obtaining the second set of data objects using the metadata graph, training the second artificial intelligence model using the set of data objects, and applying the second artificial intelligence model to generate the intended result.
[0274] Clause 97. The medium of clause 96, wherein the instructions, when executed by one or more processors, further cause the medium to perform operations including: accessing a governance database to obtain a set of policies that prescribe usage measures corresponding to the second set of data objects; using the set of policies that prescribe usage measures corresponding to the second set of data objects to determine whether the second set of data objects is approved for use in training a second artificial intelligence model; using the second set of policies that prescribe usage measures corresponding to the artificial intelligence model predictions to determine whether output of the second artificial intelligence model is approved for provision to one or more computing systems; and, in response to (i) the second set of data objects being approved for use in training the second artificial intelligence model, and (ii) the output of the second artificial intelligence model being approved for provision to one or more computing systems, applying the second artificial intelligence model to generate an intended result.
[0275] Clause 98. The medium of clause 94, wherein determining the set of phrases corresponding to the user-specified query further includes parsing the user-specified query for a set of keywords, each keyword in the set of keywords being associated with a set of data objects; determining, for each keyword in the set of keywords associated with the set of data objects, a set of semantically similar phrases corresponding to each keyword in the set of keywords; and determining the set of phrases corresponding to the user-specified query using the set of semantically similar phrases corresponding to each keyword in the set of keywords.
[0276] Clause 99. The medium of clause 98, wherein determining semantically similar phrases corresponding to each keyword in the set of keywords further includes accessing a database indicating a mapping between the first keyword and the set of second keywords, and in response to accessing the database, using each keyword to determine a set of semantically similar phrases corresponding to each keyword.
[0277] Clause 100. The medium of clause 94, wherein accessing the metadata graph further includes traversing each node of the set of nodes to identify a metadata identifier that matches at least one phrase of the set of phrases, and determining a node corresponding to the set of phrases in response to determining that the metadata identifier matches at least one phrase of the set of phrases.
[0278] conclusion Unless the context clearly requires otherwise, throughout this description and claims, the words "comprises," "comprising," and the like, should be construed in an inclusive sense, i.e., "including, but not limited to," as opposed to an exclusive or exhaustive sense. As used herein, the terms "connected," "coupled," or any variation thereof, mean any direct or indirect connection or coupling between two or more elements, where the coupling or connection between the elements can be physical, logical, or a combination thereof. Additionally, the words "herein," "above," "below," and words of similar import, when used in this application, refer to this application as a whole and not to any particular portions of this application. Where the context permits, words in the above Detailed Description using singular or plural numbers can also include the plural or singular number, respectively. The word "or," in reference to a list of two or more items, covers all interpretations of the word: any of the items in the list, all of the items in the list, and any combination of the items in the list.
[0279] The above detailed description of examples of the present technology is not intended to be exhaustive or to limit the technology to the precise form disclosed above. Specific examples of the present technology are described above for illustrative purposes, but various equivalent modifications are possible within the scope of the present technology, as those skilled in the art will recognize. For example, while processes or blocks are presented in a given order, alternative implementations may perform routines having steps or employ systems having blocks in a different order, and some processes or blocks may be deleted, moved, added, subdivided, combined, and / or modified to provide alternative or subcombinations. Each of these processes or blocks may be performed in a variety of different ways. Also, while processes or blocks may be shown as being performed sequentially, these processes or blocks may instead be performed or executed in parallel, or may be performed at different times. Furthermore, any specific numbers referred to herein are merely examples, and alternative implementations may employ different values or ranges.
[0280] The teachings of the present technology provided herein can be applied to other systems, not necessarily those described above. The elements and acts of the various examples described above can be combined to provide further implementations of the present technology. Some alternative implementations of the present technology may include additional elements to those implementations mentioned above, as well as fewer elements.
[0281] These and other changes can be made to the present technology in light of the above detailed description. While the above description describes particular examples of the technology and explains the best mode contemplated, no matter how detailed it appears in the text, the technology can be practiced in many ways. While system details are still encompassed by the technology disclosed herein, its specific implementation may vary considerably. As noted above, specific terminology used when describing particular features or aspects of the technology should not be taken to suggest that the terminology has been redefined herein to be limited to any particular characteristic, feature, or aspect of the technology with which it is associated. In general, the terms used in the following claims should not be construed to limit the technology to the specific examples disclosed herein unless the detailed description section above explicitly defines such terms. Thus, the actual scope of the present technology encompasses not only the disclosed examples but also all equivalent ways of practicing or implementing the technology in the claims.
[0282] To reduce the number of claims, certain aspects of the present technology are presented below in certain claim forms, but the applicant contemplates various aspects of the technology in any number of claim forms. For example, while only one aspect of the present technology is recited as a computer-readable medium claim, other aspects may equally be embodied as a computer-readable medium claim or in other forms, such as means-plus-function claims. Any claim intended to be covered by 35 U.S.C. 112(f) will begin with the words "means for," but use of the term "for" in any other context is not intended to invoke 35 U.S.C. 112(f). Accordingly, the applicant reserves the right to pursue additional claims after filing this application to pursue such additional claim forms in either the present application or any continuing application. [Explanation of symbols]
[0283] 100 User Interface 102 User-specified query input 104 Result Output 106 At least one data object 108 Data Lineage Information 200 Computer Systems and Other Devices 204 Input Components 206 Output Components 208 processors 210 Storage 212a Application 212N Applications 214a model 214N model 216 Network Connection Components 218 Persistent Storage Devices 220 Computer-Readable Media Drive 300 Environment 302 Client Computing Device 302a Client Computing Device, Computing Device 302b Client Computing Device, Computing Device 302c Client computing device, computing device 302d Client Computing Device, Computing Device 304 Network 306 Server Computing Devices 310 Server Computing Device 310a Server 310b Server 310c Server 308 Database 312 databases 312a Database 312b database 312c database 400 processes 500 Metadata Graphs 502 nodes 502a node 502b node 502c node 502d node 502e node 502f node 502g node 502h node 502i node 502j node 502k nodes 502l node 502m node 504 Edge 504a Edge 504b Edge 504c Edge 504d Edge 504e Edge 504f Edge 504g Edge 504h Edge 504i Edge 504j Edge 504k Edge 504l Edge 504m Edge 504n Edge 504o Edge 504p Edge 504q Edge 600 Metadata Graph Close-Up 602 nodes 602a first node, node 602b second node, node 602c Third node, node 602d Fourth Node, Node 604 Edge 604a First Edge 604b Second Edge 604c Third Edge 606 File-Level Metadata Identifiers 606a File-Level Metadata Identifiers 608 Container-Level Metadata Identifier 608a Container-Level Metadata Identifier 610 Location Identifier 610a Location Identifier 700 Artificial Intelligence Model Diagram 702 model 704 Input 706 Output 800 processes 900 LLM Prompts 902 prompt 1 903 level 904 First Metadata Identifier 905 First prompt text 906 LLM 907 Second prompt text 908 First intermediate output 910 Prompt 2 912 Second intermediate output 914 Prompt 3 916 Prompt 4 918 Third prompt text 920 Metadata Graph 922 Prompt 5 1000 Subsystem Diagram 1002 User Interface 1004 Metadata Graph 1006 LLM 1008 Domain Ontology Components 1010 Raw Data Components 1012 Feedback Components 1014a Communication Link 1014b communication link 1014c communication link 1014d Communication Link 1014e communication link 1014f Communication Link 1014g communication link 1014h communication link 1014i communication link 1014j communication link 1014k communication link 1014l communication link 1014m communication link 1014n communication link 1014o communication link 1014p communication link 1015 Data Silos 1016 Crawler 1018 Parser 1020 Profiler 1022 Concept 1024 Domain Thesaurus 1026 Versioning Components 1028 Generator 1030 results 1032 Third Party Input 1034 Update Components 1100 Metadata Graph 1102 nodes 1102a 5th node, node 1102b 6th node, node 1102c Seventh Node, Node 1102d 8th Node, Node 1104a Second Edge 1106a File-level metadata identifier 1108a Container-level metadata identifier 1110a Location Identifier 1112a Domain-Specific Metadata Identifiers 1200 AI Sandbox 1210 Model Review Assistant 1220 Data Processor 1230 Model Generator 1240 Model Governor 1250 Model Auto Meter 1310 datasets, raw datasets 1320 Data Preprocessor 1322 Application Data, Application Data Set 1330 Data Analyzer 1340 Data Profile Generator 1350 Sampler 1342 Data Profile 1352 training samples 1354 Test Sample 1405 User Input 1410 Feature Set Generator 1420 Target Selector, Selector 1430 Model Selector 1440 Model Builder 1445 model 1510 Large-Scale Language Models (LLM) 1520 RAG-based process 1522 Industry Knowledge Repository 1524 Corporate Knowledge Repository 1530 Explainability Layer 1600 processes 1620 Process 1700 Chat Interface 1705 Natural Language Input 1710 User Input 1715 Plot 1720 input 1810 Model Expandable Engine, Engine 1812 Model Deployment Engine Updater 1814 Orchestrator 1816 Data Handler 1820A #1 Public Cloud Provider 1820B Second Public Cloud Provider 1830A On-Premise Machine 1830B On-Premise Elastic Resource 1900 processes
Claims
1. 1. A system for reducing computational resource usage when accessing siloed data across disparate locations via a unified metadata graph, comprising: at least one hardware processor; When executed by the at least one hardware processor, receiving, at a graphical user interface (GUI), a user-specified query indicating a request to access a set of data objects, wherein each data object of the set of data objects is stored in a respective data silo of a set of data silos among heterogeneous locations; parsing the user-specified query for a set of keywords, each keyword in the set of keywords being associated with the set of data objects; performing natural language processing on the set of keywords to determine a set of semantically similar phrases corresponding to each keyword in the set of keywords; accessing a metadata graph to determine nodes corresponding to the set of semantically similar phrases, the metadata graph comprising: (i) a set of nodes indicating (a) metadata of internal data objects stored in a data silo and (b) a location identifier of the data silo, and (ii) edges indicating data lineage between first and second nodes of the set of nodes, the metadata graph being generated using a metadata data structure based on file-level and container-level metadata identifiers; determining a data silo that stores the at least one data object of the set of data objects using the location identifier corresponding to the determined node to retrieve the at least one data object of the set of data objects via the data silo; and generating a visual representation of the at least one data object for display on the GUI, the visual representation of the at least one data object comprising lineage information of the at least one data object. at least one non-transitory memory storing instructions that cause the system to A system comprising:
2. the metadata graph comprises: retrieving from a second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier in the set of file-level metadata identifiers indicates metadata for a given data object stored in a respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates metadata for the respective data silo in the second set of data silos; generating a set of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, respectively; generating the metadata data structure for mapping each semantically similar metadata identifier in the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; generating the metadata graph using the generated metadata data structure; and The system of claim 1 , wherein the system is generated by:
3. receiving a second user-specified query via a second GUI indicating a request for generating an intended result; providing the second user-specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising (i) a second artificial intelligence model to be used to generate the intended result, and (ii) a second set of data objects to be used in training the second artificial intelligence model; In response to receiving a user selection indicating acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model, and (ii) obtaining the second set of data objects using the metadata graph; training the second artificial intelligence model using the set of data objects; applying the second artificial intelligence model to generate the intended result; and The system of claim 1 , further comprising the instructions for:
4. accessing a governance database to obtain a set of policies that dictate usage metrics corresponding to said second set of data objects; using the set of policies indicating usage metrics corresponding to the second set of data objects to determine whether the second set of data objects is approved for use in training the second artificial intelligence model; determining whether an output of the second artificial intelligence model is authorized to be provided to one or more computing systems using a second set of policies that dictate usage metrics corresponding to the artificial intelligence model predictions; and (i) in response to the second set of data objects being approved for use in training the second artificial intelligence model, and (ii) in response to the output of the second artificial intelligence model being approved for provision to the one or more computing systems, applying the second artificial intelligence model to generate the intended results; The system of claim 3 , further comprising instructions for:
5. 1. A computer-implemented method for reducing computational resource usage when accessing siloed data across disparate locations via a unified metadata graph, the method comprising: receiving a user-specified query at a graphical user interface (GUI) indicating a request to access a set of data objects, each data object of the set of data objects being stored in a respective data silo of a set of data silos among disparate locations; performing natural language processing on the user-specified query to determine a set of phrases corresponding to the user-specified query; accessing a metadata graph to determine nodes corresponding to the set of phrases, the metadata graph comprising: (i) a set of nodes comprising: (a) metadata indicating internal data objects stored in a data silo; and (b) a location identifier of the data silo; and (ii) edges indicating data lineage of the set of nodes, the metadata graph being generated using a metadata data structure based on file-level and container-level metadata identifiers; determining a data silo that stores the at least one data object of the set of data objects using the location identifier corresponding to the determined node, to retrieve the at least one data object of the set of data objects via the data silo; generating a visual representation of the at least one data object for display on the GUI; A method comprising:
6. the metadata graph comprises: retrieving from a second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier in the set of file-level metadata identifiers indicates metadata for a given data object stored in a respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates metadata for the respective data silo in the second set of data silos; generating a set of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, respectively; generating the metadata data structure for mapping each semantically similar metadata identifier in the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; generating the metadata graph using the generated metadata data structure; The method of claim 5 , wherein the compound is produced by
7. receiving a second user-specified query via a second GUI indicating a request for generating an intended result; providing the second user-specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising (i) a second artificial intelligence model to be used to generate the intended result, and (ii) a second set of data objects to be used in training the second artificial intelligence model; In response to receiving a user selection indicating acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model, and (ii) obtaining the second set of data objects using the metadata graph; training the second artificial intelligence model using the set of data objects; applying the second artificial intelligence model to generate the intended result; The method of claim 5 further comprising:
8. accessing a governance database to obtain a set of policies that dictate usage metrics corresponding to said second set of data objects; using the set of policies indicating usage metrics corresponding to the second set of data objects to determine whether the second set of data objects is approved for use in training the second artificial intelligence model; determining whether an output of the second artificial intelligence model is authorized to be provided to one or more computing systems using a second set of policies that dictate usage metrics corresponding to the artificial intelligence model predictions; (i) in response to the second set of data objects being approved for use in training the second artificial intelligence model, and (ii) in response to the output of the second artificial intelligence model being approved for provision to the one or more computing systems, applying the second artificial intelligence model to generate the intended result; The method of claim 7 further comprising:
9. determining the set of phrases corresponding to the user-specified query, parsing the user-specified query for a set of keywords, each keyword in the set of keywords being associated with the set of data objects; for each keyword in the set of keywords associated with the set of data objects, determining a set of semantically similar phrases corresponding to each keyword in the set of keywords; determining the set of phrases corresponding to the user-specified query using the set of semantically similar phrases corresponding to each keyword in the set of keywords; The method of claim 5 further comprising:
10. determining semantically similar phrases corresponding to each keyword in the set of keywords, accessing a database indicating a mapping between a first set of keywords and a second set of keywords; responsive to accessing the database, using each of the keywords to determine the set of semantically similar phrases corresponding to each of the keywords; 10. The method of claim 9, further comprising:
11. accessing the metadata graph includes: traversing each node in the set of nodes to identify a metadata identifier that matches at least one phrase in the set of phrases; determining the node corresponding to the set of phrases in response to determining that the metadata identifier matches the at least one phrase of the set of phrases; The method of claim 5 further comprising:
12. accessing the metadata graph includes: traversing each node in the set of nodes to identify a metadata identifier that matches at least one phrase in the set of phrases; determining a first node corresponding to the set of phrases in response to determining that the metadata identifier matches the at least one phrase of the set of phrases; responsive to determining that the first node corresponds to the set of phrases, performing a second traversal of the nodes of the set of nodes using an edge pointing to a first data pedigree of the first node, the first data pedigree of the first node pointing to a second node comprising information that is a source of information associated with the first node; determining a second data silo that stores the second data object of the set of data objects using the location identifier corresponding to the second node to retrieve the second data object of the set of data objects via the second data silo; generating a second visual representation of the second data object for display on the GUI; The method of claim 5 further comprising:
13. The method of claim 5 , wherein the visual representation of the at least one data object comprises lineage information of the at least one data object.
14. When executed by one or more processors, receiving, at a graphical user interface (GUI), a user-specified query indicating a request to access a set of data objects, wherein each data object of the set of data objects is stored in a respective data silo of a set of data silos among heterogeneous locations; performing natural language processing on the user-specified query to determine a set of phrases corresponding to the user-specified query; accessing a metadata graph to determine nodes corresponding to the set of phrases, the metadata graph comprising: (i) a set of nodes comprising: (a) metadata indicating internal data objects stored in a data silo; and (b) a location identifier of the data silo; and (ii) edges indicating data lineage of the set of nodes, the metadata graph being generated using a metadata data structure based on file-level and container-level metadata identifiers; determining a data silo that stores the at least one data object of the set of data objects using the location identifier corresponding to the determined node to retrieve the at least one data object of the set of data objects via the data silo; generating a visual representation of the at least one data object for display on the GUI; One or more non-transitory computer-readable media having stored thereon instructions that cause the computer to perform operations including:
15. the metadata graph comprises: retrieving from a second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier in the set of file-level metadata identifiers indicates metadata for a given data object stored in a respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates metadata for the respective data silo in the second set of data silos; generating a set of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, respectively; generating the metadata data structure for mapping each semantically similar metadata identifier in the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; generating the metadata graph using the generated metadata data structure; and The medium of claim 14 produced by
16. The instructions, when executed by the one or more processors, receiving a second user-specified query via a second GUI indicating a request for generating an intended result; providing the second user-specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising (i) a second artificial intelligence model to be used to generate the intended result, and (ii) a second set of data objects to be used in training the second artificial intelligence model; In response to receiving a user selection indicating acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model, and (ii) obtaining the second set of data objects using the metadata graph; training the second artificial intelligence model using the set of data objects; applying the second artificial intelligence model to generate the intended result; and The medium of claim 14 , further comprising:
17. The instructions, when executed by the one or more processors, accessing a governance database to obtain a set of policies that dictate usage metrics corresponding to said second set of data objects; using the set of policies indicating usage metrics corresponding to the second set of data objects to determine whether the second set of data objects is approved for use in training the second artificial intelligence model; determining whether an output of the second artificial intelligence model is authorized to be provided to one or more computing systems using a second set of policies that dictate usage metrics corresponding to the artificial intelligence model predictions; and (i) in response to the second set of data objects being approved for use in training the second artificial intelligence model, and (ii) in response to the output of the second artificial intelligence model being approved for provision to the one or more computing systems, applying the second artificial intelligence model to generate the intended results; The medium of claim 16 , further comprising:
18. determining the set of phrases corresponding to the user-specified query; parsing the user-specified query for a set of keywords, each keyword in the set of keywords being associated with the set of data objects; determining, for each keyword in the set of keywords associated with the set of data objects, a set of semantically similar phrases corresponding to each keyword in the set of keywords; determining the set of phrases corresponding to the user-specified query using the set of semantically similar phrases corresponding to each keyword in the set of keywords; The medium of claim 14 further comprising:
19. determining semantically similar phrases corresponding to each keyword in the set of keywords; accessing a database indicating a mapping between a first set of keywords and a second set of keywords; responsive to accessing the database, using each of the keywords to determine the set of semantically similar phrases corresponding to each of the keywords; The medium of claim 18 further comprising:
20. accessing the metadata graph, traversing each node in the set of nodes to identify a metadata identifier that matches at least one phrase in the set of phrases; determining the node corresponding to the set of phrases in response to determining that the metadata identifier matches the at least one phrase of the set of phrases; The medium of claim 14 further comprising:
Citation Information
Patent Citations
Prescribed navigation using topology metadata and navigation path
JP2006190261A
Graphic representation of data relationships
JP2011517352A
Generating, accessing, and displaying lineage metadata
JP2022033825A
Systems, methods, and apparatuses for executing a graph query against a graph representing a plurality of data stores
US20200226156A1
System and method for querying multiple data sources
US20220292092A1