Automating efficient deployment of artificial intelligence model

The integrated metadata graph system optimizes data access across heterogeneous locations by using natural language processing and large language models to generate efficient metadata graphs, addressing inefficiencies and resource wastage in existing systems, ensuring accurate and timely data retrieval.

JP2025111385AActive Publication Date: 2025-07-30CITIBANK N A

Patent Information

Application Number
JP2024225656
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-17
Filing Date
2024-12-20
Publication Date
2025-07-30
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Existing systems face inefficiencies and resource wastage in accessing siloed storage data across heterogeneous locations, leading to increased computing resource utilization, data integrity threats, and prolonged downtime due to the creation of new data silos and reconfiguration of computing systems.

Method used

An integrated metadata graph system leverages natural language processing and large language models to generate optimized metadata graphs, reducing the need for new data silos and reconfiguration by providing a unified access mechanism through a graphical user interface, utilizing domain-specific ontologies to enhance data search efficiency and accuracy.

Benefits of technology

This approach minimizes computing resource usage, reduces data search time, and maintains data integrity by eliminating the need for redundant data storage, thereby enhancing user experience and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111385000001_ABST
    Figure 2025111385000001_ABST
Patent Text Reader

Abstract

To provide facilitating a process for automatically deploying artificial intelligence (AI) models.SOLUTION: A system receives, for a first artificial intelligence (AI) model used by an entity, a first request to deploy a first AI model to make the first AI model available for use in a production environment to process input data and generate corresponding outputs. A first model deployment location for a first model is selected based on a model deployment engine. The system generates scripts to deploy the first AI model to the selected location, then monitors operations parameters associated with the deployment of the first AI model. Based on the values of the operations parameters, the system updates the model deployment engine. In response to a second request to deploy a second AI model, the system uses the updated model deployment engine to select a second model deployment location for a second model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims priority to U.S. Patent Application No. 18 / 390,916, filed on December 20, 2023, entitled "GENERATING A UNIFIED METADATA GRAPH VIA A RETRIEVAL - AUGMENTED GENERATION (RAG) FRAMEWORK SYSTEMS AND METHODS", which is a continuation - in - part of U.S. Patent Application No. 18 / 627,332, filed on April 4, 2024, entitled "AUTOMATING EFFICIENT DEPLOYMENT OF ARTIFICIAL INTELLIGENCE MODELS", and U.S. Patent Application No. 18 / 888,151, filed on September 17, 2024. This application also claims priority to U.S. Patent Application No. 18 / 617,305, filed on March 26, 2024, entitled "ACCESSING SILOED DATA ACROSS DISPARATE LOCATIONS VIA A UNIFIED METADATA GRAPH SYSTEMS AND METHODS", and U.S. Patent Application No. 18 / 888,088, filed on September 17, 2024, entitled "ARTIFICIAL INTELLIGENCE SANDBOX FOR AUTOMATING DEVELOPMENT OF AI MODELS". The contents of the foregoing applications are hereby incorporated by reference in their entirety into this specification.

Background Art

[0002] As computing systems become increasingly complex, the data used by such computing systems often requires its own data silos (databases, data warehouses, or data lakes) to efficiently process the data. Nevertheless, having each computing system have its own data silo can result in copies of data existing between different silos for each computing system. As a result, a large amount of computing capacity is required to read and maintain the single true version among the copies of data residing in different data silos distributed across one or more computer systems. Moreover, the data in one data silo may be similar to the data in other data silos. For example, although variable names and technical implementation forms may differ for each data silo because each computing system may require its own unique variable names, sequencing keys, and integrity constraints, the underlying data is the same. New applications are built using the latest technologies and techniques, but these quickly become obsolete with the emergence of newer and better implementation systems. These new and emerging silos compromise time and complexity in business processes, and building coordinated and integrated silos that can replace all existing silos requires a huge amount of effort to build and verify. Left unchecked, this results in similar data residing in multiple data silos that could otherwise be used to store new information, causing a further waste of a large amount of computing resources.

Brief Description of the Drawings

[0003]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16A

Figure 16B

Figure 17A

Figure 17B

Figure 17C

Figure 17D

Figure 18

Figure 19

Best Mode for Carrying Out the Invention

[0004] In the drawings, for the sake of discussion of some embodiments of the present technology, some components and / or operations may be separated into different blocks or combined into a single block. Moreover, the present technology is capable of accommodating various modifications and alternative forms, but a particular embodiment is shown by way of example in the drawings and will be described in detail below. Nevertheless, the intention is not to limit the present technology to the particular embodiments described. On the contrary, the present technology is intended to cover all modifications, equivalents, and alternative forms that fall within the scope of the present technology as defined by the appended claims.

[0005] To maintain data consistency between computing systems, modern computing systems can have data silos created to store data for a given computing system or software application. For example, each data silo can be configured with a unique variable name, access protocol (e.g., SQL, AMQP, etc.), data format (e.g., relational, non-relational, etc.), or other unique characteristics. Having a data silo specifically configured for a given computing system or software application not only allows the computing system / software application to communicate with a given data silo, but also makes it possible to maintain the data consistency of the data within the given data silo such that only the data within the silo can be modified, thereby protecting the data stored in other data silos.

[0006] Data silos bring such benefits, but these also result in many drawbacks. For example, one such drawback is that computing systems / software applications not configured to communicate with a given data silo are blocked by the data silo from obtaining or receiving data from this data silo. Since each data silo can be configured for a specific software application or computing system, when a new software application is built, or when a computing system is scaled up or down, data scientists must reconfigure the data silo or the software application / computing system. Another drawback is that a data silo may store the same or similar information about other data silos. For example, due to the configuration of such data silos (e.g., variable names, access protocols, data formats, or other characteristics), one data silo can store information associated with a first variable name, and another data silo can store the same information associated with a second variable name, where the first variable name and the second variable name are different. Although the variable names are different, the underlying data can be the same (or similar). This causes a large amount of computer memory to be wasted across computing systems because various copies of the data exist between different data silos. Yet another drawback is that it is often difficult to search for data that can be stored within a data silo due to its configuration. For example, since each data silo is separated from other data silos, there is no common interface for searching all available data silos at once, which requires the user to manually search each and every data silo repeatedly until the user finds the data that the user needs to obtain. This is not only a waste of time, but such repeated searches consume a large amount of wasted computing resources because it is required that hundreds, if not thousands, of queries be provided to each and every data silo. Data retrieval from distributed silos becomes even more complex when it is not known where the data is stored.

[0007] Existing systems have sometimes tried to address such drawbacks by leveraging computers and data scientists to create new data silos that can (i) remove copies of data and (ii) communicate with all computing systems / software applications that utilize such data. Still, the manual creation of new data silos is almost infeasible to implement. For example, due to the vast scale of modern computing systems, there can be hundreds, if not thousands, of data silos and corresponding computing systems / software applications that would need to be modified to communicate and utilize such data. Since such computing systems / software applications rely on large amounts of data stored within such data silos that are to be processed in real-time (or near real-time), reconfiguring such systems, applications, or data silos can lead to significant downtime of the computing systems, thereby impacting the user experience.

[0008] Furthermore, even when computers and data scientists manually create new data silos, there is a threat of impacting the data integrity of the data stored in the data silos. For example, when creating a new data silo, the computer / data scientist not only has to remove copies of the data but may also need to reformat the data so that the intended computing system / software application can effectively communicate with the data within the data silo. Such modifications to the data can corrupt the data and render such valuable data unusable. If the data stored within a given data silo is corrupted, even when a data scientist creates a copy of the data silo, this exacerbates the problem of wasted computer memory as even more copies of the data would need to be created.

[0009] Moreover, creating new data silos or reconfiguring existing computing systems / software applications creates further problems of wasting the computing resources (e.g., computer processing and computer memory resources) of a given system. For example, when each data silo, computing system, or software application has to be reconfigured / created, computing resources are wasted because each new data silo or new computing system / software application occupies a large amount of memory. Therefore, creating these new data silos, computing systems, or software applications exacerbates these problems.

[0010] For these and other reasons, there is a need to stop copying data and simplify data access patterns when accessing siloed storage data that spans heterogeneous locations via an integrated metadata graph. There is a further need to access siloed storage data that spans heterogeneous locations to enable user access to such siloed storage data without creating new data silos, databases, or reconfiguring existing computer systems and / or software applications. There is a further need to maintain data integrity of the data stored within the data silos without requiring that multiple copies of the data be stored within the data silos.

[0011] For example, as described above, existing systems lack a mechanism for accessing siloed storage data across heterogeneous locations without creating new computing components. Since existing systems rely on creating new data silos, databases, computing systems, software applications, and the like to access siloed storage data, such new computing components require a large amount of resources to effectively access the data. Further, since these existing systems rely on creating new data silos, the time and energy expended can lead to long-term computing system downtime. Moreover, since existing systems are prone to damaging data during the creation process of such computing components, existing systems rely on creating various copies of the data silos themselves, which can further exacerbate the problem of wasting valuable computer memory resources.

[0012] To overcome these and other deficiencies of existing systems, the inventors have developed a system and method for reducing the amount of computing resources used when accessing siloed storage data spanning heterogeneous locations via an integrated metadata graph. For example, the system can receive a user-specified query in a graphical user interface (GUI) that instructs a request to access a set of data objects, where each data object in the set of data objects is stored in a respective data silo in a set of data silos in heterogeneous locations. For example, the system can receive a user query to access data stored across various data silos. The system can then perform natural language processing on the user-specified query to determine a set of clauses corresponding to the user-specified query. For example, to enable a non-technically proficient user to access the data they desire, the system can determine a contextually accurate set of clauses (e.g., based on the user query) to provide the data that the non-technically proficient user is attempting to access.

[0013] The system then accesses the metadata graph to determine the nodes corresponding to the set of phrases. The metadata graph can comprise (i) a set of nodes comprising (a) metadata indicating internal data objects stored in a data silo, and (b) a location identifier of the data silo, and (ii) edges indicating data lineages between the set of nodes. For example, by using the metadata graph, the system can traverse a metadata graph that indicates where data (e.g., data objects) are stored and which data within different data silos are available. Thus, since the metadata graph can provide an abstraction layer about which data is stored where, data scientists do not need to create new data silos and / or reconfigure existing computing systems / software applications, thereby reducing the utilization of computing resources. Moreover, since the metadata graph includes data lineages between the set of nodes (e.g., representations of data stored within the data silo itself), the system can further provide information about where copies of data that the user intends to access can reside, which the system can utilize to efficiently find where the copied data is hosted. The system then determines the data silo storing at least one data object of the set of data objects by using the location identifier corresponding to the determined node to obtain at least one data object of the set of data objects via the data silo. The system then generates a visual representation of the at least one data object for display on the GUI. For example, the system can then provide data intended to be accessed by non-technically proficient users.Accordingly, by accessing siloed data by leveraging the capabilities of the metadata graph, the system can reduce the utilization of computing resources caused by generating new data silos, computing systems, or software applications to access data stored across different data silos in disparate locations.

[0014] Using a metadata graph reduces the data search time when accessing siloed data that spans heterogeneous locations (e.g., data silos hosted in various locations), but there is a further need to optimize the generation of such metadata graphs. For example, traditional approaches to searching for data can involve manually generating tables that contain the metadata of the siloed data, but generating these tables is inefficient because computer scientists first have to find the metadata, normalize the data (e.g., based on mere opinion), and then create the tables, wasting a large amount of computing resources (e.g., computer memory and processing power). Creating such tables is not only inefficient, but also error-prone considering the vast amount of data to be considered and the various copies of data inherent in many copies of data stored in different data silos. To reduce errors and overcome the inherent inefficiencies of traditional approaches, the inventors have developed optimized data structures (e.g., metadata graphs) that reduce data search time compared to parsing error-prone metadata tables. The inventors have further developed an optimized method for generating metadata that is less error-prone by leveraging large language models, the metadata itself, and domain-specific languages to reduce the time taken to generate such data structures while ensuring correct labeling of the metadata and enhancing accuracy.

[0015] For example, the system can select a first LLM prompt corresponding to a first metadata identifier from a set of metadata identifiers. The LLM prompt can correspond to the first metadata identifier based on a data profile of the metadata identifier (e.g., data schema, data format, etc.). The system can then expand the first LLM prompt with the first metadata identifier that is to be provided to the LLM, and the LLM is configured to generate a first intermediate output that indicates a second set of metadata identifiers corresponding to the first metadata identifier. For example, the system can provide the first metadata identifier to the first LLM prompt to cause the LLM to generate a set of semantically similar metadata identifiers. The set of semantically similar metadata identifiers can represent variations of the first metadata identifier (e.g., to "ask" the LLM what it thinks the first metadata identifier represents).

[0016] The system can then extend the first LLM prompt with a first intermediate output (e.g., a second set of metadata identifiers) to be provided to the LLM, and the LLM is configured to generate a second intermediate output that indicates filtered domain-specific metadata identifiers by accessing a set of domain-specific ontologies. For example, by providing the extended LLM prompt to the LLM (e.g., communicatively coupled to the set of domain-specific ontologies), the LLM can utilize the contextual knowledge provided by the domain-specific ontologies to generate normalized domain-specific metadata identifiers. The domain-specific ontologies include relationships between phrases, words, or descriptions of data that exist within the computing system of the entity, thereby providing a level of contextual knowledge to the entity. The LLM can utilize such contextual knowledge to generate filtered domain-specific metadata identifiers. Moreover, by using an LLM communicatively coupled to the domain-specific ontologies, the system can reduce the amount of computational resources required to generate the metadata graph by reducing the dataset of metadata identifiers to be considered (e.g., via access to the domain-specific ontologies).

[0017] The system can then generate a domain-specific integrated metadata graph via an LLM using (i) a first metadata identifier and (ii) a second intermediate output that indicates a filtered domain-specific identifier. For example, the filtered domain-specific metadata identifier may be a traversable identifier, and the first metadata identifier may be a non-traversable identifier within the domain-specific integrated metadata graph. By generating a domain-specific integrated metadata graph with traversable and non-traversable identifiers, the system can (e.g., by storing unfiltered non-domain-specific first metadata identifiers associated with filtered domain-specific metadata identifiers) reduce the amount of information to be traversed when identifying where data is (e.g., within data silos) while maintaining the verifiability and accuracy of the metadata graph, thereby reducing data search time. In this way, the system maintains data consistency of metadata from heterogeneous data silos by converting metadata into a verifiable metadata graph to efficiently search for and determine the available underlying data stored between data silos. Finally, to ensure data search time efficiency, the system determines the performance metrics of the generated domain-specific integrated metadata graph against the previous performance metrics of another version of the domain-specific integrated metadata graph. If the performance metrics of the generated domain-specific integrated metadata graph cannot meet the performance measure against the previous performance metrics of other versions of the domain-specific integrated metadata graph, the system performs an update process on the domain-specific integrated metadata graph. In this way, the system can ensure that the data search time is minimal and accurate when generation, update, or modification of the domain-specific integrated metadata graph occurs.

[0018] In various implementations, the methods and systems described herein can reduce the utilization of computing resources when accessing siloed storage data spanning heterogeneous locations via an integrated metadata graph. For example, the system can receive (e.g., via a GUI) a query that instructs a request to access a set of data objects, where each data object in the set of data objects is stored in a respective data silo in a set of data silos in heterogeneous locations. The system can perform natural language processing on the query to determine a corresponding set of clauses. The system can then access a metadata graph to determine nodes corresponding to the set of clauses, where the metadata graph comprises (i) a set of nodes comprising (a) metadata that indicates internal data objects stored in the data silos, and (b) location identifiers of the data silos, and (ii) edges that indicate data lineages of the set of nodes, and the metadata graph is generated using a metadata data structure based on file-level and container-level metadata identifiers. The system can then use the location identifiers corresponding to the determined nodes to determine the data silos storing at least one data object in the set of data objects in order to retrieve at least one data object in the set of data objects via the data silos. The system can then generate a visual representation of the at least one data object for display on the GUI.

[0019] In various implementations, the methods and systems described herein can reduce data search time when accessing siloed storage data across heterogeneous locations by generating an integrated metadata graph via a Retrieval-Augmented Generation (RAG) framework. For example, the system selects a first LLM prompt corresponding to a first metadata identifier from a set of metadata identifiers. The system then expands the first LLM prompt with the first metadata identifier to be provided to the LLM, and the LLM is configured to generate a first intermediate output that indicates a second set of metadata identifiers corresponding to the first metadata identifier. The system then expands the first LLM prompt with the second set of metadata identifiers corresponding to the first metadata identifier to be provided to the LLM, and the LLM is configured to generate a second intermediate output that indicates filtered domain-specific metadata identifiers by accessing a set of domain-specific ontologies. The system can then generate a domain-specific integrated metadata graph via the LLM using (i) the first metadata identifier and (ii) the second intermediate output that indicates the filtered domain-specific metadata identifiers. The filtered domain-specific metadata identifiers can be traversable identifiers, and the first metadata identifier can be a non-traversable identifier within the domain-specific integrated metadata graph. In response to determining that a first performance metric of the domain-specific integrated metadata graph fails to meet a performance measure with respect to a second performance metric of another version of the domain-specific integrated metadata graph, the system performs an update process on the domain-specific integrated metadata graph.

[0020] This domain-specific integrated metadata graph also enables the system to automatically construct or apply an artificial intelligence (AI) model. The AI sandbox according to the implementation forms in this specification provides a low-code or no-code environment in which data from heterogeneous locations, as represented by the metadata graph, is used to automatically generate an AI model for data analysis or to automatically apply an existing AI model to the data.

[0021] In some implementations, a computer system generates a dataset for training or applying an AI model. The computer system can receive a first natural language input from a user that includes a set of phrases and instructions for using an artificial intelligence (AI) model to analyze data associated with the set of phrases. In response to the first natural language input, the computer system accesses the metadata graph to determine nodes corresponding to the set of phrases, where the metadata graph includes (i) a set of nodes including (a) metadata indicating internal data objects stored in a data silo and (b) a location identifier of the data silo, and (ii) edges indicating data lineages of the set of nodes. The system processes the internal data objects indicated by the determined nodes to generate a first set of application data and applies the AI model to the first set of application data to generate one or more first outputs. For example, the first output can include a classification of data items within the first set of application data or a prediction made based on the first application data. The representation of the one or more outputs is sent for display to the user. Subsequently, a second natural language input can be received from the user, and the second natural language input includes instructions for modifying the first set of application data. Based on the second natural language input, the computer system generates a second set of application data. The AI model can then be applied to the second set of application data.

[0022] In some implementations, a computer system automates the deployment of an AI model. The computer system can receive, from an entity, a first request for a first artificial intelligence (AI) model for use in a production environment to process input data and generate a corresponding output, so that the first AI model can be deployed and utilized. Based on a model deployment engine, the computer system selects a first model deployment location for the first AI model, where the first model deployment location can be selected from a set of one or more cloud provider environments or an on-premises environment operated by the entity. The computer system generates a script for deploying the first AI model to the first model deployment location and, after deploying the model, monitors the operational parameters associated with the deployment of the first AI model at the selected model deployment location when the first AI model processes the input data and generates a corresponding output. The computer system can update the model deployment engine based on the monitored operational parameters and, in response to a second request for a second AI model, select a second model deployment location for the second AI model based on the updated model deployment engine.

[0023] In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present technology. However, it will be apparent to one of ordinary skill in the art that the implementations of the present technology may be practiced without some of these specific details.

[0024] Phrases such as "in some implementations", "in several implementations", "in some embodiments", "in the illustrated embodiments", "in other embodiments", and the like generally mean that the particular features, structures, or characteristics that follow the phrase are included in at least one implementation of the present technology and may be included in one or more implementations. Additionally, such phrases do not necessarily refer to the same or different implementations.

[0025] System Overview FIG. 1 illustrates a graphical user interface (GUI) representation for reducing the usage of computing resources when accessing silo storage data spanning heterogeneous locations via an integrated metadata graph, according to some embodiments of the present technology. For example, the user interface 100 can include a user-specified query input 102, a result output 104, a visual representation of at least one data object 106, and data lineage information 108 (e.g., 108a - 108b) of at least one data object. For example, the user-specified query input 102 can be a data field configured to receive a user-specified query as input. The user can provide a query for accessing data that can be stored across heterogeneous data silos of a computing system to the user-specified query input 102. The result output 104 can include one or more visual representations of at least one data object 106 and data lineage information 108 corresponding to at least one data object 106. In the context of a technically unsophisticated user attempting to find or otherwise access data that can be stored between a set of data silos related to one or more computing systems, the user interface 100 provides a mechanism for enabling the user to find the data that the user desires or needs.

[0026] Often, a user does not know which data silo (e.g., database) hosts the data the user intends to obtain, nor does the user know precisely which data may be needed for a given use. For example, a non-technically savvy user, such as a business user, may want a list of all the first names of users who were active last month. Thus, the user can provide a query instructing "I want all the first names of users who were active last month" to the user-specified query input 102, and the system can generate a result output 104. As will be explained later, the system can perform natural language processing on the user-specified query to obtain a set of terms (e.g., keywords, semantically similar terms, etc.) for searching the metadata graph. The metadata graph may be a graph indicating where the data is stored and which data is available. For example, when the user-specified query can be in a question format, the system can determine a set of terms for accessing the metadata graph by removing unnecessary terms in the user-specified query. The set of terms can be not only a "cleaned-up" version of the user-specified query, but can also be useful for targeting which data the user intends to obtain. By leveraging access to the metadata graph, the system can display the result output 104, which can include a visual representation of at least one data object 106 (e.g., the data the user is attempting to access, the location of the data the user is attempting to access, the format regarding how the data the user is attempting to access is stored, etc.), and can also include a visual representation of data lineage information 108 (e.g., where copies of the data or similar data can be stored, the format regarding how the data is stored, etc.).In this way, for users who are not technically proficient, an integrated and user-friendly user interface can be provided that improves the user experience while providing a central access point for accessing data stored between different data silos in different locations.

[0027] In some implementations, the visual representation of at least one data object 106 may be interactive. For example, the visual representation of at least one data object 106 may be a two-way link (e.g., a hyperlink) that enables a user to access data associated with at least one data object (e.g., by generating a visual representation of a table storing at least one data object, generating a window showing at least one data object, etc.) when the user selects the visual representation of at least one data object 106. In this way, the user is enabled to quickly and efficiently browse the data that the user intends to access.

[0028] Suitable computing environment FIG. 2 is a block diagram showing a part of components typically incorporated into at least a part of a computer system and other devices in which the disclosed system operates. In various implementations, these computer systems and other devices 200 can include server computer systems, desktop computer systems, laptop computer systems, netbooks, mobile phones, personal digital assistants, televisions, cameras, automotive computers, electronic media players, web services, mobile devices, wristwatches, wearables, glasses, smartphones, tablets, smart displays, virtual reality devices, augmented reality devices, and the like.In various implementations, a computer system and device include an input component 204 that includes a keyboard, microphone, image sensor, touch screen, button, touch screen, trackpad, mouse, CD drive, DVD drive, 3.5 mm input jack, HDMI (registered trademark) input connection, VGA input connection, USB input connection, or other computing input components; an output component 206 that includes a display screen (e.g., LCD, OLED, CRT, etc.), speaker, 3.5 mm output jack, light, LED, tactile motor, or other output-related components; a processor 208 that includes a central processing unit (CPU) for executing a computer program and a graphics processing unit (GPU) for executing a computer graphics program and handling computing graphical elements; a storage 210 that includes at least one computer memory for storing programs (e.g., applications 212a - 212N, models 214a - 214N, and other programs), data used while the programs are in use, including facilities and associated data, an operating system including a kernel, and device drivers; a network connection component 216 for the computer system to communicate with other computer systems, send, and / or receive data via the Internet or another network, and networking hardware such as switches, routers, repeaters, electrical cables and optical fibers, light emitters and light receivers, wireless transmitters and receivers, and the like; a permanent storage device 218 such as a hard drive or flash drive for permanently storing programs and data; and a computer-readable media drive 220 (e.g., at least one non-transitory computer-readable media), which is a tangible storage means that does not include a temporary propagation signal, for reading programs and data stored on a computer-readable medium, including zero or more of each of these.The computer system configured as described above is typically used to support the operation of equipment, but those skilled in the art will understand that the equipment can be implemented using devices of various types and configurations, as well as having various components.

[0029] Figure 3 is a system diagram illustrating an example of a computing environment in which the disclosed system operates in some implementations. In some implementations, environment 300 includes one or more client computing devices 302a - d, which can host, for example, metadata graph 500 (Figure 5) (or other system components). For example, computing devices 302a - d can each include distributed entities a - d. Client computing devices 302 operate in a networked environment using a logical connection through network 304 to one or more remote computers, such as server computing devices. In some implementations, client computing device 302 can correspond to device 200 (Figure 2).

[0030] In some implementation forms, the server computing device 306 is an edge server that receives client requests and coordinates the execution of these requests through other servers such as servers 310a - c. In some implementation forms, the server computing devices 306 and 310 comprise a computing system. Each of the server computing devices 306 and 310 is logically represented as a single server, but the server computing devices can each be a distributed computing environment that includes a number of computing devices in the same or geographically heterogeneous physical locations. In some implementation forms, each server computing device 310 corresponds to a group of servers. In some implementation forms, the server computing devices 306 and 310 host large language models, a set of domain-specific ontologies, artificial intelligence models, user interfaces, web servers, or other computing components.

[0031] The client computing device 302 and the server computing devices 306 and 310 can each function as a server or a client with respect to other servers or client devices. In some implementations, the server computing devices (306, 310a - c) are connected to corresponding databases (308, 312a - c). As described above, each server computing device 310 can correspond to a group of servers, and each of these servers can share a database or can have its own database (e.g., a data silo). The databases 308 and 312 store (e.g., hold) information such as predefined ranges, predefined thresholds, error thresholds, graphical representations, machine learning models, artificial intelligence models, natural language processing models, LLMs, LLM prompts, keywords, metadata graphs, location identifiers, lineage information, semantically similar phrases, file - level metadata identifiers, container - level metadata identifiers, system - level metadata identifiers, governance policies, usage metrics, machine learning model training data, artificial intelligence model training data, performance measurement criteria, data schemas, data profiles, or other information. In some implementations, the databases 308 and 312 may be data silos.

[0032] Although the databases 308 and 312 are logically shown as a single unit, the databases 308 and 312 can each be a distributed computing environment that includes multiple computing devices, can be within their corresponding servers, or can be in the same or geographically different physical locations.

[0033] Network 304 can be a Local Area Network (LAN) or a Wide Area Network (WAN), but can also be another wired or wireless network. In some implementations, network 304 is the Internet, or some other public or private network. Client computing device 302 is connected to network 304 through a network interface by means such as wired or wireless communication. Although the connections between server computing device 306 and server computing device 310 are shown as separate connections, these connections can be of any type of local, wide area, wired, or wireless network, including network 304, or a separate public or private network.

[0034] Accessing silo storage data across heterogeneous locations FIG. 4 is a flowchart illustrating a process 400 for reducing the usage of computing resources when accessing silo storage data spanning heterogeneous locations via an integrated metadata graph, according to some implementations of the present technology.

[0035] In act 402, process 400 receives a user-specified query that instructs a request to access a set of data objects. For example, the system receives, in a GUI, a user-specified query that instructs a request to access a set of data objects, and each data object in the set of data objects is stored in a respective data silo of a set of data silos in heterogeneous locations. A data object can be any object, piece of data, or information that can be stored in a data silo, such as a file, information contained within the file (e.g., first name, last name, e-mail address, home address, business address, financial information, account identifier, account number, value, percentage, ratio, alphanumeric string, sentence, etc.), table, data structure, or other data object.

[0036] (For example, the) data object that the user is attempting to access can be stored across various data silos (e.g., databases) within a computing environment (e.g., environment 300 (FIG. 3)). For example, a user may want to access account-related data for one or more user accounts. Nevertheless, the account-related data can be stored in one or more data silos within the computing environment. For example, one data silo can indicate how many accounts are currently open / active (e.g., the first data object), and another data silo can indicate the name of the user who opened the account (e.g., the second data object). The user may not be aware of where such data is, if available. Thus, the user can provide a user-specified query that indicates a request to access a set of data objects, and as will be described later, the system can return the data (e.g., the set of data objects) to the user. In this way, the system improves the user experience because the user can access such data without the need for prior knowledge of where the data can or cannot reside.

[0037] In Act 404, Process 400 can perform natural language processing to determine a set of phrases. For example, the system can perform natural language processing on a user-specified query to determine the set of phrases corresponding to the user-specified query. Since the data stored between data silos can include the same data (e.g., a copy of the data) or similar data, the system can determine the set of phrases corresponding to the user-specified query for efficiently searching the data stored between data silos. As an example, one data silo storing user account information such as the user's last name can store the user's last name as a variable called "last_name". Nevertheless, another data silo storing user account information can store the user's last name as a variable called "given_name". The data stored is the same (e.g., each silo stores the user's last name), but the variable names can be different. Therefore, when searching for data, the system can determine the set of phrases corresponding to the user-provided query for accessing the data.

[0038] In some implementations, the system determines a set of semantically similar phrases corresponding to the user-specified query. For example, the system parses the user-specified query for a set of keywords. The set of keywords can correspond to a set of data objects stored in the data silo. For example, the user provides a query (e.g., "I want the first names of all users who were active last month"). The system parses the user-provided query for a set of keywords (e.g., first name, active, etc.). For each keyword, the system can determine a set of semantically similar phrases.

[0039] For example, since data can be stored in different silos for each computer application across an entity's computing systems, the same or similar data can be stored in various formats. For example, a database storing a table of user account information can store the user's first name as the variable "name_first", "account_ID", "first_name", "name", or otherwise. Thus, the system determines a set of semantically similar phrases corresponding to each keyword in a set of keywords for searching a metadata graph to obtain the data intended for the user to receive.

[0040] The system can then use the set of semantically similar phrases corresponding to each keyword in the set of keywords to determine a set of phrases corresponding to the user-specified query. For example, continuing with the above example, if the user-specified query is "I want the first names of all users who were active last month", the system can determine a first set of semantically similar phrases for "first name" (e.g., "name_first", "account_ID", "first_name", "name") that would be used when accessing the metadata graph to determine nodes (e.g., metadata of data objects stored in silos and lineage data of such data objects). Thus, the system can more efficiently determine the location of the data needed based on a set of semantically similar phrases (e.g., when traversing the metadata graph), as opposed to being limited to a single phrase, keyword, or variable name, so the system can reduce the amount of computing resources used to access silo-stored data using the metadata graph.

[0041] In some implementations, the system can determine semantically similar phrases by accessing a database. For example, the database can indicate a mapping between a set of a first keyword and a second keyword. In some implementations, the database can store a set of predetermined keywords generated by subject matter experts (SMEs) in a particular field. In this way, the SMEs can create such a database to accurately determine which keywords are semantically similar to other keywords, thereby improving the accuracy with which semantically similar phrases are determined.

[0042] In some implementations, the database can be based on an artificial intelligence model. For example, based on a large number of user-specified queries, the amount of semantically similar phrases, and the unique data that can be searched within a data silo, the system can use the artificial intelligence model to determine a set of semantically similar phrases or generate a database to determine semantically similar phrases. The artificial intelligence model can be a machine learning model configured to receive keywords (e.g., phrases) as input and output a set of semantically similar keywords (e.g., semantically similar phrases). Due to the nature of the machine learning model (or other artificial intelligence model) having the ability to learn the associations between training data (e.g., labeled instances of keywords and semantically similar phrases), the model is not limited to a predefined set of keywords and phrases. For example, the machine learning model can generate new, undiscovered instances of semantically similar phrases corresponding to a given keyword that might not seem plausible to the human mind. Thus, the system can use the machine learning model to determine a set of semantically similar phrases corresponding to each respective keyword. In this way, since the machine learning model is not limited to a predetermined set of keywords, the system can determine more robust semantically similar phrases, thereby expanding the range of possible semantically similar phrases that can be generated.

[0043] In response to accessing the database, the system can determine a set of semantically similar sentences corresponding to each keyword by using each keyword. For example, the system can parse the database using each keyword to determine a match between (i) each keyword and (ii) the keywords in the database. When identifying a match, the system can obtain a set of semantically similar sentences corresponding to the keyword. Thus, the system can reduce the amount of computational resources used when determining semantically similar sentences by using matches, as opposed to performing natural language processing on each keyword to determine a set of semantically similar sentences.

[0044] In act 406, process 400 can access the metadata graph to determine nodes corresponding to the set of sentences. For example, the system can access the metadata graph to determine nodes corresponding to the set of sentences. The metadata graph can include (i) a set of nodes and (ii) edges indicating the data lineage of the set of nodes. The set of nodes can include (a) metadata indicating internal data objects stored in the data silo and (b) a location identifier of the data silo. By way of example, the metadata graph can be a graph data structure indicating the metadata of information stored in a set of data silos of environment 300 (FIG. 3).

[0045] As described above, when accessing data that may be stored in data silos at different locations, each data silo may be associated with a unique configuration for accessing the data stored within the data silo. When designing a computing system / software application, data scientists and computer scientists may carefully design the data silos, computing systems, and software applications to effectively communicate with each other via one or more communication protocols. Still, this creates scalability issues when scaling the computing system because the data required for a given computing system / software application may become inaccessible due to the configuration of either the computing system / software application or the data silo itself. Further, the information stored in one data silo may be the same underlying data as that in another data silo, even though it has different variable names (e.g., variable identifiers, metadata identifiers, etc.), making it difficult to search for the required data. When searching for such required data for a given computing system / software application, existing systems can parse each and every data silo available for a given match between the data stored within the data silo and the data that is intended to be accessed (e.g., the required data). Still, parsing each and every data silo in the environment wastes valuable computer processing and memory resources, which is caused by determining whether there is a match between each and every data silo and the information stored in the data silo.

[0046] To address these technical deficiencies, accessing a metadata graph to determine nodes corresponding to a set of clauses (e.g., clauses corresponding to a user-specified query, keywords, alphanumeric strings) is utilized to quickly and efficiently identify and access data, thereby reducing the amount of computing resources used.

[0047] Referring to FIG. 5 showing an illustrative representation of a metadata graph, the metadata graph 500 can include a set of nodes 502a - 502m and edges 504a - 504q. Each node 502 can be linked or connected to one or more other nodes via one or more edges 504. Each node 502 can indicate metadata of internal data objects stored within a given data silo (e.g., file-level metadata), metadata of the data silo itself (e.g., container-level metadata), and metadata of one or more data silos such as a location identifier of a given data silo (e.g., a location where the data silo is located such as a compute component node identifier, a server identifier, etc.). Each edge 504 can indicate a data lineage of a set of nodes. For example, each edge can represent a lineage relationship between a first node and a second node. That is, each edge can indicate whether a node is a data source of another node or a derivative of another node.

[0048] For example, FIG. 6 shows an enlarged view of a metadata graph. In some implementations, the enlarged view 600 of the metadata graph may correspond to a portion of the metadata graph 500 according to some implementations of the present technology. As an example, the first node 602a can indicate the metadata of one or more data objects stored within a data silo. For example, the first node 602a can include a file-level metadata identifier 606a, a container-level metadata identifier 608a, and a location identifier 610a of the data silo. The file-level metadata identifier 606a may be a variable name, a file name, or other identifier that indicates one piece of data stored within a given silo. For example, the file-level metadata identifier may be any identifier that describes the data stored within a file stored in a data silo (e.g., a variable identifier, a file format, a file size, an access time stamp, or other file-level metadata). The container-level metadata identifier 608a may be an identifier that identifies the format that the data stored in a given data silo can take (e.g., a table format, a spreadsheet format, a graphical format, a dictionary, etc.), one or more configurations of the data silo (e.g., a communication protocol, accessibility parameters, etc.), or other container-level metadata associated with a given data silo. The location identifier 610a may be an identifier that indicates the location of a given data silo. For example, the location identifier can indicate a computer node with which the data silo is associated (e.g., stored, hosted, connected, etc.), a computer system with which the data silo is associated, the location of a server on which the data silo is hosted or otherwise associated, or other location identifier.

[0049] Each node of the set of nodes (e.g., nodes 602a - 602d) can have its own file - level metadata identifier 606, container - level metadata identifier 608, or location identifier 610, respectively. Since each node of the set of nodes can represent an abstract diagram of how data is derived from each other, where the data is, and which data is available, the system can utilize the metadata graph to efficiently find where the data is, along with the lineage information of the data itself. That is, a node can represent an abstracted diagram of how data is stored across data silos, the relationships between the data stored in the data silos, and where the data is stored between the data silos. For example, the first node 602a can be linked to the second node 602b via the first edge 604a. In some implementations, the first edge 604a can indicate the lineage information of the nodes, such as when the second node 602b is the data source of the first node 602a. Still, in other implementations, the first edge 604a can indicate the lineage information, such as when the first node 602a is the data source of the second node 602b according to some implementations of the present technology. It should be recognized by those skilled in the art that each node 602 may be linked to other nodes via an edge 604, and each edge can indicate the lineage information between one or more of the nodes of the set of nodes.By representing data objects (e.g., data stored in data silos) through a metadata graph that indicates (i) where the data object is (e.g., where it is stored in a data silo), (ii) the metadata of the data object itself, (iii) the metadata of the data silo storing the data object, and (iv) the location of such a data silo, the system can more efficiently traverse the metadata graph to access data stored in data silos at heterogeneous locations, as opposed to existing systems that rely on manually parsing each and every data silo for matches between the data the user is trying to access and the data stored within the silos themselves, thereby reducing the amount of computational resources used when accessing siloed data spanning heterogeneous locations.

[0050] In some implementations, the system can determine nodes corresponding to a set of clauses by traversing the metadata graph. For example, the system can traverse each node in a set of nodes of the metadata graph. The system can compare the metadata identifier of a given node with each clause in the set of clauses. For example, the metadata identifier can be a file level, container level, or other identifier indicating that a given data silo contains data regarding the clause. For example, the metadata identifier can be "first_name" (e.g., a file level metadata identifier) indicating that the data silo contains the user's first name. In response to determining that the metadata identifier matches at least one clause in the set of clauses, the system can determine the node corresponding to the set of clauses.

[0051] In contrast to traversing a metadata graph using a single term, for example, the system traverses the metadata graph and compares each term in a set of terms to the metadata identifier of a given node. That is, in contrast to existing techniques for traversing a graph (e.g., a metadata graph or other graph) using a given keyword, the system traverses the graph using a set of terms. In this way, the system can more efficiently determine the nodes corresponding to the set of terms since the system does not need to perform multiple traversals of the graph using a different term each time, thereby reducing the amount of computing resources used.

[0052] In some implementations, the system can determine another data silo that stores a second data object. For example, the system can traverse each node in a set of nodes (e.g., of a metadata graph) to identify a metadata identifier that matches at least one term in a set of terms. In response to determining that a metadata identifier matches at least one term in the set of terms, the system determines a first node corresponding to the set of terms. Still, even though the system may have determined a first node corresponding to the set of terms (e.g., thereby determining a data silo that stores a data object associated with the set of terms), the system can continue to traverse the metadata graph to determine other locations (e.g., of a data silo) that host a given data object.

[0053] For example, if a user-specified query instructs "I want all locations where the user's first name resides", the system can continue to traverse a set of nodes using the edges connected to a given node. For example, in response to determining that a first node corresponds to a set of phrases, the system can perform a second traversal of the nodes in the set of nodes to determine a second node using an edge that instructs a first data lineage of the first node. The first data lineage of the first node can instruct a second node that includes information that is a source of information associated with the first node. For example, each edge of the metadata graph can instruct a lineage of data objects. Since each node in the set of nodes instructs metadata (e.g., data about data), the edges between nodes can instruct that one node is a source (or alternatively, a derived data source of another node) of another node.

[0054] For illustration, referring to FIG. 6, the system can determine a first node corresponding to a set of phrases, such as the first node 602a. The system can traverse to a second node 602b using a first edge 604a, to a third node 602c using a second edge 604b, or to a fourth node 602d using a third edge 604c. In some implementations, after performing a first traversal (e.g., a traversal from the first node 602a to the second node 602b), the system can perform a second traversal (e.g., from the first node 602a to the third node 602c). The system can repeat such traversals until each node is traversed or until there are no remaining nodes corresponding to the set of phrases after the traversal is performed. In this way, the system can efficiently access siloed data by traversing the metadata graph, as opposed to parsing each and every data silo in the computing environment for a match.

[0055] Accordingly, the system can determine a second data silo that stores a second data object (e.g., the same data object or a similar data object related to at least one data object) by obtaining the second data object of a set of data objects via the second data silo using a location identifier corresponding to the second node. That is, the system can determine alternative locations (e.g., data silos) where a given data object can be stored by traversing the metadata graph using an edge connected to a determined node. In this way, the system can determine all locations where the same or similar data can be stored. In some implementations, the system can then generate a visual representation of the second data object on the GUI. In this way, additional data that the user is interested in can be provided to the user.

[0056] In some implementations, in response to determining each data silo where a given data object is stored, the system can perform one or more data aggregation techniques. For example, the system can remove unnecessary instances of the data itself. For example, since the metadata graph is an abstraction that indicates where the data is and which data a given silo can contain, the system can remove all but one instance of the data (e.g., data object) to reduce the amount of computer memory utilized.

[0057] In some implementations, process 400 can generate a metadata graph using the generated metadata data structure. For example, the system can extract from each data silo within a given environment (e.g., environment 300): (i) a set of file-level metadata identifiers, and (ii) a set of container-level metadata identifiers. Each file-level metadata identifier in the set of file-level metadata identifiers indicates the metadata of a given data object stored within its respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates the metadata of each respective data silo in the set of data silos within a given environment. The system can generate a set of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, respectively. For example, the system can perform natural language processing on the file-level and container-level metadata identifiers, respectively, to determine a set of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier. For example, for a file-level metadata identifier of "first_name", the system can generate a set of semantically similar metadata identifiers such as "name_first", "account_ID", "user_id", "name", or others.

[0058] The system can then generate a metadata data structure for mapping each semantically similar metadata identifier among a set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier. For example, to enable the system to efficiently search data spanning a metadata graph, the system can generate a normalized metadata identifier corresponding to each of the semantically similar phrases (e.g., by using natural language processing, machine learning models, artificial intelligence models, etc.). For example, the normalized metadata identifier for a set of semantically similar metadata identifiers such as "name_first", "account_ID", "user_id", and "name" can be "first_name_ID", and the metadata data structure maps "first_name_ID" to each of the semantically similar metadata identifiers. In some implementations, the system can use the generated metadata data structure (e.g., normalized metadata identifiers, sets of semantically similar metadata identifiers, etc.) to generate a metadata graph. Additionally or alternatively, the system can generate a metadata graph based on an artificial intelligence model. In this way, the system can optimize the metadata graph by using the normalized container-level and file-level metadata identifiers associated with the nodes of the metadata graph to enable more efficient data searching.

[0059] Referring again to FIG. 4, in act 408, process 400 can determine a data silo storing at least one data object. For example, the system can use a location identifier corresponding to a determined node (e.g., of act 406) to determine a data silo among a set of data silos storing at least one data object of a set of data objects in order to obtain at least one data object of the set of data objects via the data silo. Since each node among the set of nodes of the metadata graph includes a location identifier corresponding to a data silo (e.g., indicating which data silo stores a given data object), the system can access the data silo using the location identifier to obtain a data object. For example, the system can use the location identifier of the determined node to determine which data silo hosts at least one data object of the set of data objects. In some implementations, the system can use location identifiers corresponding to other determined nodes to determine each data silo storing each data object of the set of data objects in order to obtain the set of data objects. Using the location identifier, the system can determine a communication protocol associated with the determined data silo to obtain at least one data object. For example, since each data silo can be associated with a unique communication protocol, the system can identify which communication protocol the determined data silo is associated with, select this communication protocol (e.g., query language, access protocol, configuration, etc.) to communicate with the data silo, and provide a query to the data silo. Thus, the system can obtain at least one data object via the query. In this way, the system can reduce the amount of computational resources wasted when accessing silo-stored data spanning heterogeneous locations via the metadata graph.

[0060] In act 410, process 400 can generate a visual representation of at least one data object for display. For example, the system can generate a visual representation of at least one data object for display on a GUI. In some implementations, the visual representation of at least one data object includes lineage information of at least one data object. For example, referring to FIG. 1, the visual representation of at least one data object 106 can be presented for display within user interface 100 along with lineage information 108 corresponding to the at least one data object. One of ordinary skill in the art should recognize that user interface 100 can include one or more visual representations (e.g., of data objects and / or lineage information) that can correspond to a set of data objects according to one or more implementations of the present technique.

[0061] In some implementations, the system can use an artificial intelligence model to generate an intended result. For example, the system can receive, via a second GUI, a second user-specified query that instructs a request to generate an intended result. For example, a user can provide a query that instructs a request to generate an intended result (e.g., a prediction) using an artificial intelligence model. The intended result can be any user-specified prediction that the user wants to receive. In the context of a non-technically-savvy user, such a user may be ignorant of which artificial intelligence model / machine learning model to select to generate a given prediction, which data to use to train a given artificial intelligence model / machine learning model, or which other components / data to use to generate a given prediction. Nevertheless, the user may know what the user wants to discover (e.g., how many accounts will be opened within the next three months, in which week the company is likely to receive the influx of opened accounts, what the expected cost is to monitor fraud for a given set of accounts over a given period, how many users / accounts are active, how many users / accounts are inactive, etc.). To enable such a non-technically-savvy user to obtain an intended result, the system can provide a GUI (which can be the same GUI or similar to the GUI described in FIG. 1) that enables the user to provide a query to generate an intended result, and can provide recommendations on which artificial intelligence model / machine learning model to use to generate the intended result and which training data to use to train the artificial intelligence model / machine learning model to generate the intended result.

[0062] The system can provide a second user-specified query to an artificial intelligence model to generate a recommendation, and the recommendation includes (i) a second artificial intelligence model that is to be used to generate the intended result, and (ii) a second set of data objects that is to be used when training the second artificial intelligence model. As an example, the system can provide a user-specified query to an artificial intelligence model (e.g., a machine learning model, model 702 (Figure 7)) trained to generate recommendations. The artificial intelligence model can generate a recommendation indicating a given artificial intelligence model to be used to generate the intended result and which training data the given artificial intelligence model should be trained with to generate the intended result. For example, the recommended artificial intelligence model can be an artificial intelligence model or a machine learning model that can be configured to generate the intended result. Such a recommended artificial intelligence model / machine learning model can be a deep learning model, a neural network, a convolutional neural network, a recurrent neural network, a support vector machine, a natural language processing model, a KNN model, a linear regression model, a logistic regression model, a random forest model, a Bayesian model, or other artificial intelligence / machine learning mode. Thus, the system provides recommendations on which artificial intelligence model to use to generate the intended result and which training data to use to train the artificial intelligence model, thereby reducing the utilization of computing resources that would otherwise be wasted by users who are not technically proficient in performing the many incorrect iterations of training a machine learning model to generate the intended result.

[0063] In response to receiving a user selection that instructs acceptance of a recommendation, the system can, according to some implementations of the present technology, (i) access a database to obtain a second artificial intelligence model, and (ii) use a metadata graph to obtain a second set of data objects. For example, the system can generate a message (such as a notification, a user-selectable object, etc.) to enable the user to accept the recommendation (e.g., via a button, a text-based command, a checkbox, etc.). In some implementations, the system can automatically accept a recommendation even without a user selection to accept the recommendation. In this way, the system can automatically select a recommended artificial intelligence model and training data for generating the intended result, thereby improving the user experience. The system can then access a database (such as an artificial intelligence model database) storing an untrained or pre-trained artificial intelligence / machine learning model and obtain the recommended artificial intelligence model (e.g., via an artificial intelligence model identifier, a machine learning model identifier, etc.). The system can also access the metadata graph to obtain a second set of data objects (which, for example, will be used as training data for the recommended artificial intelligence / machine learning model). For example, the second set of data objects can be training data stored within one or more data silos of the environment 300 that will be used as training data for the artificial intelligence model. In response to obtaining the recommended artificial intelligence model and the second set of data objects, the system can use the second set of data objects (e.g., training data) to train the recommended artificial intelligence model and apply the recommended artificial intelligence model (e.g., to input data) to generate the intended result. For example, the system can provide new input data (such as new data obtained via the metadata graph) as input to the recommended artificial intelligence model for generating the intended result (e.g., based at least in part on a user-specified query).In this way, users who are not technically proficient are enabled to use an artificial intelligence model to generate one or more intended results, thereby improving the user experience.

[0064] In some implementations, the system can determine whether the output of the artificial intelligence model is approved to be provided to one or more computing systems. For example, since artificial intelligence models and machine learning models are used in various domains related to entities (e.g., companies, businesses, etc.), the use of such artificial intelligence models / machine learning models may be required to comply with one or more governance standards when using such models for one or more functions. Since non-technically proficient users can use such models to generate predictions, discover new relationships between existing data, or for other functions, the system can ensure that the use of such models, the data provided to the models, and the output generated by the models comply with one or more industry, government, or internal standards. In this way, the system can reduce the likelihood of data breaches, thereby improving data security.

[0065] For example, the system can access a governance database to obtain a set of policies that dictate usage metrics for a set of data objects. The governance database can store policies (e.g., governance policies, industry standards, internal company policies, etc.) that dictate usage metrics (e.g., definitions or other metrics regarding how data can be used, generated, provided to other computing systems, provided to an external computing environment, published, etc.). The system can access the governance database to obtain a set of policies that dictate usage metrics for a second set of data objects (e.g., data used to train a recommended artificial intelligence model), and can use the set of policies to determine whether it is approved for the second set of data objects to be used to train a recommended artificial intelligence model. For example, in some implementations, the system can provide (i) the second set of data objects, and (ii) the obtained set of policies (e.g., corresponding to the second set of data objects) to another artificial intelligence / machine learning model (e.g., model 702 (FIG. 7)) configured to generate a prediction as to whether it is approved for the second set of data objects to be used to train a second artificial intelligence model. The system can also determine whether the output of the second artificial intelligence model (e.g., the recommended artificial intelligence model) is approved to be provided to one or more computing systems using a second set of policies that dictate usage metrics corresponding to the artificial intelligence model prediction. For example, the second set of policies can include information regarding which types of artificial intelligence model predictions can be sent, provided, published, or otherwise sent to internal or external computing systems.(i) It is approved that a second set of data objects is used to train a second artificial intelligence model, and (ii) in response to it being approved that the output of the second artificial intelligence model is provided to one or more computing systems, the system can apply the second artificial intelligence model (e.g., the recommended artificial intelligence model) to generate the intended result. In this way, before generating the intended result, the system can examine the training data and the output that can be generated by the artificial intelligence model, thereby reducing the likelihood of data infringement caused by providing such output to one or more computing systems.

[0066] Referring to FIG. 7, FIG. 7 shows a diagram 700 of an artificial intelligence model according to some implementations of the present technology. Model 702 can take in input 704 and provide output 706. The input can include a number of data sets, such as a training data set and a test data set. Each of the plurality of data sets (e.g., input 704) can include a data subset related to user data, predicted expectations and / or errors, and / or actual expectations and / or errors. In some embodiments, output 706 may be fed back to model 702 as input for training model 702 (e.g., alone or with a user indication regarding the accuracy of output 706, a label associated with the input, or other reference feedback information). For example, the system can receive a first labeled feature input, and the first labeled feature input is labeled with a known prediction for the first labeled feature input. The system can then train a first machine learning model to classify the first labeled feature input with the known prediction (e.g., a response to a user-provided query).

[0067] In various implementations, the model 702 can update its configuration (e.g., weights, biases, or other parameters) based on the evaluation of its predictions (e.g., output 706) and reference feedback information (e.g., user instructions regarding accuracy, reference labels, or other information). In various implementations, if the model 702 is a neural network, the combined weights can be adjusted to reconcile the difference between the neural network's prediction and the reference feedback. In further use cases, one or more neurons (or nodes) of the neural network may require that their respective errors be sent backward through the neural network (e.g., backpropagation of errors) to facilitate the update process. The update to the combined weights can, for example, reflect the magnitude of the error propagated backward after the forward pass is complete. Thus, for example, the model 702 can be trained to produce better predictions.

[0068] In some implementations, model 702 can include an artificial neural network. In such implementations, model 702 can include an input layer and one or more hidden layers. Each neural unit of model 702 can be connected to many other neural units of model 702. Such connections can be such that the effect on the activation state of the connected neural units can be either forcing or inhibitory. In some implementations, each individual neural unit can have a summary function that combines all the values of its inputs. In some implementations, each connection (or the neural unit itself) can have a threshold function such that a signal must exceed it before propagating to other neural units. Model 702 may be self - learning and trained rather than being explicitly programmed and can perform quite well in certain areas of problem - solving compared to traditional computer programs. During training, the output layer of model 702 can correspond to the classification of model 702, and inputs known to correspond to this classification can be input into the input layer of model 702 during training. During testing, inputs without known classifications may be input into the input layer, and a determined classification may be output.

[0069] In some implementations, model 702 can include a number of layers (e.g., the signal path traverses from the front layer to the back layer). In some implementations, backpropagation techniques may be utilized by model 702, and forward stimuli are used to reset the weights for the "front" neural units. In some implementations, the stimuli and inhibitions for model 702 can be more fluid, and the connections interact in a more disorderly and complex manner. During testing, the output layer of model 702 can indicate whether a given input corresponds to the classification of model 702 (e.g., a response to a user - provided query).

[0070] In some implementations, a model (e.g., model 702) can automatically perform an action based on output 706. In some implementations, the model (e.g., model 702) may not perform any action. The output of the model (e.g., model 702) can be used to generate a metadata graph, determine a set of phrases, determine semantically similar phrases, provide recommendations for an artificial intelligence / machine learning model, determine whether a data object is approved for use in training an artificial intelligence / machine learning model, determine whether an artificial intelligence / machine learning model output is approved for being provided to one or more computing systems, generate a response, or generate other information, in accordance with one or more implementations of the present technology, or be used to instruct to do so otherwise.

[0071] In some implementation forms, a model (e.g., model 702) can be trained based on training information stored in database 308 or database 312 to generate recommendations. For example, the recommendations can be recommendations for a given artificial intelligence / machine learning model to generate an intended result, and recommendations on which training data should be used when training a given artificial intelligence / machine learning model. Model 702 can take in a first set of training information as input 704 and generate an output (e.g., one recommendation, multiple recommendations) as output 706. The first set of training information can include a user-specified query indicating a request to generate an intended result (e.g., a prediction), an artificial intelligence / machine learning model identifier used to generate the intended result, training data used to train the artificial intelligence / machine learning model used to generate the intended result, or other information. For example, model 702 can learn the associations among the first set of training information to generate recommendations as output 706. Output 706 can also be recommendations on which artificial intelligence model should be selected to generate the intended result and which training data should be used to train the artificial intelligence model to generate the intended result. In some embodiments, output 706 can be fed back to model 702 to update one or more configurations (e.g., weights, biases, or other parameters) based on the evaluation of model 702 of the prediction of model 702 (e.g., output 706) and reference feedback information (e.g., user instructions regarding accuracy, reference labels, ground truth information, known recommendations, etc.). The first set of training information can also be historical training information that has been used to train a previous artificial intelligence / machine learning model to generate a given intended result.In this way, the model 702 is trained to generate one or more recommendations as to which artificial intelligence / machine learning model can generate the intended result, and the training data required to train such an artificial intelligence model / machine learning model, thereby enabling a non-technically proficient user to utilize the artificial intelligence / machine learning model.

[0072] In some implementations, a model (e.g., model 702) can be trained to determine approvals based on training information stored in database 308 or database 312. For example, model 702 can be trained to determine whether training data for a given artificial intelligence / machine learning model is approved for use in training the artificial intelligence / machine learning model and whether the output of the artificial intelligence / machine learning model is approved to be disclosed, transmitted, or provided to one or more computing systems. For example, as described above, with the increasing use of artificial intelligence and machine learning models in business contexts, such models are under scrutiny and must be vetted before being applied to sensitive user data. To vet such models, model 702 can take a second set of training information as input 704 and generate an output (e.g., one approval, multiple approvals) as output 706. The second set of training information can include predictions generated by the artificial intelligence / machine learning model, artificial intelligence / machine learning model identifiers used to generate the predictions, training data used to train the artificial intelligence / machine learning model used to generate the predictions, a set of policies indicating usage metrics corresponding to data objects (e.g., training data) used to train the artificial intelligence / machine learning mode used to generate the predictions, a second set of policies indicating usage metrics corresponding to artificial intelligence model predictions, or other information. For example, model 702 can learn associations between the second set of training information to generate an approval as output 706. Output 706 can be an approval indicating whether a second set of data objects (e.g., training data) is approved for use in training the artificial intelligence / machine learning model and whether the output (e.g., prediction) of the artificial intelligence model / machine learning model is approved to be provided to one or more computing systems.In some embodiments, the output 706 can be fed back to the model 702 to update one or more configurations (e.g., load, bias, or other parameters) based on an evaluation of the model 702's predictions (e.g., output 706) and reference feedback information (e.g., user instructions regarding accuracy, reference labels, ground truth information, known recommendations, etc.). The second set of training information may be historical information that has been used to provide recommendations for different data objects (e.g., training data) and machine learning models. Thus, the model 702 can be trained to scrutinize the artificial intelligence model / machine learning model, its input data, its training data, and its output data prior to being used, according to one or more implementations of the present technology.

[0073] Generating an integrated metadata graph FIG. 8 illustrates a process for generating an integrated metadata graph via a retrieval augmented generation (RAG) framework, according to some implementations of the present technology.

[0074] At act 802, the process 800 selects a first LLM prompt. For example, the system selects a first LLM prompt corresponding to a first metadata identifier from a set of metadata identifiers from a set of LLM prompts. Each LLM prompt in the set of LLM prompts can be associated with a data schema, data format, data type, or other characteristic of the metadata identifier. For example, an LLM prompt associated with a data type of a metadata identifier can refer to a file-level metadata identifier, a container-level metadata identifier, a system-level metadata identifier, or other metadata identifiers. For example, the type of the metadata identifier can instruct the structure of the LLM prompt that will be selected for use.

[0075] LLM prompts can be structured with respect to the data type of metadata identifiers. For example, a structured LLM prompt can refer to an input configured to be interpreted by an LLM in a structured format. A structured LLM prompt is a prompt for a text-to-text language model (e.g., an LLM) that is structured such that the text included in the structured LLM prompt is interpreted and understood by the LLM. Since each LLM prompt can be structured for a data schema, data format, data type, data profile, or other characteristics of metadata identifiers, the LLM prompt can provide additional information to the LLM when generating an output. For example, a structured LLM prompt structured for the data type of a metadata identifier can include one or more attributes, tags, labels, or other information that indicates that the metadata identifier included in the LLM prompt is of a particular type. In this way, the LLM can generate a more accurate first intermediate output that indicates a set of metadata identifiers corresponding to the first metadata identifier.

[0076] Referring to FIGS. 9A - 9B, which illustrate examples of LLM prompts according to some implementations of the present technology, an exemplary LLM prompt 900 can include prompt 1 902, prompt 2 910, prompt 3 914, prompt 4 916, and prompt 5 922. Additionally, for illustration purposes, an LLM 906 is shown. Prompt 1 902 can include level 903, a first metadata identifier 904, and a first prompt text 905. For example, level 903 can refer to the data schema, data format, data type, or other characteristics of the first metadata identifier 904 to provide additional information to the LLM 906 about what kind or type of metadata identifier the LLM should consider. The first prompt text 905 can be structured text associated with level 903. For example, the first prompt text 905 can be unique to level 903. For example, the first prompt text 905 is shown to instruct "provide a set of similar identifiers", and can be text corresponding to level 903, where level 903 indicates a file - level metadata identifier type and the first metadata identifier 904 indicates a file - level metadata identifier. In some implementations, the first prompt text 905 can vary based on the metadata identifier of the first metadata identifier 904. For example, if the first metadata identifier 904 is a container - level metadata identifier, the first prompt text 905 can instead enumerate "provide a set of similar container - level identifiers", and level 903 indicates the first prompt text 905. That is, when determining the type of the metadata identifier of the first metadata identifier 904, prompt 1 902 can be selected. Prompt 1 902 is associated with level 903 that indicates the type of the metadata identifier, and the first metadata identifier 904 further includes the correct first prompt text 905.In this way, an LLM prompt can be structured based on the data schema, data format, data type, or other characteristics of the first metadata identifier in order to obtain more accurate results from the LLM, in contrast to general LLM prompts for existing systems that do not rely specifically on the generated LLM prompt.

[0077] To generate a metadata graph, the system can utilize the RAG framework. For example, the RAG framework, or alternatively, RAG, can refer to a framework that enables an artificial intelligence model (e.g., a large language model) to access a data source that may contain information to be updated without requiring retraining of the entire LLM. Traditionally, an LLM is trained on a large corpus of data to provide an output based on an input prompt. Nevertheless, an LLM is often limited to the training data on which it is trained, and the training process for an LLM is exceptionally computationally intensive. To overcome such drawbacks of an LLM, the RAG method can be adopted to ensure that up-to-date data is provided to the LLM without requiring complete retraining of the LLM.

[0078] Moreover, using RAG, the LLM can query for additional information that the LLM has not been previously trained on. For example, the LLM can often produce outputs that may seemingly and factually appear correct, but the LLM has no mechanism to decipher between what is true and what is not. Rather, the LLM defines that the output interpreted by the LLM is the most correct output for the input (e.g., the prompt). To provide a mechanism that not only allows the LLM to access information that the LLM has not been previously trained on, but also enables the LLM to provide a source of truth (e.g., verifiable information on which the LLM can base its output generation), the LLM can be communicatively coupled to one or more data sources. For example, in the context of generating a metadata graph via the RAG framework, the LLM can be communicatively coupled to a set of domain ontologies of entity domain ontology components 1008 for returning raw data components 1010 (e.g., for "retrieving" metadata identifiers) and filtered domain-specific metadata identifiers generated (based on an extended LLM prompt that includes the retrieved metadata identifiers).

[0079] For example, referring to FIG. 10A, FIG. 10A shows a subsystem diagram 1000 that illustrates an example of a RAG framework environment for generating an integrated metadata graph. The subsystem diagram 1000 can provide an example of a RAG framework environment according to some implementations of the present technology. The RAG framework environment can include a user interface 1002, a metadata graph 1004, an LLM 1006, a domain ontology component 1008, a raw data component 1010, a feedback component 1012, and communication links 1014a - 1014p. For example, the user interface 1002 is a user interface for receiving a user - specified query that instructs a request to access a set of data objects (as described, for example, in act 402 of FIG. 4). The metadata graph 1004 can be a metadata graph (as described, for example, in act 406 of FIG. 4). The LLM 1006 can be any LLM (such as BERT, Claude, Cohere, Ernie, Falcon40B, Galactica, etc.) configured to provide an output in response to an input. For example, the LLM 1006 can receive an LLM prompt and be configured to provide a response (such as text, graphical, etc.) in response to the LLM prompt.

[0080] Communication links 1014a - 1014p can enable communication between the user interface 1002, the metadata graph 1004, the LLM 1006, the domain ontology component 1008, the raw data component 1010, the feedback component 1012, or other components (shown or not shown). For example, communication links 1014a - 1014p can include the Internet, a mobile phone network, a mobile voice or data network (e.g., 5G or LTE network), a cable network, the public switched telephone network, or other types of communication networks, or combinations of communication networks. Communication links 1014a - 1014p can separately or together include one or more communication paths such as satellite paths, optical fiber paths, cable paths, paths that support Internet communication (e.g., IPTV), free - space connections (e.g., for broadcasts or other wireless signals), or any other suitable wired or wireless communication path, or combinations of such paths.

[0081] The domain ontology component 1008 can be a database, server, or other computing component configured to store a set of domain ontologies regarding entities. For example, the domain ontology component 1008 can store a domain ontology that indicates a set of concepts and categories in a given subject area (e.g., domain) that provides information about the nature of the concepts / categories and the relationships between the concepts / categories. According to one or more implementations of the present technology, the domain ontology component 1008 can store a set of domain ontologies specific to the entities of the system. For example, if the entity is a company, the domain ontology can reflect domain - specific knowledge (e.g., terminology, taxonomy, morphology) about the terms used in the entity's domain. For example, if the entity is a bank, the domain ontology component 1008 can include an ontology that relates financial terms to other financial terms to infer the context in which a given financial term is used.

[0082] Such domain-specific (e.g., entity-specific) contextual knowledge can be used to generate normalized and filtered domain-specific metadata identifiers that conform to the terminology system used by the entity in its daily operations, and is thus advantageous for utilization in generating a metadata graph. By communicatively connecting an LLM to a set of domain ontologies specific to the entity, the system can generate normalized and filtered domain-specific metadata identifiers that will be used in generating the metadata graph. By doing so, the system can enable the users of the system to extract, identify, and retrieve data that they intend to retrieve from the metadata graph based on a common terminology system, context, or domain. Moreover, by leveraging the domain ontology associated with the entity, the system can reduce errors when generating normalized and filtered domain-specific metadata identifiers as the entity's terminology system and contextual knowledge are maintained. For example, different entities may have different meanings for a given term. By leveraging the domain-specific ontology, the system can reduce errors when generating normalized and filtered domain-specific metadata identifiers that will be used in the metadata graph as the LLM can "refer" to a set of domain-specific ontologies to validate the LLM's output (e.g., the normalized and filtered domain-specific metadata identifier). In addition to validating the output of the LLM (which may include the metadata graph 1004 itself), the feedback component 1012 can be used to update, validate, or verify the metadata graph 1004 (e.g., additions thereto) during the metadata graph generation process. For example, the feedback component 1012 can include one or more user inputs or automated inputs for validating the accuracy of the metadata graph 1004 (to be described in more detail later).

[0083] The raw data component 1010 can be a data source that provides raw data to the LLM. For example, when generating the metadata graph 1004, the raw data component 1010 provides metadata identifiers, (e.g., metadata identifiers, data silos, data objects, system) data profiles, or other raw data to the LLM 1006. For example, the raw data component 1010 can obtain raw data from the data silos of the system. For example, the system can receive raw data with a set of metadata identifiers that indicate (i) file-level metadata identifiers, (ii) container-level metadata identifiers, or (iii) system-level metadata identifiers, from a set of data silos. A file-level metadata identifier can indicate the metadata of a data object stored in one of the data silos in the set of data silos, a container-level metadata identifier can indicate the metadata of one of the data silos in the set of data silos, and a system-level metadata identifier can indicate the metadata of the computing system that hosts one of the data silos in the set of data silos. For example, a file-level metadata identifier can indicate the label of a data object stored in a data silo, a container-level metadata identifier can indicate the label of the data format in which the data silo stores data, and a system-level metadata identifier can indicate the label of the operating system or the system identifier that hosts the data silo.

[0084] In some implementations, the system can perform a crawling process across a set of data silos. For example, to obtain raw metadata, the system can perform a crawling process across a set of data silos associated with entities (e.g., companies, merchants, legal persons, businesses, computing environments, etc.) to obtain raw data with a set of metadata identifiers. For example, referring to FIG. 10B showing a subsystem diagram of raw data component 1010, the crawling process may be performed by crawler 1016, which can be any database crawling service configured to extract metadata from data silo 1015. Data silo 1015 may be the same or similar to the data silos described in acts 402-410 of process 400 (FIG. 4). For example, crawler 1016 can generate a set of crawl queries to obtain file-level, container-level, or system-level metadata (e.g., metadata values, metadata identifiers, etc.). Parser 1018 can parse a set of crawl queries to extract file-level, container-level, or system-level metadata identifiers. In this way, the system can obtain all available metadata from data silos for use in generating a more robust and accurate metadata graph, as opposed to existing methods that rely on manual labeling techniques. In some implementations, the metadata is obtained via a combination of crawler 1016 and parser 1018, and via manually labeled data entities.

[0085] In some implementations, the system can generate data profiles for each data silo in a set of data silos. When generating a domain-specific integrated metadata graph, a data silo itself can store various data with different schemas, types, and formats, and can also have different contexts. Profiling such data silos is advantageous because these data profiles can indicate valuable context information that can impact the given structure of an LLM prompt, thereby impacting the ultimate results received by the LLM. For example, each LLM prompt can be a key, particularly for obtaining (e.g., normalized metadata identifiers with a particular context or domain) to achieve the intended result. Thus, when providing a prompt to the LLM, the structure of the prompt can include various data elements to achieve more efficient and accurate results. As an example, an LLM prompt extended with a metadata identifier and the data type corresponding to this metadata identifier can result in more accurate results, as opposed to an LLM prompt with only the metadata identifier (e.g., because additional context information may be lacking). Therefore, the system can generate data profiles for each data silo in a set of data silos to extend or select a structured LLM prompt for processing.

[0086] For example, the parser 1018 can extract a first value from each data silo of a set of data silos. Since each data silo can store a unique set of data, the system only needs to extract at least one value from each of the set of data silos. Nevertheless, in other implementations, the system can extract one or more values from each data silo of the set of data silos. The profiler 1020 can then determine the data type corresponding to each first value extracted from each of the set of data silos. For example, the profiler 1020 can be a logical component that can determine the data type corresponding to the first value. The data type can relate to the data schema of the first value, the format of the first value, whether the first value is an integer, a character, a floating point number, a double-precision floating point number, or some other data type. Using the data type of the first value, the profiler 1020 can generate a data profile for each data silo of the set of data silos that indicates the data type of the values stored in the data silo. For example, the system can generate a data profile (e.g., a file, a text file, a tag, etc.) associated with each data silo (e.g., a container) within the computing system of the entity that indicates the data type of the values stored in each of the data silos. Such data profiles can be stored in a database for later retrieval and can be associated with their respective data silos. In this way, the system can create an index for the data types associated with each data silo to accurately select a structured LLM prompt with context information (e.g., a data profile).

[0087] In some implementations, to select a first LLM prompt from a set of LLM prompts, the system can filter the set of LLM prompts. For example, as described above, the data profile (e.g., data type) of each data silo can add context information that is advantageous to use when selecting structured LLM prompts to generate a domain-specific integrated metadata graph. For example, by augmenting LLM prompts specifically designed with context information (e.g., the data profile of a data silo) from which the metadata identifier arises, the system can add such context information, rather than relying on the learned knowledge of the LLM itself, and can achieve more accurate results as opposed to existing systems.

[0088] Accordingly, the system can determine the data silo that stores the data corresponding to the first metadata identifier. For example, the system can compare the first metadata identifier to each metadata identifier stored in each of the data silos for a match. In other implementations, still, the system can refer to a database that stores a mapping between the metadata identifier and the data silo storing the data associated with the metadata identifier. The system can then retrieve the data profile corresponding to the data silo that stores the data corresponding to the first metadata identifier. For example, as described above, the system can retrieve the generated data profile for the data silo.

[0089] The system can then filter a set of structured LLM prompts to generate a set of filtered LLM prompts using the retrieved data profile. For example, each LLM prompt in the set of LLM prompts can be tagged with one or more tags indicating (i) a metadata identifier, (ii) a data profile (e.g., data type), (iii) the architecture of the LLM prompt, and / or (iv) other tags (e.g., data schema, data format, or other characteristics). The system can filter the set of structured LLM prompts to a subset of LLM prompts (e.g., a filtered set of LLM prompts) to reduce the amount of computing resources utilized when comparing LLM prompts. Filtering not only reduces the utilization of computing resources (e.g., computer memory and processing power) but can also provide a reduced set of LLM prompts to select from based on the data profile of the data silo associated with the metadata identifier, thereby improving LLM prompt selection accuracy. The system can then select a first structured LLM prompt corresponding to a first metadata identifier from the set of filtered LLM prompts. For example, the system selects the first structured LLM prompt based on a match between the tag of the LLM prompt indicating the data format of the LLM prompt and the data format of the metadata identifier.

[0090] Referring back to FIG. 8, in act 804, process 800 extends the first LLM prompt with a first metadata identifier. For example, the system extends the first LLM prompt (e.g., a structured LLM prompt) with the first metadata identifier that is to be provided to the LLM. The LLM can be configured to generate a first intermediate output that indicates a second set of metadata identifiers corresponding to the first metadata identifier. For example, the LLM can be communicatively coupled to (i) raw data components, and (ii) domain ontology components, and the LLM is configured to generate a first intermediate output that indicates a second set of metadata identifiers corresponding to the first metadata identifier without accessing the domain ontology components.

[0091] For example, referring to FIG. 10A, LLM 1006 is communicatively coupled to both raw data components 1010 and domain ontology components 1008. The LLM is communicatively coupled to each of raw data components 1010 and domain ontology components 1008, but the LLM can communicate with raw data components 1010 to generate a first intermediate output. For example, the first intermediate output can be an intermediate output such that the first intermediate output is not the final output of the LLM. For example, consistent with the RAG framework, the system can extend the first LLM prompt with the first metadata identifier that is to be provided to LLM 1006 to generate a second set of metadata identifiers corresponding to the first metadata identifier.

[0092] Referring back to FIG. 9, for example, prompt 1902 may reflect a first LLM prompt. The system can extend (e.g., add, update, place, etc.) prompt 1902 with a first metadata identifier (e.g., first metadata identifier 904). The extended version of prompt 1902 can then be provided by the system as an input to LLM 906 to generate a first intermediate output 908. For example, LLM 906 may be the same as or similar to LLM 1006 (FIG. 10A) according to some implementations of the present technology. LLM 906 processes prompt 1902 to generate a first intermediate output 908. The first intermediate output 908 can be a set of metadata identifiers corresponding to the first metadata identifier 904. For example, to determine what the metadata identifier (e.g., first metadata identifier 904) means, what it could be, or what it is similar to, the system can provide the first metadata identifier to an LLM that will receive a pre-generated set of metadata identifiers, descriptions, narratives, or other information that corresponds to or is otherwise associated with the first metadata identifier. In some implementations, the first intermediate output may be the same as or similar to semantically similar phrases as described in act 404 of process 400 (FIG. 4). By doing so, the system generates a set of semantically similar phrases corresponding to the first metadata identifier, thereby expanding the scope of context information that will be considered by the LLM to later generate more accurate normalized domain-specific metadata identifiers that will serve as keys to the unique entities.

[0093] Referring back to FIG. 8, in act 806, process 800 extends the first LLM prompt with a set of metadata identifiers corresponding to the first metadata identifier. For example, the system extends the first LLM prompt with a second set of metadata identifiers corresponding to the first metadata identifier that is to be provided to the LLM. The LLM can be configured to generate a second intermediate output that indicates filtered domain-specific metadata identifiers by accessing a set of domain-specific ontologies. For example, referring to FIG. 9A, the system extends prompt 2 910 with a first intermediate output 908 corresponding to the first metadata identifier 904 that is to be provided as input to LLM 906. In some implementations, prompt 2 910 may be the same or similar to prompt 1 902, but in other embodiments, prompt 2 910 may be different from that of prompt 1 902. For example, prompt 3 914 can represent a single prompt that combines the information of prompt 1 902 and prompt 2 910 into a single updatable prompt. That is, in contrast to having two separate prompts to achieve a given goal, the system can extend the prompt multiple times with respect to receiving respective outputs from LLM 906.

[0094] For example, referring back to prompt 2910, prompt 2910 can include a second prompt text 907 and a first intermediate output 908. The second prompt text 907 may be structured text associated with level 903 or the first intermediate output 908. For example, the second prompt text 907 may be unique to level 903. For example, the second prompt text 907 shown to instruct "return the domain-specific identifier for" may be text corresponding to level 903, where level 903 indicates a file-level metadata identifier type and the first metadata identifier 904 indicates a file-level metadata identifier. In some implementations, the second prompt text 907 may vary based on the metadata identifier of the first metadata identifier 904. For example, if the first metadata identifier 904 is a container-level metadata identifier, the second prompt text 907 can alternatively enumerate "return the domain-specific for the container-level identifier of", and level 903 indicates the second prompt text 907. That is, when determining the type of the metadata identifier of the first metadata identifier 904, prompt 2910 may be selected. Prompt 2910 is associated with level 903 that indicates the type of the metadata identifier, and the first metadata identifier 904 further includes the correct second prompt text 907. In this way, the LLM prompt can be structured based on the data schema, data format, data type, or other characteristics of the first metadata identifier in order to obtain more accurate results from the LLM, in contrast to general LLM prompts of existing systems that do not rely on specifically generated LLM prompts.

[0095] Additionally or alternatively, the second prompt text 907 may be associated with the first intermediate output 908. For example, the second prompt text 907 may be expanded to the prompt 2 910 when the system receives the first intermediate output 908. For example, the system may change, update, expand, or otherwise alter the prompt 1 902 to reflect the prompt 2 910 (e.g., including the second prompt text 907 and the first intermediate output 908). An illustrative example is provided by the prompt 3 914. The prompt 3 914 may be a combined prompt (e.g., of the prompt 1 902 and the prompt 2 910). In some implementations, the prompt 3 914 may be the resulting prompt. For example, when the system originally selects the prompt 1 902 that is to be provided to the LLM to generate the first intermediate output 908, the system can only provide the information of the prompt 1 902 to the LLM to generate the first intermediate output 908. When the system receives the first intermediate output from the LLM, the system can expand the original prompt (e.g., the prompt 1 902) to generate the prompt 3 914 that includes the second prompt text 907 and the first intermediate output 908. In some implementations, the system can provide the entire prompt 3 914 to the LLM906 to generate a second intermediate output 912 that indicates a filtered domain-specific metadata identifier by accessing a set of domain-specific ontologies. In yet other implementations, the system can provide only the new information of the prompt 3 914 to the LLM906 to generate the second intermediate output 912. For example, the system can only provide the second prompt text 907 and the first intermediate output 908 to the LLM906 to generate the second intermediate output 912, thereby reducing the amount of computational resources required by the LLM to process the input data (e.g., prompt information).

[0096] The second intermediate output 912 may be a filtered domain-specific metadata identifier. For example, in order to reduce the data search time when accessing data stored in various heterogeneous data silos through users who are not technically proficient, it is necessary to maintain domain-specific information in the context of the entity's system that enables the user to quickly search for the data the user needs without the burden of knowing the correct terminology system of the data. For example, a user who is not technically proficient may try to search for the name of an account. Nevertheless, due to the fact that computer engineers, data scientists, and other more technically proficient users are the ones who set up, create, or otherwise maintain the data silos, values, identifiers, terms, or other data markers may be different from what non-technically proficient users are familiar with. Non-technically proficient users have business thinking and can understand the domain-specific language officially used by the entity, but computer engineers and data scientists often do not understand it, and as a result, label the data without considering the domain-specific context of the business (e.g., of the entity). To overcome this, the system can provide the LLM communicatively coupled to the domain ontology component 1008 with a prompt 2 910 (or, alternatively, a prompt 3 914) to generate the second intermediate output 912 (e.g., a filtered domain-specific metadata identifier).

[0097] For example, referring to FIG. 10, the LLM 1006 can be provided with an extended LLM prompt that instructs a second set of metadata identifiers (e.g., the first intermediate output) to generate a second intermediate output that instructs filtered domain-specific metadata identifiers by accessing the domain ontology component 1008. The LLM 1006 can extract one or more domain ontologies, including a set of domain ontologies specific to the entities of the system. For example, as described above, if the entity is a company, the domain ontology can reflect domain-specific knowledge (e.g., terminology, taxonomy, morphology) about the terms used in the entity's domain. For example, if the entity is a bank, the domain ontology component 1008 can include an ontology that relates financial terms to other financial terms to infer the context in which a given financial term is used. The LLM 1006 can use one or more domain-specific ontologies included in the domain ontology component 1008 to validate the first intermediate output (e.g., the second set of metadata identifiers corresponding to the first metadata identifiers). For example, the LLM 1006 can compare each metadata identifier in the second set of metadata identifiers with keywords, phrases, character strings, or other domain-specific values of the domain-specific ontology to determine (i) the meaning of each metadata identifier in the second set of metadata identifiers, or (ii) the filtered domain-specific metadata identifiers.

[0098] Since LLM1006 can be an unsupervised artificial intelligence model, LLM1006 can be trained to determine the meaning of each metadata identifier in a second set of metadata identifiers by accessing a domain-specific ontology. The domain-specific ontology may be a pre-defined ontology generated by one or more domain experts for a given entity. LLM1006 can determine filtered domain-specific metadata identifiers by accessing the domain-specific ontology. For example, during a comparison process (e.g., where the LLM compares or otherwise processes a first intermediate output), LLM1006 can determine that the first intermediate output (e.g., the second set of metadata identifiers corresponding to the first metadata identifier) corresponds to (e.g., is associated with, matches, etc.) a common filtered domain-specific metadata identifier that exists within the domain-specific ontology. For example, the domain-specific metadata identifiers are "filtered" because they are filtered to a single representative domain-specific metadata identifier corresponding to a potential match to the first intermediate output as generated via LLM1006. In this way, the system can reduce the amount of computational resources included when generating the metadata graph since the filtered domain-specific metadata identifiers are used to generate the metadata graph.

[0099] Referring to FIG. 10C showing a subsystem diagram of the domain ontology component 1008, the domain ontology component 1008 can be communicatively coupled to the domain thesaurus 1024 and the concept 1022. For example, the domain ontology component 1008 can host a set of domain-specific ontologies generated at least in part based on domain experts, although the domain thesaurus 1024 and the concept 1022 can contribute to the generation of the domain ontology component 1008 of the domain-specific ontology. The concept 1022 can include a set of entity-specific terms, phrases, or other concepts that are commonly used throughout the system of entities (e.g., FIG. 3). The thesaurus 1024 can include a data structure that maps a set of entity-specific terms, phrases, or concepts to other terms, phrases, or other concepts used throughout the system of entities. For example, the thesaurus 1024 can represent a digital thesaurus of terms, phrases, or concepts. In some implementations, the ontology component 1008 can aggregate the information stored in the thesaurus 1024 and the concept 1022 to automatically generate one or more domain-specific ontologies (e.g., via one or more ontology creation models). Additionally or alternatively, the ontology component 1008 can utilize SMEs to create a set of domain-specific ontologies. Thus, the system can maintain the accuracy of the domain-specific context on which the metadata identifiers depend when generating the metadata graph, thereby maintaining the domain-specific language of the system of entities.

[0100] In act 808, process 800 generates a metadata graph. For example, the system can generate a domain-specific integrated metadata graph via an LLM using (i) a first metadata identifier and (ii) a second intermediate output that indicates a filtered domain-specific metadata identifier. The LLM can be configured to generate a graph (e.g., an undirected graph, a directed graph, a directed acyclic graph, etc.) using the first metadata identifier and the filtered domain-specific metadata identifier. In some implementations, the LLM may be provided with a prompt that instructs it to generate a graph (e.g., a metadata graph), and the prompt includes the first metadata identifier, the filtered domain-specific metadata identifier, and prompt text that instructs to generate a metadata graph. In some implementations, acts 802-808 may be repeatedly repeated until all of the metadata in data silo 1015 (FIG. 10B) is processed by the system.

[0101] Referring to FIG. 9B, the prompt 4916 can include a third prompt text 918 that instructs to generate a graph 920, a first metadata identifier 904, and a second intermediate output 912. The third prompt text 918 may be structured text associated with level 903 or the second intermediate output 912. For example, the third prompt text 918 may be unique to level 903, the first metadata identifier 904, and the second intermediate output 912, and the LLM906 is for generating a domain-specific integrated metadata graph based on level 903, the first metadata identifier 904, or the second intermediate output 912. The LLM can generate the metadata graph 920 using at least a portion of the information included in the prompt 4916. For example, the LLM can be trained to generate a graph data structure including the act 406 of process 400 (FIG. 4) and, as described in FIGS. 5-6, the first metadata identifier, the filtered domain-specific metadata identifier, and / or other information (e.g., file-level metadata identifier, container-level metadata identifier, system-level metadata identifier, location identifier, data lineage, etc.).

[0102] In some implementations, it is possible to provide a prompt 5922 to the LLM, and the prompt 5922 can represent a single prompt that combines the information of prompt 1902, prompt 2910, and prompt 4916 into a single updatable prompt. That is, in contrast to having three separate prompts to achieve a given goal, the system can expand the prompt multiple times with respect to receiving each output from the LLM906. When a prompt (e.g., prompt 4916 or prompt 5922) is provided as input to the LLM906, the LLM906 can generate the metadata graph 920.

[0103] Referring to FIG. 11 which shows an illustrative representation of the generated metadata graph, the LLM 906 can generate a metadata graph 1100. The metadata graph 1100 may be the same as or similar to the metadata graph 920 (FIG. 9B), the metadata graph 1004 (FIG. 10A), the metadata graph 500 (FIG. 5), or the metadata graph 600 (FIG. 6). The metadata graph 1100 can represent a domain-specific integrated metadata graph according to some implementations of the present technology. The metadata graph 1100 can include nodes 1102a-1102d, such as a fifth node 1102a, a sixth node 1102b, a seventh node 1102c, and an eighth node 1102d. Each of the nodes 1102a-1102d can indicate the metadata of one or more data objects stored within a given data silo. For example, the fifth node 1102a can include a file-level metadata identifier 1106a, a container-level metadata identifier 1108a, a location identifier 1110a, and a domain-specific metadata identifier 1112a. Additionally or alternatively, the fifth node 1102a can include a system-level metadata identifier or other information, not shown. The domain-specific metadata identifier 1112a may be the same as or similar to the second intermediate output 912 that indicates a filtered domain-specific metadata identifier.

[0104] Each of nodes 1102a - 1102d can be linked to one or more other nodes. For example, the fifth node 1102a can be linked to the sixth node 1102b via the second edge 1104a. In some implementations, the second edge 1104a can indicate lineage information of the nodes, such as when the sixth node 1102b is a data source of the fifth node 1102a. Still, in other implementations, the second edge 1104a can indicate lineage information, such as when the fifth node 1102a is a data source of the sixth node 1102b, according to some implementations of the present technology. Each node 1102 may be linked to other nodes via edges 1104, and it should be recognized by those skilled in the art that each edge can indicate lineage information among one or more of the nodes in a set of nodes.

[0105] In some implementations, one or more of the identifiers included within nodes 1102a - 1102d are traversable. To efficiently traverse metadata graph 1100, the system can traverse the metadata graph based on a single traversable identifier while ignoring other identifiers included in the nodes. For example, the traversable identifier can be the filtered domain - specific metadata identifier 1112a. As referred to herein, the traversable identifier is the identifier that the system looks for when a user provides a query in an attempt to search for data, while a non - traversable identifier is an identifier that the system stores associated with node 1102a and does not look for when traversing metadata graph 1100. Thus, the system reduces the amount of computational resources traditionally utilized when a large table is searched by a character string, since the metadata graph is (i) a graph that provides direction (e.g., a directed graph), and (ii) uses entity - specific, domain - specific, contextually accurate metadata identifiers to find the same instance of a data object stored across the entire entity system. For example, when the system traverses the metadata graph, the system can compare a set of terms (such as described in act 406 of process 400 (FIG. 4)) to the traversable metadata identifiers of metadata graph 1100. The system can traverse metadata graph 1100 based on the filtered domain - specific metadata identifier, but metadata graph 1100 can still store other metadata identifiers (which can be any of, for example, file - level metadata identifier 1106a, container - level metadata identifier 1108a, or system - level metadata identifier) associated with nodes 1102a - 1102d to maintain information for future use.For example, when the metadata graph searches for a given data object using the filtered domain-specific metadata identifier 1112a, the system can then retrieve file-level, container-level, or system-level location information, or other information about the given data object.

[0106] Referring again to FIG. 8, in some implementations, the system can perform a validation process on the generated metadata graph 1100 (FIG. 11). For example, in some implementations, the system calculates query-to-result performance metrics and accuracy performance metrics. For example, the validation process can include providing an automatically generated or user-provided test query to the domain-specific integrated metadata graph to measure the performance of the domain-specific integrated metadata graph.

[0107] For example, referring to FIG. 10D showing a subsystem diagram of feedback component 1012, feedback component 1012 can include a versioning component 1026, a generator 1028, a result 1030, and an update component 1034. The versioning component 1026 can store a previous version of the metadata graph 1100 (FIG. 11). For example, the versioning component can store the most recent version of the metadata graph 1100 before an updated version of the metadata graph 1100 (FIG. 11) is generated. The generator 1028 can generate a test query to provide the metadata graph 1100. For example, the test query can be the same as or similar to a user-specified query as discussed in act 402 of process 400 (FIG. 4). Additionally or alternatively, the test query can be a historical user-specified query as discussed in act 402 of process 400 (FIG. 4). The test query can be utilized to generate one or more results. For example, the result 1030 can determine (e.g., generate) one or more performance measurement criteria based on a test query such as that generated by the generator 1028. For example, the result 1030 can store historical performance measurement criteria of other versions of the metadata graph 1100 and performance measurement criteria of the current version of the metadata graph 1100 (FIG. 11). The performance measurement criteria can be query-to-result performance measurement criteria, accuracy measurement criteria, or other performance measurement criteria.

[0108] Once the performance measurement criteria for the metadata graph 1100 are generated, a decision may be made to update the metadata graph 1100 (FIG. 11). The update component 1034 can automatically trigger the update process to the metadata graph 1100 (FIG. 11). In some implementations, however, the update component 1034 can also trigger the update process to the metadata graph 1100 in conjunction with the third-party input 1032. For example, the third-party input 1032 may be a third-party source of information (e.g., a website, a computing device, etc.). As another example, the third-party input 1032 may be the input of an expert in a particular field. In this way, by leveraging the input of experts in a particular field, humans can verify the accuracy of the LLM-generated metadata graph before publishing such metadata graphs for use across systems. Adding the opinions of experts in a particular field to the generation process of the metadata graph 1100 can enhance the accuracy with which the metadata graph is generated in order to avoid any unintended LLM-related errors.

[0109] To perform a verification process, the system can provide a first query (e.g., a test query) that requests the location of a first data item for each of (i) a domain-specific integrated metadata graph and (ii) other versions of the domain-specific integrated metadata graph. As an example, the system can test the most recent iteration of the domain-specific integrated metadata graph to find the location of a given data item (e.g., stored in a data silo). Nevertheless, to ensure that the most recent modification to the domain-specific integrated metadata graph results in a better metadata graph, the system compares the performance metrics of the domain-specific metadata graph to a previous version (or other version) of the metadata graph. For example, the system can calculate query-to-result performance metrics that indicate the period between when a query is provided to each metadata graph and when a result is received from each metadata graph. The period can be measured in epoch time, seconds, milliseconds, ticks, microseconds of UNIX® time, etc. Such query-to-result performance metrics can be generated for each of the domain-specific integrated metadata graph and other versions of the domain-specific integrated metadata graph (e.g., a previous version of the integrated metadata graph).

[0110] In some implementations, the system can calculate accuracy metrics (e.g., performance metrics) for a domain-specific integrated metadata graph and other versions of the domain-specific integrated metadata graph. For example, the accuracy metric can be a degree of accuracy of the result (e.g., percentage, decimal value, ratio, integer, binary value, numerical value, alphanumeric value, etc.) generated from the domain-specific integrated metadata graph and other versions of the domain-specific metadata graph. The accuracy metric can be generated based on a human evaluation of the results (e.g., the results returned from providing a query to each metadata graph). For example, experts in a particular field (e.g., data scientists, software developers, computer engineers) can verify the accuracy of the results for each of the domain-specific metadata graph and other versions of the domain-specific metadata graph. Since each of the metadata graphs integrates the "domain", "context", and "terminology system" of the system for a given entity, experts in a particular field can verify the accuracy of the generated results returned by each metadata graph when finding the location of a given data item. In this way, the experts can verify the accuracy of the results, thereby enabling more accurate generation of the domain-specific integrated metadata graph. Still, in other implementations, the accuracy metric may be automatically generated without human intervention. For example, the accuracy metric can be based on a comparison of the results generated from each metadata graph with historical results according to one or more implementations of the present technology.

[0111] The system can calculate the accuracy metric by sampling and inspecting one or more parts of the result, or all of the result. For example, the system can select a sample set of the results to determine the accuracy metric and can vary the size of the sample set until a desired accuracy metric threshold is met.

[0112] In some implementations, the system can determine whether to perform an update process on the metadata graph. For example, the system can determine whether the performance measurement criteria of a metadata graph (e.g., metadata graph 1100) meet the performance metrics regarding the second performance measurement criteria of other versions of the metadata graph (e.g., the previous version). In some implementations, determining whether the performance measurement criteria of the metadata graph meet the performance metrics can be based on: (i) whether the query-to-result performance measurement criteria of the metadata graph can exceed the query-to-result performance measurement criteria of other versions of the metadata graph; (ii) whether the result accuracy measurement criteria of the metadata graph meet or exceed the result accuracy measurement criteria of other versions of the metadata graph. Thus, when the performance metrics fail to be met, the system can perform an update process on the metadata graph. Put simply, if the metadata graph (i) returns results faster than the previous version of the metadata graph and (ii) returns more accurate results than the previous version of the metadata graph, the metadata graph should not be updated. However, if the metadata graph (i) returns results slower than the previous version of the metadata graph or (ii) returns less accurate results than the previous version of the metadata graph, the metadata graph will be updated.

[0113] Referring again to FIG. 8, in act 810, process 800 performs an update process on the metadata graph. For example, in response to determining that the first performance metric of the domain-specific integrated metadata graph cannot meet the performance measure for the second performance metric of another version of the domain-specific integrated metadata graph, the system can perform an update process on the domain-specific integrated metadata graph (e.g., metadata graph 1100 (FIG. 11)). As described above, the metadata graph will be updated if (i) it returns results later than the previous version of the metadata graph, or (ii) it returns results that are less accurate than the previous version of the metadata graph.

[0114] In some implementations, the update process can be performed by updating the nodes and edges of the metadata graph to the nodes and edges of a previous version of the metadata graph. For example, the system can determine a set of mismatches between a domain-specific integrated metadata graph and another version of the domain-specific integrated metadata graph. The set of mismatches can reflect mismatches between (i) the nodes of the domain-specific integrated metadata graph and another version of the domain-specific integrated metadata graph, and (ii) the edges connected to at least one node of the domain-specific integrated metadata graph and another version of the domain-specific integrated metadata graph. For example, the system can traverse each graph to determine newly added nodes, edges, metadata identifiers, or other information. For example, the system can first traverse the current version of the domain-specific integrated metadata graph and store a tabular representation of the current version of the domain-specific integrated metadata graph in a database. The system can then retrieve a tabular version of another version (e.g., a previous version) of the domain-specific integrated metadata graph, if available. In some implementations, the system can traverse another version of the domain-specific integrated metadata graph and store a tabular representation of the other version of the domain-specific integrated metadata graph in the database. The system can then compare the two tabular versions of the metadata graph with each other to identify mismatches between the two versions. For example, the system can identify one newly added node (e.g., and the metadata identifier or location identifier included in the node), as well as two newly added edges that connect this node to other nodes in the current version of the metadata graph (when compared to the previous version of the metadata graph). The system can then update the domain-specific integrated metadata graph with the updated nodes and edges of the other version of the domain-specific integrated metadata graph corresponding to the set of mismatches.In this way, when the system fails to meet the performance measurement criteria, it can revert back to the previous version of the integrated metadata graph.

[0115] In some implementations, the system can initiate an update process based on detecting the addition of a data silo. For example, in some implementations, the system can perform an update process on the domain-specific integrated metadata graph when a data silo is added to the computing environment associated with an entity. The system can use one or more network discovery tools, such as SNMP, LLDP, CDP, or others, to monitor the computing environment (e.g., FIG. 3) associated with the entity to identify when a new device is added to the entity's computing environment. The system can then communicate with the new device using the IP address determined via the network discovery tool to verify the addition of the new device (e.g., using the "ping" command) and to determine the type of the new device. The system can further compare the determined IP address of the new device with a database storing device information for the computing environment. For example, the table can store the IP address of the device and the device type for the entity's computing system. The system can use the table to determine whether the newly added device is a data silo (e.g., a data source, a database, etc.). In response to detecting the addition of a data silo, the system can cause an update process to be performed on the domain-specific integrated metadata graph. For example, the system can extract metadata from the newly added data silo to update or regenerate the domain-specific integrated metadata graph 1100 (FIG. 11). For example, the system can repeat one or more of the above processes until all metadata for the data silo (and newly added data silos) has been processed by the system, thereby generating a system-wide domain-specific integrated metadata graph.

[0116] Artificial Intelligence Sandbox The above domain-specific integrated metadata graph enables a computer system to efficiently retrieve silo-stored data spanning different locations. An artificial intelligence (AI) sandbox enables the use of data from these different locations to generate AI models for data analysis or to automatically apply existing AI models. The sandbox provides a low-code or no-code environment and enables users to build and apply models to extract insights or predictions from data, even when the user lacks the expertise or time to build the AI models and the data pipelines for the data to be processed through the models, or when the user is unaware that the model should be useful for solving the problem. As described herein, the AI sandbox includes automated tools for deploying and managing models that ensure the efficient operation of the models without requiring human intervention. The AI sandbox can be used to automate any or all steps in the model life cycle, such as training new models, fine-tuning or improving existing models, deploying models, or generating data pipelines appropriate for analysis by the models when deployed. By integrating these functions, the sandbox provides a practical solution that transforms the user interaction with data and AI models, enhancing accessibility and usability while maintaining operational efficiency.

[0117] FIG. 12 is a block diagram illustrating components within an AI sandbox 1200 according to some implementations. As shown in FIG. 12, the AI sandbox 1200 can include a model reviewer assistant 1210, a data processor 1220, a model generator 1230, a model governor 1240, and a model automator 1250. Other implementations of the AI sandbox 1200 can include additional, fewer, or different components, or the functions can be divided differently among the components.

[0118] The Model Reviewer Assistant 1210 interacts with the user and orchestrates the other components of the AI Sandbox 1200 to generate and apply an AI model. The Model Reviewer Assistant 1210 can utilize a large language model (LLM) to both interact with the user and perform tasks related to generating, applying, or improving an AI model when the user interacts with the Assistant 1210. In some implementations, the Model Reviewer Assistant 1210 generates a chat-style interface where user input is received and information is output to the user.

[0119] The Data Processor 1220 identifies relevant data objects applicable to the model that the user requests to train or apply. The Data Processor 1220 processes the data to ensure cleanliness and consistency, making the data suitable for analysis by the AI model. The data can be prepared for various stages of AI development, including generating, training, and testing a dataset to train, fine-tune, or evaluate the model. Additionally, the Data Processor 1220 can provide an archival ability to securely store the clean data and the generated training samples. Using this archived data, the Data Processor 1220 can generate documents that facilitate the inspection of the model and its use, ensuring transparency and compliance with regulatory standards. The Data Processor 1220 is further described with respect to FIG. 13.

[0120] The model generator 1230 constructs an AI model based on user input and data output by the data processor 1220. The model generator 1230 is configured to train a machine learning model using the cleaned and processed data set from the data processor 1220. The model generator 1230 can analyze the data set output by the processor 1220, for example, by creating a set of features from the data set and identifying the target variable of the model. The model generator 1230 can also recommend the type of machine learning model to be constructed from a list of available models or based on model performance metrics. Using the processed data set and the selected model type, the model generator 1230 constructs a model, which can include training a new model, fine-tuning an existing model, or preparing a model for deployment. The model generator 1230 is further described with respect to FIG. 14.

[0121] The model governor 1240 performs higher-level analysis and governance evaluation tasks to ensure that the model complies with a set of governance metrics. These governance metrics can relate to policies or procedures within the organization in which the model reviewer assistant 1210 operates, or to policies from external organizations that the organization must follow, such as policies or regulations implemented by a government or standards-setting body, or a social covenant signed by the organization. If the model does not comply with the governance metrics, the model governor 1240 can cause the model to be modified until it complies. The model governor 1240 can further generate documentation for the model that can be archived and stored for subsequent analysis and inspection. The model governor 1240 is further described with respect to FIG. 15.

[0122] The Model Automator 1250 manages the deployment of AI models. The models can have different requirements for their deployment. For example, some models require a large and significant amount of computing resources. Some models are used for applications where results are needed quickly. Entities using AI models can access various different locations for deploying the models, such as one or more cloud providers or on-premises machines or scalable resources. The Model Automator 1250 can orchestrate the deployment of models across these various different locations. The Model Automator 1250 can provide a central system that can access any applicable APIs for deploying and invoking the models, access to the data used to train these models, as well as access to the application data processed by the models, and thus the Model Automator 1250 can determine how to efficiently deploy the models such that the models can effectively implement and manage computing resources. The Model Automator is further described with respect to FIG. 18.

[0123] FIG. 13 is a block diagram illustrating components within the data processor 1220 according to some implementations. The data processor 1220 can generally prepare data for various purposes associated with an AI model, including generating training data for training or fine-tuning a model, generating test data for testing a model, or preparing data for application to a model.

[0124] When a request for data (e.g., for training, fine-tuning, or testing a model or for applying to a model) is received, data processor 1220 accesses dataset 1310 for which the entity requesting access is authorized. Dataset 1310 can be accessed using any of the techniques described above, including using a metadata graph to identify data objects within data silos of a data repository.

[0125] Data processor 1220 can include a data preprocessor 1320 that performs various operations on the accessed dataset 1310 to prepare data for a desired use. The data preprocessor can perform operations to generate a training dataset or a test dataset for training, fine-tuning, and / or testing a model. Such operations can include, for example, normalizing data, converting data from one type to another, converting data to an appropriate format or structure, or aggregating or reducing data. In some implementations, data preprocessor 1320 can apply anonymization operations to remove, encrypt, or obfuscate personally identifiable or private information within the accessed dataset 1310. Some implementations of data preprocessor 1320 also apply operations to bring data into compliance with policies or regulations. For example, data preprocessor 1320 can remove data determined to be inaccurate from dataset 1310. An example process for detecting and removing noise from a dataset is described in U.S. Patent Application No. [18 / 736,407] filed on June 6, 2024, which is incorporated herein by reference in its entirety.

[0126] The processed data output by the data preprocessor 1320 can be passed to other elements of the data processor 1220 to generate a set of data for training or testing the AI model. The data preprocessor 1320 can, additionally or alternatively, generate a set of application data 1322 for applying to an existing AI model. The data preprocessor 1320 can perform cleaning operations, normalization operations, conversion operations, or other processing operations on the accessed data set 1310 to generate the application data. These operations can include anonymizing the data as described above, or, in another example, applying a preprocessing model to the data set 1310 that cancels out distortions in the model to which the application data set 1322 will be applied. The preprocessing model can include one or more data correction operators that modify the raw data set 1310, for example, by adding features, removing features, or changing values within the data set 1310, to create a set of corrected data. When applied to the AI model, the corrected data items from the set of corrected data cause the model to produce an output that does not exhibit distortion in the model. A process for detecting and canceling model distortion via a preprocessing model is described in U.S. Patent Application No. 18 / 783,409, filed July 25, 2024, which is incorporated herein by reference in its entirety.

[0127] The data analyzer 1330 analyzes the data to understand the data type and how the data is used, and generates a set of available data (e.g., training data, test data, or speculative data sets). The data analyzer 1330 receives the raw data and / or the preprocessed data set output by the preprocessor 1320, and evaluates the nature, such as the data type, the range of the data, the linkages or dependencies between the data, or other information that helps the system understand what the data is, how the data is used, and how the data can be used in the model. Based on the analysis of the data analyzer 1330 about the nature of the data, the data analyzer 1330 can further generate synthetic data to supplement the existing data. In some implementations, the data analyzer 1330 includes a data profile generator 1340 and / or a sampler 1350, which are described below.

[0128] The data profile generator 1340 generates a data profile 1342 of the data set output by the data preprocessor 1320. The data profile generator 1340 analyzes the structure of the data set and the interrelationships between the data within the set to ensure that the data set is appropriate for the intended artificial intelligence application. Such analysis can include, for example, calculating the statistics of the data set (e.g., mean or standard deviation), checking the data types within the data set and correcting the data types if necessary, checking for data anomalies (e.g., missing values, duplicates, or outliers), detecting patterns or correlations between data items within the data set, or detecting asymmetry or imbalance in the distribution of the data within the data set that can affect model performance. The data profile 1342 output by the data profile generator 1340 can further include metadata that describes the data set and that can be used to generate model documentation or to perform inspections and compliance checks.

[0129] In some implementation forms, the data processor 1220 generates a set of training data that can be used to train a machine learning model. Therefore, the data processor 1220 can further include a sampler 1350 that samples the data sets output by the data pre-processor 1320 and the data analyzer 1330 to generate training samples 1352 or test samples 1354. For example, the sampler 1350 can generate representative random samples of data items from the processed data sets for each of the training data set and the test data set. The training sample 1352 can be further input to a synthetic fabricator for generating additional training or test cases.

[0130] FIG. 14 is a block diagram illustrating components within the model generator 1230 according to some implementation forms. As described above, the model generator 1230 trains a machine learning model using the cleaned and / or processed data sets output by the data processor 1220. The model training procedure can be performed based in part on the input 1405 received from the user (e.g., via the model reviewer assistant 1210).

[0131] As shown in FIG. 14, the model generator 1230 can include a feature set generator 1410 that collects the data profile 1342 and the training samples 1352 generated by the data processor 1220. The feature set generator 1410 creates a set of features from the data sets that can be used to train a machine learning model. When creating features, the feature set generator 1410 can convert the data (e.g., within the training sample 1352) into an appropriate machine-readable format by, for example, converting non-numeric values to numeric values, converting numeric formats to the same type, vectorizing the data, or the like.

[0132] The features output by the feature set generator 1410 can be defined based in part on user input. For example, when a user is creating a model using the AI sandbox 1200, the user can directly specify the features of interest in the model or provide the information that the feature set generator 1410 uses to identify features within the dataset.

[0133] In some implementations, the feature set generator 1410 further learns over time how to use feedback to select relevant features for a given application. For example, when a user interacts with the AI sandbox 1200 to generate a model, the feature set generator 1410 can evaluate features such as the data types input to the model, the target variable of the model, or the performance of the model, which are used in each model, to build a robust mapping between the features and attributes of the model. The feature set generator 1410 can then use this mapping to recommend features for other models or train a feature selection model based on the mapping. In another example, the feature set generator 1410 uses model performance feedback to recommend or select features for a given model. The feature set generator 1410 can, for example, select a first set of features and receive feedback indicating the performance (e.g., accuracy) of a model trained using the first set of features. The generator 1410 can then select a second set of features, receive feedback indicating the performance of a model trained using the second set of features, and compare the performance measurements to determine which of the first and second sets of features resulted in better results.

[0134] The target selector 1420 selects the target variable of the machine learning model based on the features output by the feature set generator 1410 and / or based on user instructions received via the model reviewer assistant 1210. The target variable specifies the output that the machine learning model will predict or classify. To identify the target variable, the target selector 1420 can receive the identification of the variable from the model reviewer assistant 1210 generated based on the user input 1405. For example, the model reviewer assistant 1210 uses an LLM to identify the user-specified target variable in the natural language user input. The target selector 1420 compares the variable identified in the user's input with the features output by the feature set generator 1410 to determine whether the user-specified variable exists in the set of features or can be derived from the set of features. In some cases, the target selector 1420 can use an LLM to evaluate the set of features and identify the closest match to the user-specified target variable. The target selector 1420 can also, additionally or alternatively, use patterns of user behavior to identify the target variable. For example, if the user has recently used the AI sandbox 1200 to generate a model based on a specific target variable within the corresponding dataset, the target selector 1420 can determine that the user may be interested in generating a model based on the same target variable in a different dataset. Similarly, the target selector 1420 detects similarities between the actions of the user within the enterprise (such as other models built by the user, models the user has used, or data the user has created or accessed) and the actions of other users who have used the AI sandbox 1200 to build models. Based on these similarities, the target selector 1420 can determine that it is likely that the user will build a model for the specified target variable since other users have built models for the same specified target variable.

[0135] The target selector 1420 can, additionally or alternatively, use feedback from the user or other systems when selecting a target variable. This feedback can be used instead of the process for selecting the target variable from the user's natural language input or based on the user's past activities described above, or to improve the selector 1420's ability to identify the correct target variable based on the natural language input or user activities. In one example, after a model has been initially trained for one target variable, feedback from the user or an external system can be used to determine that the model should be trained for a different target variable (e.g., because the user provides direct input indicating that the trained model is not producing the desired output, or because another system identifies an error in the output of the trained model). In other cases, the target selector 1420 generates a mapping between the target variables used in other models and the attributes of the models (such as the features input to the other models, the data types used in the other models, the user who created the other models, or the performance of the other models) to predict target variables that are likely to be relevant to the developed model.

[0136] The model selector 1430 evaluates whether a model should be generated or deployed and, if so, recommends the type of machine learning model that the model generator 1230 should build. The model type can be one of a set of different model architectures such as neural networks, random forests, or support vector machines. The model type can alternatively be selected from a set of commercial or existing models that can be used as-is or fine-tuned for a particular purpose. Similarly, the model selector 1430 can select from pre-trained models that can be models previously developed within the organization where the model reviewer assistant 1210 operates or received from external sources. For each use case, the model selector 1430 can recommend an individual model or a set of multiple models to achieve the user's desired goal. For example, the model selector 1430 can recommend generating a series of models that can be used together in an ensemble method. When recommending multiple models, the model selector 1430 can recommend generating all new models using a selected set of pre-trained or fine-tuned models or combining new models with pre-trained or fine-tuned models. The model selector 1430 can also recommend specific ensemble learning techniques that enable these models to be used together. For example, the model selector 1430 can recommend generating a series of three models in a voting ensemble where the predictions for each model are combined by majority vote or averaging. In another example, the model selector 1430 recommends a stacking ensemble where the outputs of several base models are used as inputs to a meta-learning model that makes the final prediction or a boosting ensemble where the models are sequentially trained focusing one by one on correcting the errors of its predecessors to improve the overall performance.

[0137] When selecting a model, the model selector 1430 can receive an explicit model selection from the user. For example, the model reviewer assistant 1210 can provide the user with a list of available types of models, and the user can select a model from the list. Alternatively, the model selector 1430 can recommend a model type for a given application. To recommend a model, the model selector 1430 can apply a set of input data to each of a number of types of models and calculate a measure for each type of model. The measures can include, for example, a measure of how quickly the model generates an output (e.g., latency, model output speed, or variance of model response time) for a set of input data, a measure of how accurate the model's output is, the amount of memory used by the model, or a measure of the cost of using the model. Based on the measures, the model selector 1430 can recommend one or more model types that achieve a particular goal, such as the fastest execution or the most accurate results. In other cases, rather than applying the input data to a number of models and calculating the measures used for model selection, the model selector 1430 can recommend a model type based on historical model performance. For example, the model selector 1430 can evaluate historical data that indicates, for example, that one type of model is typically faster but less accurate than another type of model, to enable the model selector 1430 to make a model type recommendation based on whether speed or accuracy is more important for a given application. Some implementations of the model selector 1430 output an identifier of the recommended model to the model builder 1440 to enable the recommended model to be built. In other implementations, the model selector 1430 outputs the recommended model type to the user, such that the user can then select between the recommended model types or select another type of model.

[0138] The model builder 1440 constructs a model 1445, which can include training a new model, fine-tuning an existing model, or packaging or refining a deployment model without further training. The model builder 1440 receives a model type selected from the model selector 1430 and can determine whether to perform training, fine-tuning, packaging, or other tasks based on the selected model. When training a new model or fine-tuning an existing model, the model builder 1440 trains the model type selected by the user or the model selector 1430 based on the target variable identified by the target selector 1420 and the set of training samples 1352 generated by the data processor 1220. The model builder 1440 then tests the model using the test samples 1354 generated by the data processor 1220 and can retrain as necessary, for example, until the model converges or reaches a specified accuracy threshold for the test data set.

[0139] FIG. 15 is a block diagram illustrating the functions of the model governor 1240 according to some implementations. As shown in FIG. 15, the model governor 1240 can utilize the model reviewer assistant 1210 to evaluate the model 1445, including interfacing with the large language model (LLM) 1510 using the model reviewer assistant 1210 and evaluating the model 1445 for higher-level analysis and governance tasks. In some implementations, the LLM 1510 uses the RAG-based process 1520 to retrieve policies, procedures, know-how, or other applicable governance information from a data repository such as the industry knowledge repository 1522 or the enterprise knowledge repository 1524. The industry knowledge repository 1522 can store, for example, scientific models, national regulations, or data standards. The enterprise knowledge repository 1524 can store information such as organizational policies, organization-specific classifications, or the organization's business mission. The model reviewer assistant 1210 can generate queries to the LLM 1510 that cause the LLM 1510 to evaluate the model 1445 against the applicable governance information.

[0140] Based on the evaluation, the Model Governor 1240 can update the model 1445, create additional models or data to make the model 1445 compliant with the governance metrics, or generate a document describing whether and to what extent the model complies with the governance information. For example, an organization may be subject to regulations that require a particular type of model (e.g., a model for approving loan applications) to produce output that does not skew towards a particular background (e.g., race, gender, age, sexual orientation, or geographical location or loan applicant). If the Model Governor 1240 detects that such a loan approval model is unfairly skewing its conclusions based on one or more of these backgrounds, the Model Governor 1240 can cause the model to be retrained or fine-tuned to reduce the skewed conclusions. Alternatively, the Model Governor 1240 can cause a second model to be generated, which is configured to preprocess the application data before it is input into the model 1445 to correct for the model's skew. In another example, if the Model Governor 1240 determines that the model complies with each of a set of governance metrics, the Model Governor 1240 can generate (optionally using the LLM 1510) a document describing the governance metrics against which the model 1445 was evaluated and how the model was determined to comply with each of the metrics.

[0141] The model governor 1240 can further generate an explainability layer 1530 for model 1445. The explainability layer 1530 includes data associated with model 1445 that explains what the model is doing, how the model is making decisions, which data was used to train the model, which data is an input or output from the model, any modifications applied to the input data before it is processed through the model, or any other characteristics specified by the enterprise, regulatory, or other standards for the organization for which the model was generated. The model governor 1240 can generate the explainability layer using the LLM 1510 by analyzing the model 1445 itself and / or governance documents retrieved from repositories such as the industry knowledge repository 1522 or the enterprise knowledge repository 1524.

[0142] Automated Generation of AI Models FIG. 16A is a flowchart illustrating a process 1600 for automatically generating an artificial intelligence (AI) model according to some implementations. The process 1600 can be executed by a computer system, such as one or more systems implementing the components of the artificial intelligence sandbox 1200.

[0143] As shown in FIG. 16A, at 1602, the computer system receives natural language input from a user. The natural language input can include explicit instructions for the computer system to train an AI model or general queries that the computer system can process as instructions to generate a model. The input can include a set of phrases that implicitly or explicitly indicate desired properties of the model, such as data to be processed by the model and / or the intended output from the model. In one example, the user can provide the input "I need help analyzing my investment portfolio." The computer system processes this input to determine whether the model would be useful in answering the query or whether the query can be answered without the model. For example, the computer system determines that the model should be useful for analyzing the investment portfolio, while a more direct query (e.g., "Do I own any shares of XYZ Corp?") should not be required or should be complicated by the model. When entering the natural language input, the user may not be aware that the model should be useful, let alone the type of model that should produce the best results, the data that should be used to train or input to the model, or how to approach training and deploying the model.

[0144] The input can be received via, for example, a model reviewer assistant 1210 that can provide a chat-style interface where the input can be received from a user and information can be output to the user. User input to the chat interface can be received as natural language input and / or as other types of input, such as selection of an item from a list. Similarly, the model reviewer assistant 1210 can output information to the user in natural language format (e.g., using an LLM to generate the output), in graphical format such as a data plot, or in other formats. An example of a chat interface 1700 where user input can be received is depicted in FIG. 17A. In the example, the user provided the natural language input 1705 "I need help analyzing my investment portfolio." Other users can provide natural language input that more directly instructs the computer system to generate a model (e.g., "Create a model for analyzing my investment portfolio"). The phrase "my investment portfolio" can be processed by the computer system as indicating the data to be processed by the model.

[0145] In 1604, the computer system accesses a metadata graph based on user input. The metadata graph can comprise (i) a set of nodes comprising (a) metadata indicating internal data objects stored in data silos, and (b) location identifiers of the data silos, and (ii) edges indicating data lineages between the sets of nodes. As described above, the computer system can traverse a metadata graph that indicates where data is stored, which data in different data silos is available, and the data lineages between the nodes of the graph, thereby enabling the computer system to efficiently find data within a set of data silos. By traversing the metadata graph, the system can determine nodes corresponding to a set of phrases in the natural language input received from the user. In the example interface illustrated in FIG. 17A, the computer system selects four candidate data sources associated with one or more nodes of the metadata graph identified based on the phrase "my investment portfolio", and each data source includes one or more data items. For example, a data source for the implementation of Investment A can include a set of measurements of the value of Investment A at different times. In some implementations, such as the example shown in FIG. 17A, the user can provide additional input for selection from among the candidate data sources shown in FIG. 17B. Alternatively, the computer system can proceed to train a new model with data objects identified based on the user's natural language input.

[0146] In 1606, the computer system processes data objects retrieved using a metadata graph to generate a training dataset. The computer system can preprocess the data objects, such as by removing skewed, incorrect, or irrelevant values from the dataset. Once the data objects are cleaned, the system can sample a training dataset from the data objects and / or input the data objects into a synthetic data generator to generate synthetic training data. Similarly, the computer system can generate a set of test data for testing the model once the model is trained.

[0147] In 1608, the computer system selects a model type for the model to be trained. The model type can be selected from different model architectures and / or from commercial packages or existing models using model metrics associated with each available model type. In some cases, the model type can be output to the user via a chat interface that allows the user to select the type of model to train. The computer system then, in 1610, trains the selected type of model using the generated set of training data.

[0148] After the model is trained, regardless of whether by the process of FIG. 16A or another process, the computer system can also interact with the user to automatically generate a pipeline for application data to be processed through the model. FIG. 16B is a flowchart illustrating a process 1620 for automatically generating a data pipeline for an AI model according to some implementations. Process 1620 can be at least partially implemented by the same computing system that implements process 1600 of FIG. 16A or can be implemented by one or more different computing systems.

[0149] As shown in FIG. 16B, at 1612, the computer system receives, from a user, a first natural language input that may include a set of clauses and instructions for analyzing data associated with a set of clauses using an AI model. Like the input for training the model, the input for deploying the model can be received via a chat interface generated by the model reviewer assistant 1210. For example, FIG. 17C illustrates that a user input 1710 instructing the computer system to "look at last year's investment mix" is received via the chat interface. The clause "last year's investment mix" can be used to identify a set of data to which the model will be applied, while "look at" is interpreted in the context of the chat session as an instruction for deploying the model generated during the chat session.

[0150] At 1624, the computer system uses the natural language input to access a metadata graph that directs internal data objects stored in a data silo. The metadata graph can be the same graph used to identify the data objects for generating the training data set described with respect to FIG. 16A. Using the metadata graph, the system can determine the nodes corresponding to the set of clauses in the first natural language input.

[0151] In 1626, the computer system processes the internal data object indicated by the determined node to generate a first set of data. For example, the computer system can remove personally identifiable or private information from the internal data object. In another example, the system applies a set of data modification operators to the internal data object to generate a set of modified data based on the internal data object, and the data modification operators are configured to remove skews from the internal data object or remove or compensate for inaccurate or irrelevant data within the object.

[0152] In 1628, the computer system applies an AI model to a first set of application data to generate one or more outputs. For example, the computer system uses the model to classify data items in the first set of application data or make predictions based on one or more of the application data items.

[0153] In 1630, the computer system sends to the user a representation of one or more outputs for display. For example, FIG. 17D illustrates that the computer system generates plot 1715 to illustrate the results produced when the model is applied to a set of data (e.g., "last year's investment mix") specified by the user.

[0154] Based on the presented output, the user can determine that modifications to the model or the data processed by the model are required to obtain the desired result. For example, the presented output can indicate that the data input to the model was incomplete or incorrect. Thus, the user can iteratively interact with the model reviewer assistant 1210 to modify the input to the model until the expected output is received. These iterative interactions can include further natural language input from the user via the chat interface provided by the model reviewer assistant 1210.

[0155] For example, at 1612, the computer system receives a second natural language input including instructions for modifying a first set of application data (e.g., by adding data to the first set, removing data from the first set, or modifying values within the first set). FIG. 17D further illustrates an example input 1720 where the user instructs the computer system to "please add data source 4". The computer system then, at 1614, generates a second set of application data based on the second natural language input. Generating the second set of application data in response to the instructions in the second input can involve retrieving additional data objects using the metadata graph, processing the data objects to obtain application data, modifying data values within the first set of application data (e.g., by modifying the way the data objects were processed to produce the first set of application data), or removing data from the first set of application data. The AI model can then, at 1616, be applied to the second set of application data, and the output generated from the application of the model can be displayed to the user. This iterative process can be repeated until the user is satisfied with the output of the model.

[0156] Deploying the AI model Once the AI model is trained and determined to comply with the model's governance parameters, the model can be deployed to make predictions based on the application data input to the model. The model automator 1250 determines how to deploy the model for use in a production environment and orchestrates the deployment.

[0157] FIG. 18 is a schematic diagram illustrating the operation of the model automator 1250 according to some implementations. As shown in FIG. 18, the model automator 1250 can include a model deployment engine 1810, a model deployment engine updater 1812, an orchestrator 1814, and a data handler 1816.

[0158] The model automator 1250 uses the model deployment engine 1810 to select a model deployment location for the AI model, and the model deployment location can be selected from among several available computing environments. For example, as illustrated in FIG. 18, the model automator 1250 can access a number of public cloud environments (e.g., from a first public cloud provider 1820A and a second public cloud provider 1820B), as well as one or more on-premises environments (such as an on-premises machine 1830A or an on-premises scalable resource 1830B).

[0159] The various deployment environments for AI models can present various advantages and disadvantages, particularly regarding cost efficiency, computing capacity, and privacy. On-premises environments are often limited in their capabilities as they are restricted by the hardware available within an enterprise. Still, they can offer the benefit of privacy, which is not available in some cloud environments. On the other hand, cloud environments provide scalable resources, but the cost can vary significantly depending on the cloud provider and the amount of computing power used. For example, cloud providers typically charge based on usage and sometimes use a tiered pricing system where the price of the marginal amount of computing resources varies according to the total amount of resources used within a given time frame. Cloud and on-premises environments can also differ in terms of ease of integration with existing systems, flexibility of resource allocation, and the likelihood of downtime or service interruptions.

[0160] The model automator 1250 can monitor the operating parameters associated with the deployment environments available to the automator. The operating parameters can include relatively dynamic data such as the amount of computing capacity available for use by a given model, the computing resources used during the execution of a model deployed in the environment, the response time from a model deployed in the environment, or the accuracy of the results produced by the model when deployed in the environment. Other examples of operating parameters can include data that is static or changes over a longer period, such as the privacy policy of the operator of the deployment environment, the average amount of downtime of the environment, or the number of service interruptions to the environment. The model automator 1250 can measure some of the operating parameters when the automator deploys the model to various available environments. Alternatively, the model automator 1250 can obtain operating parameters from other sources, such as the operator of the environment or other systems that perform computing tasks within the environment.

[0161] Additionally, the Model Automator 1250 can facilitate the publication of models for use by other users. Within some organizations, it may be desirable for models generated by one user to be made available to other users. For example, many users within an organization may perform similar tasks and should benefit from models created by other users to assist with these tasks. When a model is published, the Model Automator 1250 can manage access rights or entitlements to the model. For example, the Model Automator 1250 can link the model to access rights that specify that only users within a particular department of the organization can use the model, or that the model can only be used by users with permission to access certain data (e.g., the data on which the model was trained).

[0162] The model deployment engine 1810 includes rules, models, or other logic tools that enable the model automator 1250 to select a model deployment location. For example, the model deployment engine 1810 can include trained decision models such as one or more decision trees or random forests, knowledge graphs, or rule engines. When determining the deployment location for a given model, the model deployment engine 1810 can input information such as parameters of the model itself (e.g., the size of the model, privacy considerations associated with the model, or information indicating whether the model will be used in real-time or in a batch processing flow), parameters of the data that will be processed through the model when deployed (e.g., the amount of data processed per iteration of the model, the location of the data, or privacy considerations associated with the data), or operational parameters associated with the available deployment locations. Based on one or more of these inputs, the model deployment engine 1810 selects a model deployment location for the model. In some implementations, the model deployment engine 1810 includes an explainability layer that enables the engine to output an explanation about its selected model deployment location. For example, if the user is using the chat-based interface from the model reviewer assistant 1210 to create and deploy a model, the model deployment engine 1810 can generate an explanation about its determination that can be output to the user via the model reviewer assistant 1210.

[0163] In one example, the model deployment engine 1810 evaluates the cost efficiency of available environments and deploys the model to an environment where the cost efficiency is greater than a specified threshold. The cost efficiency can be measured based on the amount of computing resources the model is expected to use and the expected cost for using these resources for each available environment. The model deployment engine 1810 can determine the cost efficiency as a single metric associated with each available model deployment environment or as a differential metric comparing the cost of deploying the model to one environment to the cost of deploying the model to another environment. In other cases, the model deployment engine 1810 selects the deployment location based at least in part on the privacy policy associated with the model, the input data that will be processed by the model, or the output produced when the model processes the input. For example, the model is deployed to an on-premises environment if the model has a privacy policy that limits the use of the model to an on-premises system, but is deployed to a cloud environment if there is no such limitation. In still other cases, the model deployment engine 1810 can select the model deployment location based in part on whether the model will be used to process data in real time or in a batch process. For example, if the model is used in a real-time processing flow, the engine 1810 can select a cloud environment to deploy the model based on the determination that the cloud environment can more easily scale resources to guarantee the availability of the model than an on-premises environment. On the other hand, if the model is used in a batch processing flow, the engine 1810 can cause the model to be deployed to an on-premises environment based on the determination that the execution of the batch process can be delayed until computing resources are available if necessary.

[0164] The model deployment engine update data 1812 can update the model deployment engine 1810 based on the continued observation of the operating parameters. For example, when the model deployment engine 1810 includes a trained decision model, the model deployment engine update data 1812 can retrain the decision model based on the updated operating parameters. When the model deployment engine 1810 includes explicit rules, the model deployment engine update data 1812 can update the rules by instructing a large language model (LLM) to modify the existing rules or create new rules based on the observed operating parameters. For example, when a cloud service provider updates its privacy policy, the model deployment engine update data 1812 can prompt the LLM to evaluate the updated privacy policy to determine whether the privacy policy of the organization operating the model automator 1250 complies with the updated privacy policy. Based on the evaluation, the LLM can then update the rules for acceptable deployment locations for a particular type of AI model, for example, if the first cloud service provider has modified its privacy policy such that the privacy policy no longer complies with the privacy policy associated with a particular AI model or the data processed by a particular AI model.

[0165] The Model Automator 1250 can use the model deployment engine 1810 to both select a model deployment location for a new model or a new instance of a model, and move the model from one deployment location to another. For example, since the engine 1810 is updated in response to observed operating parameters, the Model Automator 1250 can periodically use the engine 1810 to re-evaluate whether the model is deployed to a location that meets the engine's metrics (e.g., whether the cost efficiency of the environment is still greater than a corresponding threshold, or whether the environment still complies with the privacy policy associated with the data). In another example, the Model Automator 1250 can determine that a new instance of the model should be deployed to a second environment when the computing resources of the first environment exceed a given threshold (e.g., a specified pricing tier from a cloud computing provider operating the first environment), while the original instance of the model is in use.

[0166] The orchestrator 1814 facilitates the deployment of the AI model to the location selected by the Model Automator 1250 based on the model deployment engine 1810, making the model available for use in a production environment. When deploying the model, the orchestrator 1814 can configure the model for deployment to the selected infrastructure (including transferring the model to the selected infrastructure), and configure the infrastructure for the model (e.g., by spinning up the resources required for the model). The orchestrator 1814 is configured to automatically and seamlessly deploy the model in any of the available environments using platform-specific APIs. The orchestrator 1814 can also leverage containerization technologies such as Docker and orchestration frameworks such as Kubernetes to package the model with all the necessary dependencies and scale and manage the containerized application.

[0167] The orchestrator 1814 can determine deployment attributes required for the deployment of each model or to improve the performance of the deployed model. These deployment attributes can include, for example, the language of the model (such as Python or R), the infrastructure mechanisms of the model (such as available memory, processing speed, available parallelization, or hardware type), or configuration parameters for Docker file or Kubernetes deployment. In some implementations, the orchestrator 1814 maintains a set of patterns or templates that each specify the corresponding type of deployment attributes for the model. Some of these patterns or templates can be initially provided by the user. Other patterns or templates can be automatically generated by the orchestrator 1814, for example, by identifying similar types of deployment attributes of the model. The orchestrator 1814 can automatically update the patterns or templates over time when the orchestrator 1814 deploys the model using the attributes in the patterns or templates.

[0168] After the model is deployed, the orchestrator 1814 can monitor the deployed model to verify that the deployment was successful. Success can be measured, for example, by an indicator specifying whether the model is producing results, by measured values of model operation parameters (e.g., latency, memory utilization, or CPU or GPU utilization), by measured values of model performance (e.g., accuracy or precision), or by a combination of factors. The orchestrator 1814 can determine that the model was not successfully deployed if, for example, the model is not producing results, the model operation parameters are outside the specified range, or the model performance is different from the expected model performance by at least a threshold amount. When the orchestrator 1814 determines that the model was not successfully deployed, the orchestrator 1814 can decide to roll back the configuration to the previous configuration, deploy the model on a different infrastructure, completely stop the model deployment until the error is corrected, or take other corrective actions to improve the model deployment. The orchestrator 1814 can also update the deployment template based on a successful or unsuccessful deployment. For example, if it is determined that the configuration parameters in the Docker file were set outside the expected range for a particular type of model's model operation parameters, the orchestrator 1814 updates the deployment template for the type of model to ensure that the correct configuration parameters are used in future Docker files for the same type of model.

[0169] The data handler 1816 enables the deployed AI model to access data at the location where the AI model is deployed. In some implementations, the data handler 1816 identifies a pipeline of application data for the deployed model that is processed by the data processor 1220 described above. Using an API associated with the environment in which the model is deployed, the data handler 1816 can generate a script that makes the application data available to the environment for processing by the model. The data handler 1816 can further handle data privacy policies, check data for compliance with ethical or fairness standards, and / or obtain or generate approvals for data that can be used in a given situation.

[0170] FIG. 19 is a flowchart illustrating a process 1900 for automating the deployment of an AI model according to some implementations. The process 1900 can be implemented by a computing system, such as a system that implements aspects of the model automator 1250.

[0171] As shown in FIG. 19, at 1902, the computer system receives a first request to deploy a first AI model to make the first AI model available for use in a production environment. In some implementations, the request can be received as a natural language input to a chat interface similar to the inputs described with respect to FIGS. 17A-17D.

[0172] In 1904, a computer system selects a first model deployment location for a first AI model based on a model deployment engine. The deployment location can be selected from among a cloud provider environment or an environment (an "on-premises environment") that is at least partially operated by an entity that controls the first AI model. For example, an entity can contract with multiple cloud providers to use computing environments maintained by the cloud providers, but can also maintain part of its own computing infrastructure. The computer system can select whether to deploy the first model to an on-premises infrastructure or to a cloud environment and / or can select a specific location within the on-premises infrastructure or a specific cloud environment where the model will be deployed. To select a model deployment location, the model deployment engine can take into account the nature of the first AI model itself, the nature of the deployment location, or considerations regarding the governance of the model or entity-specific policies of the entity that controls the first AI model. The model deployment engine can include one or more trained models such as a decision tree or a random forest, a rule-based system such as one or more knowledge graphs or rule engines, or a combination of logic or tools that enable the computer system to make a determination regarding the model deployment location.

[0173] In 1906, a computer system generates a script for deploying a first AI model to a first model deployment location. For example, the computer system can generate a script that calls a platform-specific API for a selected model deployment location, utilize cloud technologies (such as containerization or orchestration) for a cloud deployment location, or file system or server management technologies for an on-premises deployment location, and ensure that the deployed model can access a data pipeline with data configured to be processed by the model.

[0174] After deploying the first AI model, in 1908, the computer system can monitor the operational parameters associated with the model deployment. The operational parameters can include, for example, the computational cost used by the first AI model, or the response time from the first AI model when deployed at a selected location, the privacy policy of the environment including the model deployment location, or measurements of the downtime or service interruptions of the environment including the deployment location.

[0175] Based on the operational parameters, in 1910, the computer system can update the model deployment engine, for example, by retraining a trained model within the model deployment engine or updating the rules within the engine. For example, the computer system can update the model when the cloud provider modifies its privacy policy, or when the actual operational parameters monitored by the computer system differ from the operational parameters used to train the model deployment engine.

[0176] The updated model deployment engine can then be used in 1912 to deploy a second AI model or redeploy the first AI model to a different location. In some cases, the second AI model can be deployed to the same environment as the first AI model or a different environment based on the operational parameters observed from the deployment of the first model. For example, if the first model is deployed to an on-premises system and the on-premises system is approaching its computing capacity, the second model can be deployed to a cloud environment. In other cases, the second AI model is a second instance of the first model deployed to a different location. For example, if an entity deploys a first AI model to a first cloud environment where the entity is approaching a specific compute cost threshold, a second instance of the first model can be automatically deployed to a second cloud environment to reduce the compute cost within the first environment. In still other cases, the computer system can determine that the first AI model should be moved from one deployment location to another. For example, if the privacy policy of the first cloud provider changes after the first model is deployed to the environment of the first provider, the computer system can move the first model from the environment of the first provider to the environment of another cloud provider.

[0177] Process 1900 can be repeated when additional models are deployed by the computer system. The computer system can thus iteratively improve its knowledge about how well the models perform in different environments, the cost to deploy the models to these environments, and how well the environments comply with governance or policy considerations. Leveraging this iteratively improved knowledge, the computer system can improve its ability to automatically and efficiently deploy AI models.

[0178] Section A system for reducing the amount of computing resources used when accessing silo storage data spanning heterogeneous locations via an integrated metadata graph, comprising at least one hardware processor and, when executed by the at least one hardware processor, receiving a user-specified query in a graphical user interface (GUI) that instructs a request to access a set of data objects, wherein each data object in the set of data objects is stored in a respective data silo of a set of data silos in heterogeneous locations, receiving, parsing the user-specified query for a set of keywords, wherein each keyword in the set of keywords is associated with the set of data objects, performing natural language processing on the set of keywords to determine a set of semantically similar phrases corresponding to each keyword in the set of keywords, accessing a metadata graph to determine nodes corresponding to the set of semantically similar phrases, wherein the metadata graph comprises (i) a set of nodes that indicate (a) metadata of internal data objects stored in the data silos and (b) location identifiers of the data silos, and (ii) edges that indicate data lineage between a first node and a second node in the set of nodes, and the metadata graph is generated using a metadata data structure based on file-level and container-level metadata identifiers, accessing, using the location identifier corresponding to the determined nodes, to determine a data silo storing at least one data object of the set of data objects in order to obtain at least one data object of the set of data objects via the data silo, and generating a visual representation of the at least one data object for display on the GUI, wherein the visual representation of the at least one data object comprises lineage information of the at least one data object, and at least one non-transitory memory storing instructions that cause the system to perform the above operations.

[0179] Section 2. The metadata graph extracts from a second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, where each file-level metadata identifier in the set of file-level metadata identifiers indicates the metadata of a given data object stored within a respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates the metadata of a respective data silo in the second set of data silos; generating respective sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier; generating a metadata data structure for mapping each semantically similar metadata identifier in the sets of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; and generating a metadata graph using the generated metadata data structure, the system of Section 1 generated thereby.

[0180] Section 3. Receiving a second user-specified query that instructs a request to generate an intended result via a second GUI; providing the second user-specified query to an artificial intelligence model to generate a recommendation, where the recommendation comprises (i) a second artificial intelligence model that is to be used to generate the intended result and (ii) a second set of data objects that is to be used when training the second artificial intelligence model; in response to receiving a user selection that indicates acceptance of the recommendation, further comprising instructions to (i) access a database to obtain the second artificial intelligence model and (ii) obtain the second set of data objects using the metadata graph; training the second artificial intelligence model using the set of data objects; and applying the second artificial intelligence model to generate the intended result, the system of Section 1.

[0181] To obtain a set of policies that instruct usage metrics corresponding to a second set of data objects, access the governance database and use the set of policies that instruct usage metrics corresponding to the second set of data objects to determine whether the second set of data objects is approved for use in training a second artificial intelligence model, and use a second set of policies that instruct usage metrics corresponding to artificial intelligence model predictions to determine whether the output of the second artificial intelligence model is approved for providing to one or more computing systems, and in response to (i) the second set of data objects being approved for use in training a second artificial intelligence model and (ii) the output of the second artificial intelligence model being approved for providing to one or more computing systems, further comprising instructions for applying the second artificial intelligence model to generate the intended result, the system of section 3.

[0182] A method for reducing the amount of computing resources used when accessing silo storage data spanning heterogeneous locations via an integrated metadata graph, the method comprising: receiving, in a graphical user interface (GUI), a user-specified query that instructs a request to access a set of data objects, wherein each data object in the set of data objects is stored in a respective data silo of a set of data silos in heterogeneous locations; performing natural language processing on the user-specified query to determine a set of clauses corresponding to the user-specified query; accessing a metadata graph to determine nodes corresponding to the set of clauses, the metadata graph comprising: (i) a set of nodes comprising (a) metadata that indicates internal data objects stored in the data silos, and (b) location identifiers of the data silos, and (ii) edges that indicate data lineages of the set of nodes, the metadata graph being generated using a metadata data structure based on file-level and container-level metadata identifiers; using the location identifiers corresponding to the determined nodes to determine data silos storing at least one data object of the set of data objects in order to retrieve at least one data object of the set of data objects via the data silos; and generating a visual representation of the at least one data object for display in the GUI.

[0183] Section 6. The metadata graph extracts from a second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, where each file-level metadata identifier in the set of file-level metadata identifiers indicates the metadata of a given data object stored within a respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates the metadata of a respective data silo in the second set of data silos; generating respective sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier; generating a metadata data structure for mapping each semantically similar metadata identifier in the sets of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; and generating a metadata graph using the generated metadata data structure, the method of Section 5 generated thereby.

[0184] Section 7. Receiving, via a second GUI, a second user-specified query that instructs a request to generate an intended result; providing the second user-specified query to an artificial intelligence model to generate a recommendation, where the recommendation comprises (i) a second artificial intelligence model that is to be used to generate the intended result and (ii) a second set of data objects that is to be used when training the second artificial intelligence model; in response to receiving a user selection that indicates acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model and (ii) using the metadata graph to obtain the second set of data objects; training the second artificial intelligence model using the set of data objects; and applying the second artificial intelligence model to generate the intended result, the method of Section 5 further comprising.

[0185] To obtain a set of policies that indicate usage metrics corresponding to a second set of data objects, access a governance database, and use the set of policies that indicate usage metrics corresponding to the second set of data objects to determine whether the second set of data objects is approved to be used to train a second artificial intelligence model, and use a second set of policies that indicate usage metrics corresponding to artificial intelligence model predictions to determine whether the output of the second artificial intelligence model is approved to be provided to one or more computing systems, and in response to (i) the second set of data objects being approved to be used to train a second artificial intelligence model and (ii) the output of the second artificial intelligence model being approved to be provided to one or more computing systems, apply the second artificial intelligence model to generate the intended result, the method of section 7 further comprising.

[0186] Determining a set of clauses corresponding to a user-specified query is parsing the user-specified query for a set of keywords, wherein each keyword in the set of keywords is associated with a set of data objects, parsing, and for each keyword in the set of keywords associated with the set of data objects, determining a set of semantically similar clauses corresponding to each keyword in the set of keywords, and using the set of semantically similar clauses corresponding to each keyword in the set of keywords to determine a set of clauses corresponding to the user-specified query, the method of section 5 further comprising. [[ID= 5]]

[0187] Determining a semantically similar clause corresponding to each keyword in a set of keywords further comprises accessing a database that indicates a mapping between a first set of keywords and a second set of keywords, and in response to accessing the database, using each keyword to determine a set of semantically similar clauses corresponding to each keyword, the method of section 9 further comprising.

[0188] Section 11. The method of Section 5 further includes traversing each node in a set of nodes to identify a metadata identifier that matches at least one clause in a set of clauses, and in response to determining that the metadata identifier matches at least one clause in the set of clauses, determining a node corresponding to the set of clauses.

[0189] Section 12. The method of Section 5 further includes traversing each node in a set of nodes to identify a metadata identifier that matches at least one clause in a set of clauses, and in response to determining that the metadata identifier matches at least one clause in the set of clauses, determining a first node corresponding to the set of clauses, and in response to determining that the first node corresponds to the set of clauses, performing a second traversal of the nodes in the set of nodes using an edge that indicates a first data lineage of the first node, wherein the first data lineage of the first node indicates a second node that comprises information that is a source of information associated with the first node, and determining a second data silo that stores a second data object in a set of data objects using a location identifier corresponding to the second node to obtain the second data object in the set of data objects via the second data silo, and generating a second visual representation of the second data object for display on a GUI.

[0190] Section 13. The method of Section 5, wherein a visual representation of at least one data object comprises lineage information of the at least one data object.

[0191] When executed by one or more processors, receive, in a graphical user interface (GUI), a user-specified query that instructs a request to access a set of data objects, wherein each data object in the set of data objects is stored in a respective data silo in a set of data silos in heterogeneous locations; perform natural language processing on the user-specified query to determine a set of clauses corresponding to the user-specified query; access a metadata graph to determine nodes corresponding to the set of clauses, the metadata graph comprising: (i) a set of nodes comprising (a) metadata that indicates internal data objects stored in data silos, and (b) location identifiers of the data silos, and (ii) edges that indicate data lineages of the set of nodes, the metadata graph being generated using a metadata data structure based on file-level and container-level metadata identifiers; use the location identifiers corresponding to the determined nodes to determine data silos storing at least one data object in the set of data objects in order to obtain at least one data object in the set of data objects via the data silos; and store instructions that cause an operation to be performed that includes generating a visual representation of the at least one data object for display in the GUI, in one or more non-transitory computer-readable media.

[0192] Section 15. The metadata graph extracts from a second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, where each file-level metadata identifier in the set of file-level metadata identifiers indicates the metadata of a given data object stored within a respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates the metadata of a respective data silo in the second set of data silos; generating a respective set of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier; generating a metadata data structure for mapping each semantically similar metadata identifier in the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; and generating a metadata graph using the generated metadata data structure. The medium of Section 14 is generated by the above.

[0193] Section 16. When instructions are executed by one or more processors, receiving, via a second GUI, a second user-specified query that instructs a request for generating an intended result; providing the second user-specified query to an artificial intelligence model to generate a recommendation, where the recommendation comprises (i) a second artificial intelligence model that is to be used to generate the intended result and (ii) a second set of data objects that are to be used when training the second artificial intelligence model; in response to receiving a user selection that indicates acceptance of the recommendation, further causing operations to be performed, including (i) accessing a database to obtain the second artificial intelligence model and (ii) using the metadata graph to obtain the second set of data objects; training the second artificial intelligence model using the set of data objects; and applying the second artificial intelligence model to generate the intended result. The medium of Section 14 is as described above.

[0194] Section 17. When the command is executed by one or more processors, access the governance database to obtain a set of policies that instruct usage metrics corresponding to a second set of data objects, and use the set of policies that instruct usage metrics corresponding to the second set of data objects to determine whether the second set of data objects is approved to be used to train a second artificial intelligence model, and use a second set of policies that instruct usage metrics corresponding to artificial intelligence model predictions to determine whether the output of the second artificial intelligence model is approved to be provided to one or more computing systems, and further perform an operation including applying the second artificial intelligence model to generate an intended result in response to (i) the second set of data objects being approved to be used to train the second artificial intelligence model and (ii) the output of the second artificial intelligence model being approved to be provided to one or more computing systems, the medium of Section 16.

[0195] Section 18. Determining a set of clauses corresponding to a user-specified query is parsing the user-specified query for a set of keywords, wherein each keyword in the set of keywords is associated with a set of data objects, parsing, and for each keyword in the set of keywords associated with the set of data objects, determining a set of semantically similar clauses corresponding to each respective keyword in the set of keywords, and using the set of semantically similar clauses corresponding to each keyword in the set of keywords to determine a set of clauses corresponding to the user-specified query, the medium of Section 14 further comprising.

[0196] The medium of section 18 further includes accessing a database that indicates a mapping between a first set of keywords and a second set of keywords by determining semantically similar sentences corresponding to each keyword in the set of keywords, and in response to accessing the database, determining a set of semantically similar sentences corresponding to each keyword using each keyword.

[0197] The medium of section 14 includes accessing a metadata graph by traversing each node in a set of nodes to identify a metadata identifier that matches at least one sentence in a set of sentences, and in response to determining that the metadata identifier matches at least one sentence in the set of sentences, determining a node corresponding to the set of sentences. US01 Approved

[0198] A system for reducing data search time when accessing siloed storage data across heterogeneous locations by generating an integrated metadata graph via a retrieval augmented generation (RAG) framework, comprising at least one hardware processor and, when executed by the at least one hardware processor, receiving raw data comprising a set of metadata identifiers that indicate (i) file-level metadata identifiers, (ii) container-level metadata identifiers, and (iii) system-level metadata identifiers, from a set of data silos; selecting a first structured large language model (LLM) prompt corresponding to a first metadata identifier of the set of metadata identifiers from a set of structured LLM prompts; augmenting the first structured LLM prompt with the first metadata identifier that is to be provided to an LLM communicatively coupled to a set of domain-specific ontologies, wherein the LLM is configured to generate a first intermediate output that indicates a second set of metadata identifiers corresponding to the first metadata identifier without accessing the set of domain-specific ontologies; augmenting the first structured LLM prompt with the second set of metadata identifiers corresponding to the first metadata identifier that is to be provided to the LLM, wherein the LLM is configured to generate a second intermediate output that indicates filtered domain-specific metadata identifiers by accessing the set of domain-specific ontologies; generating a domain-specific integrated metadata graph via the LLM using (i) the first metadata identifier and (ii) the second intermediate output that indicates the filtered domain-specific metadata identifiers, wherein the filtered domain-specific metadata identifiers are traversable identifiers and the first metadata identifier is a non-traversable identifier within the domain-specific integrated metadata graph; performing a validation process on the domain-specific integrated metadata graph by comparing a first performance metric of the domain-specific integrated metadata graph to a second performance metric of another version of the domain-specific integrated metadata graph, and wherein the first performance metric isA system comprising at least one non-transitory memory storing instructions that cause the system to perform an update process on a domain-specific integrated metadata graph in response to determining that the domain-specific integrated metadata graph cannot meet or exceed a second performance metric of another version of the domain-specific integrated metadata graph.

[0199] The system of section 21, wherein the set of raw data is received by performing a crawling process across a set of data silos associated with an entity to obtain raw data comprising a set of metadata identifiers.

[0200] The system of section 21, wherein the instructions, when executed by at least one hardware processor, further cause the system to extract a first value from each of the set of data silos, determine a data type corresponding to each first value extracted from the set of data silos, and generate a data profile for each of the data silos in the set of data silos that indicates the data type of the values stored in the data silo.

[0201] The system of section 23, further comprising: selecting a first structured LLM prompt corresponding to a first metadata identifier from the set of structured LLM prompts to determine a data silo storing data corresponding to the first metadata identifier; retrieving a first data profile corresponding to the data silo storing data corresponding to the first metadata identifier; filtering the set of structured LLM prompts using the first data profile to generate a set of filtered LLM prompts; and selecting the first structured LLM prompt corresponding to the first metadata identifier from the set of filtered LLM prompts.

[0202] Section 25. The confirmation process provides a first query that requests the location of a first data item to each of (i) a domain-specific integrated metadata graph, and (ii) another version of the domain-specific integrated metadata graph, wherein providing the first query causes the generation of a first result that indicates the location of the first data item from the domain-specific integrated metadata graph and a second result that indicates the location of the first data item for another version of the domain-specific integrated metadata graph; calculates a first sub-performance measurement criterion for the domain-specific integrated metadata graph and a second sub-performance measurement criterion for another version of the domain-specific integrated metadata graph, wherein the first sub-performance measurement criterion and the second sub-performance measurement criterion are query-to-result performance measurement criteria; calculates a third sub-performance measurement criterion for the domain-specific integrated metadata graph and a fourth sub-performance measurement criterion for another version of the domain-specific integrated metadata graph, wherein the third sub-performance measurement criterion and the fourth sub-performance measurement criterion are result accuracy measurement criteria, the first performance measurement criterion comprises the first sub-performance measurement criterion and the third sub-performance measurement criterion, and the second performance measurement criterion comprises the second sub-performance measurement criterion and the fourth sub-performance measurement criterion; and further includes the system of Section 21.

[0203] Section 26. The system of Section 25, wherein it is determined that the first performance measurement criterion cannot meet or exceed the second performance measurement criterion, based on (i) determining that the first sub-performance measurement criterion meets or exceeds the second sub-performance measurement criterion, or (ii) determining that the third sub-performance measurement criterion cannot meet or exceed the fourth sub-performance measurement criterion.

[0204] Performing an update process on a domain-specific integrated metadata graph is to determine a set of inconsistencies between the domain-specific integrated metadata graph and other versions of the domain-specific integrated metadata graph, where the set of inconsistencies reflects (i) nodes of the domain-specific integrated metadata graph and other versions of the domain-specific integrated metadata graph, and (ii) inconsistencies between edges connected to at least one node of the domain-specific integrated metadata graph and other versions of the domain-specific integrated metadata graph, and updating the domain-specific integrated metadata graph with updated nodes and edges of other versions of the domain-specific integrated metadata graph corresponding to the set of inconsistencies, the system of Section 21.

[0205] Section 28. When the instruction is executed by at least one hardware processor, causing the system to further detect the addition of a data silo to a computing environment associated with a first entity, and in response to detecting the addition of the data silo, cause a second update process to be performed on the domain-specific integrated metadata graph, the system of Section 21.

[0206] Section 29. A file-level metadata identifier indicates the metadata of a given data object stored within each data silo of a set of data silos, a container-level metadata identifier indicates the metadata of each data silo of the set of data silos, and a system-level metadata identifier indicates the metadata of a computing system hosting each data silo of the set of data silos, the system of Section 21.

[0207] The system of section 21, wherein the domain-specific integrated metadata graph comprises (i) a first node that indicates (a) a filtered metadata identifier, (b) a first metadata identifier, and (c) a location identifier of a data silo associated with the first metadata identifier, and (ii) at least one edge that indicates a data lineage between the first node and a second node.

[0208] A method for reducing data search time when accessing siloed storage data spanning heterogeneous locations by generating an integrated metadata graph via a retrieval augmented generation (RAG) framework, the method comprising: selecting a first LLM prompt corresponding to a first metadata identifier from a set of metadata identifiers; expanding the first LLM prompt with the first metadata identifier to be provided to the LLM, wherein the LLM is configured to generate a first intermediate output that indicates a second set of metadata identifiers corresponding to the first metadata identifier; expanding the first LLM prompt with a second set of metadata identifiers corresponding to the first metadata identifier to be provided to the LLM, wherein the LLM is configured to generate a second intermediate output that indicates filtered domain-specific metadata identifiers by accessing a set of domain-specific ontologies; generating a domain-specific integrated metadata graph via the LLM using (i) the first metadata identifier and (ii) the second intermediate output that indicates the filtered domain-specific metadata identifiers, wherein the filtered domain-specific metadata identifiers are traversable identifiers and the first metadata identifier is a non-traversable identifier within the domain-specific integrated metadata graph; and performing an update process on the domain-specific integrated metadata graph in response to determining that a first performance metric of the domain-specific integrated metadata graph fails to meet a performance measure with respect to a second performance metric of another version of the domain-specific integrated metadata graph.

[0209] The method of section 31, further comprising: extracting a first value from each of a set of data silos; determining a data type corresponding to each first value extracted from the set of data silos; and generating a data profile for each of the data silos in the set of data silos that indicates the data type of the values stored in the data silo.

[0210] Selecting, from a set of LLM prompts, a first LLM prompt corresponding to a first metadata identifier of a set of metadata identifiers further includes determining a data silo storing data corresponding to the first metadata identifier, retrieving a first data profile corresponding to the data silo storing data corresponding to the first metadata identifier, filtering the set of LLM prompts using the first data profile to generate a set of filtered LLM prompts, and selecting, from the set of filtered LLM prompts, the first LLM prompt corresponding to the first metadata identifier of the set of metadata identifiers. The method of section 32.

[0211] Step 34. To perform a verification process on a domain-specific integrated metadata graph, the verification process provides a first query that requests the location of a first data item to (i) the domain-specific integrated metadata graph and (ii) each of the other versions of the domain-specific integrated metadata graph, and the provision of the first query causes the generation of a first result that indicates the location of the first data item from the domain-specific integrated metadata graph and a second result that indicates the location of the first data item for the other version of the domain-specific integrated metadata graph. Calculate a first sub-performance measurement criterion for the domain-specific integrated metadata graph and a second sub-performance measurement criterion for the other version of the domain-specific integrated metadata graph, where the first sub-performance measurement criterion and the second sub-performance measurement criterion are query-to-result performance measurement criteria. Calculate a third sub-performance measurement criterion for the domain-specific integrated metadata graph and a fourth sub-performance measurement criterion for the other version of the domain-specific integrated metadata graph, where the third sub-performance measurement criterion and the fourth sub-performance measurement criterion are result accuracy measurement criteria. The first performance measurement criterion comprises the first sub-performance measurement criterion and the third sub-performance measurement criterion, and the second performance measurement criterion comprises the second sub-performance measurement criterion and the fourth sub-performance measurement criterion. Further including performing the method of Step 31 that further includes performing the above steps.

[0212] Step 35. Determining that the first performance measurement criterion meets the performance metric for the second performance measurement criterion is based on (i) determining that the first sub-performance measurement criterion meets or exceeds the second sub-performance measurement criterion, or (ii) determining that the third sub-performance measurement criterion neither meets nor exceeds the fourth sub-performance measurement criterion. The method of Step 34.

[0213] When executed by one or more processors, selecting a first LLM prompt corresponding to a first metadata identifier from a set of metadata identifiers, and expanding the first LLM prompt with the first metadata identifier to be provided to the LLM, wherein the LLM is configured to generate a first intermediate output that indicates a second set of metadata identifiers corresponding to the first metadata identifier; expanding the first LLM prompt with a second set of metadata identifiers corresponding to the first metadata identifier to be provided to the LLM, wherein the LLM is configured to generate a second intermediate output that indicates filtered domain-specific metadata identifiers by accessing a set of domain-specific ontologies; generating a domain-specific integrated metadata graph via the LLM using (i) the first metadata identifier and (ii) the second intermediate output that indicates the filtered domain-specific metadata identifiers, wherein the filtered domain-specific metadata identifiers are traversable identifiers and the first metadata identifier is a non-traversable identifier within the domain-specific integrated metadata graph; and performing an update process on the domain-specific integrated metadata graph in response to determining that a first performance metric of the domain-specific integrated metadata graph fails to meet a performance measure with respect to a second performance metric of another version of the domain-specific integrated metadata graph. One or more non-transitory computer-readable media storing instructions that cause the operations described above to be performed.

[0214] The medium of Section 36, wherein the instructions, when executed by one or more processors, further cause operations including extracting a first value from each of a set of data silos, determining a data type corresponding to each first value extracted from the set of data silos, and generating a data profile for each of the data silos in the set of data silos that indicates the data type of the values stored in the data silo.

[0215] Selecting, from a set of LLM prompts, a first LLM prompt corresponding to a first metadata identifier of a set of metadata identifiers, determining a data silo storing data corresponding to the first metadata identifier, retrieving a first data profile corresponding to the data silo storing data corresponding to the first metadata identifier, filtering the set of LLM prompts using the first data profile to generate a set of filtered LLM prompts, and further selecting, from the set of filtered LLM prompts, a first LLM prompt corresponding to the first metadata identifier of the set of metadata identifiers, the medium of clause 37 further comprising.

[0216] Section 39. When the command is executed by one or more processors, perform a verification process on the domain-specific integrated metadata graph, the verification process comprising providing a first query that requests the location of a first data item to each of (i) the domain-specific integrated metadata graph, and (ii) another version of the domain-specific integrated metadata graph, the providing of the first query causing the generation of a first result that indicates the location of the first data item from the domain-specific integrated metadata graph and a second result that indicates the location of the first data item for another version of the domain-specific integrated metadata graph, calculating a first sub-performance measurement criterion for the domain-specific integrated metadata graph and a second sub-performance measurement criterion for another version of the domain-specific integrated metadata graph, the first sub-performance measurement criterion and the second sub-performance measurement criterion being query-to-result performance measurement criteria, calculating a third sub-performance measurement criterion for the domain-specific integrated metadata graph and a fourth sub-performance measurement criterion for another version of the domain-specific integrated metadata graph, the third sub-performance measurement criterion and the fourth sub-performance measurement criterion being result accuracy measurement criteria, the first performance measurement criterion comprising the first sub-performance measurement criterion and the third sub-performance measurement criterion, the second performance measurement criterion comprising the second sub-performance measurement criterion and the fourth sub-performance measurement criterion, and further comprising causing the medium of Section 36 to perform further operations including performing an operation that includes performing.

[0217] Section 40. The medium of Section 39, wherein determining that the first performance measurement criterion cannot meet the performance measure for the second performance measurement criterion is based on (i) determining that the first sub-performance measurement criterion meets or exceeds the second sub-performance measurement criterion, or (ii) determining that the third sub-performance measurement criterion neither meets nor exceeds the fourth sub-performance measurement criterion. US02 Approved

[0218] Receiving, in a computer system, a first natural language input from a user that includes a set of sentences, and instructions for analyzing data associated with the set of sentences using an artificial intelligence (AI) model; accessing a metadata graph to determine nodes corresponding to the set of sentences in response to the first natural language input, the metadata graph comprising: (i) a set of nodes comprising (a) metadata indicating internal data objects stored in a data silo, and (b) a location identifier of the data silo; and (ii) edges indicating data lineages of the set of nodes; processing one or more internal data objects indicated by the determined nodes to generate a first set of application data; applying the AI model to the first set of application data to generate one or more first outputs, the one or more first outputs comprising a classification of data items in the first set of application data, or one or more predictions made based on the first set of application data; transmitting a representation of the one or more first outputs for display to the user; receiving a second natural language input from the user, the second natural language input including instructions for modifying the first set of application data; generating, by the computer system, a second set of application data based on the received second natural language input; and applying the AI model to the second set of application data to generate one or more second outputs.

[0219] The method of section 41, wherein processing the internal data objects to generate the first set of application data includes removing personally identifiable or private information from the internal data objects.

[0220] The method of section 41, wherein processing an internal data object to generate a first set of application data comprises applying a set of data modification operators to the internal data object to generate a set of modified data based on the internal data object, and the first set of application data includes one or more modified data items from the set of modified data.

[0221] The method of section 41, wherein processing an internal data object to generate a first set of application data includes removing inaccurate data items from the internal data object.

[0222] The method of section 41 further includes receiving a third natural language input including instructions for generating an AI model before receiving a first natural language input, accessing a metadata graph to determine a node corresponding to the third natural language input, processing an internal data object associated with the node corresponding to the third natural language input to generate a set of training data, and training the AI model using the set of training data to generate a trained AI model, wherein applying the trained AI model to a first set of application data includes applying the trained AI model to the first set of application data.

[0223] The method of section 45, wherein training the AI model includes accessing values of one or more model measurement criteria associated with each of a plurality of model types, selecting, by a computer system, a model type of the trained AI model from among the plurality of model types based on the accessed values of the one or more model measurement criteria, and training the selected model type.

[0224] The method of section 45, wherein training the AI model includes receiving a user selection regarding the model type of the AI model and training the selected model type.

[0225] The method of section 41, further including generating, by a computer system, a chat interface for display to a user, receiving a first natural language input via the chat interface, and displaying one or more outputs via the chat interface, including sending a representation of the one or more outputs for display to the user.

[0226] The method of section 41, wherein generating a second set of application data based on a second natural language input includes adding a data item to a first set of application data, removing a data item from the first set of application data, or applying a data modification operator to a value to modify the value of a data item in the first set of application data.

[0227] Section 50. One or more processors and one or more non - transitory computer - readable storage media storing executable instructions, wherein the instructions, when executed by the one or more processors, receive a first natural - language input from a user comprising a set of clauses and instructions for analyzing data associated with the set of clauses using an artificial - intelligence (AI) model, access a metadata graph to determine nodes corresponding to the set of clauses in response to the first natural - language input, wherein the metadata graph comprises (i) a set of nodes comprising (a) metadata indicating internal data objects stored in data silos and (b) location identifiers of the data silos, and (ii) edges indicating data lineages of the set of nodes, process one or more internal data objects indicated by the determined nodes to generate a first set of application data, apply the AI model to the first set of application data to generate one or more first outputs, transmit a representation of the one or more first outputs for display to the user, receive a second natural - language input from the user, wherein the second natural - language input comprises instructions for modifying the first set of application data, generate a second set of application data based on the received second natural - language input, and cause the system to apply the AI model to the second set of application data to generate one or more second outputs, a system comprising the one or more non - transitory computer - readable storage media.

[0228] Section 51. The system of Section 50, wherein processing the internal data objects to generate the first set of application data includes removing personally - identifiable or private information from the internal data objects.

[0229] Section 52. The system of Section 50, wherein processing an internal data object to generate a first set of application data comprises applying a set of data modification operators to the internal data object to generate a set of modified data based on the internal data object, and the first set of application data includes one or more modified data items from the set of modified data.

[0230] Section 53. The system of Section 50, wherein processing an internal data object to generate a first set of application data includes removing inaccurate data items from the internal data object.

[0231] Section 54. When instructions are executed by one or more processors, the system of Section 50 is further caused to receive a third natural language input including instructions for generating an AI model before receiving a first natural language input, access a metadata graph to determine nodes corresponding to the third natural language input, process internal data objects associated with the nodes corresponding to the third natural language input to generate a set of training data, and train the AI model using the set of training data to generate a trained AI model, wherein training includes applying the trained AI model to a first set of application data.

[0232] Section 55. When instructions are executed by one or more processors, the system of Section 50 is further caused to generate a chat interface for display to a user, wherein the first natural language input is received via the chat interface, and sending one or more output representations for display to the user includes displaying one or more outputs via the chat interface.

[0233] The system of section 50, wherein generating a second set of application data based on a second natural language input includes adding a data item to a first set of application data, removing a data item from the first set of application data, or applying a data modification operator to a value to modify the value of a data item in the first set of application data.

[0234] A non-transitory computer-readable storage medium storing executable instructions that, when executed by one or more processors of a system, cause the system to: receive a first natural language input from a user that includes a set of clauses, and instructions for analyzing data associated with the set of clauses using an artificial intelligence (AI) model; in response to the first natural language input, access a metadata graph to determine a node corresponding to the set of clauses, the metadata graph comprising: (i) a set of nodes comprising: (a) metadata indicating internal data objects stored in a data silo, and (b) a location identifier of the data silo; and (ii) edges indicating data lineages of the set of nodes; process one or more internal data objects indicated by the determined node to generate a first set of application data; apply the AI model to the first set of application data to generate one or more first outputs; transmit a representation of the one or more first outputs for display to the user; receive a second natural language input from the user, the second natural language input including instructions for modifying the first set of application data; generate a second set of application data based on the received second natural language input; and apply the AI model to the second set of application data to generate one or more second outputs.

[0235] Section 58. Processing internal data objects to generate a first set of application data involves applying a set of data modification operators to the internal data objects to generate a set of modified data based on the internal data objects, where the first set of application data includes one or more modified data items from the set of modified data, the non - transitory computer - readable storage medium of Section 57 including applying.

[0236] Section 59. When instructions are executed by one or more processors, the system is further caused to receive a third natural language input including instructions for generating an AI model before receiving a first natural language input, access a metadata graph to determine a node corresponding to the third natural language input, process internal data objects associated with the node corresponding to the third natural language input to generate a set of training data, and train the AI model using the set of training data to generate a trained AI model, where applying the trained AI model to a first set of application data includes applying the trained AI model to the first set of application data, the non - transitory computer - readable storage medium of Section 57 including training.

[0237] Section 60. When instructions are executed by one or more processors, the system is further caused to generate a chat interface for display to the user, where the first natural language input is received via the chat interface, and sending one or more output representations for display to the user includes displaying one or more outputs via the chat interface, the non - transitory computer - readable storage medium of Section 57. US03 pending

[0238] Section 61. A first request for deploying a first artificial intelligence (AI) model to make the first AI model available for use in a production environment for processing input data and generating corresponding output is received by an entity for the first AI model, and based on a model deployment engine, a first model deployment location for the first AI model is selected, where the model deployment engine is configured to select the first model deployment location from a set of one or more cloud provider environments or an on-premises environment operated by the entity, generating a script for deploying the first AI model to the first model deployment location, monitoring operation parameters associated with the deployment of the first AI model at the selected model deployment location when the first AI model processes input data and generates corresponding output, updating the model deployment engine based on the values of the monitored operation parameters, and in response to a second request for deploying a second AI model, selecting a second model deployment location for the second AI model based on the updated model deployment engine. A computer-executed method comprising the steps above.

[0239] Section 62. The computer-executed method of Section 61, wherein the model deployment engine comprises a decision tree, a knowledge graph, or a rule engine, and selecting the first model deployment location based on the model deployment engine comprises selecting the first model deployment location based on one or more of the computational cost for deploying the first AI model at the selected model deployment location, the computational capacity of the selected model deployment location, the response time from the selected model deployment location, or a measured value of the accuracy of the first AI model when deployed at the selected model deployment location.

[0240] The computer-executable method of section 61, wherein selecting a first model deployment location based on a model deployment engine includes selecting the first model deployment location based on a privacy policy associated with at least one of the first AI model, input data, or corresponding output.

[0241] The computer-executable method of section 61, wherein the operating parameters include one or more of a computational cost used by the first AI model at the first model deployment location, a response time from the first AI model when deployed at the first model deployment location, a computational capacity available in an environment including the first model deployment location, a privacy policy of an environment including the first model deployment location, or a measurement of a downtime or service interruption of an environment including the first model deployment location.

[0242] The computer-executable method of section 61, wherein the model deployment engine comprises a trained decision model, and updating the model deployment engine includes retraining the trained decision model based on a difference between a monitored value of an operating parameter and a set of values of operating parameters for which the model deployment engine was trained.

[0243] The computer-executable method of section 61, wherein the second AI model is a second instance of the first AI model, and a second model deployment location for the second AI model is a location different from the first model deployment location for the first AI model.

[0244] The computer-executable method of section 61, further comprising selecting, based on the updated model deployment engine, a third model deployment location for the first AI model that is different from the first model deployment location, and generating a script for deploying the first AI model to the third model deployment location.

[0245] The computer-executed method of section 61, further comprising: selecting, based on the model deployment engine and the operation parameters, a third model deployment location for the first AI model that is different from the first model deployment location; and generating a script for deploying the first AI model to the third model deployment location.

[0246] Section 69. The computer-executed method of section 68, wherein the operation parameters include the calculation cost associated with the deployment of the first AI model at the first model deployment location, and selecting the third model deployment location includes determining to move the first AI model to the third model deployment location when the calculation cost associated with the deployment of the first AI model at the first model deployment location is greater than the predicted calculation cost associated with the deployment of the first AI model at the third model deployment location.

[0247] Section 70. One or more processors and one or more non-transitory computer-readable storage media storing executable instructions, the instructions, when executed by the one or more processors, cause the entity to receive, for a first artificial intelligence (AI) model used by the entity, a first request to deploy the first AI model and make it available for use in a production environment for processing input data and generating corresponding output, select, based on a model deployment engine, a first model deployment location for the first AI model, where the model deployment engine is configured to select the first model deployment location from a set of one or more cloud provider environments or an on-premises environment operated by the entity, generate a script for deploying the first AI model to the first model deployment location, monitor operation parameters associated with the deployment of the first AI model at the selected model deployment location when the first AI model processes input data and generates corresponding output, update the model deployment engine based on values of the monitored operation parameters, and cause the system to select, in response to a second request to deploy a second AI model, a second model deployment location for the second AI model based on the updated model deployment engine, a system comprising the one or more non-transitory computer-readable storage media.

[0248] Section 71. The system of Section 70, wherein the model deployment engine comprises a decision tree, a knowledge graph, or a rule engine, and selecting the first model deployment location based on the model deployment engine includes selecting the first model deployment location based on one or more of a computational cost for deploying the first AI model at the selected model deployment location, a computational capacity of the selected model deployment location, a response time from the selected model deployment location, or a measure of accuracy of the first AI model when deployed at the selected model deployment location.

[0249] Section 72. The system of Section 70, wherein selecting a first model deployment location based on a model deployment engine includes selecting the first model deployment location based on a privacy policy associated with at least one of the first AI model, input data, or corresponding output.

[0250] Section 73. The system of Section 70, wherein the operating parameters include one or more of the computational cost used by the first AI model at the first model deployment location, the response time from the first AI model when deployed at the first model deployment location, the computational capacity available in the environment including the first model deployment location, the privacy policy of the environment including the first model deployment location, or the measurement of downtime or service interruption of the environment including the first model deployment location.

[0251] Section 74. The system of Section 70, wherein the model deployment engine comprises a trained decision model, and updating the model deployment engine includes retraining the trained decision model based on the difference between the monitored value of the operating parameter and the set of values of the operating parameters on which the model deployment engine was trained.

[0252] Section 75. The system of Section 70, wherein the second AI model is a second instance of the first AI model, and the second model deployment location for the second AI model is a location different from the first model deployment location for the first AI model.

[0253] Section 76. The system of Section 70, wherein when the instruction is executed by one or more processors, the system is further caused to select a third model deployment location for the first AI model that is different from the first model deployment location, and generate a script for deploying the first AI model to the third model deployment location, based on the updated model deployment engine.

[0254] Section 77. When the command is executed by one or more processors, the system is further caused to select, based on the model deployment engine and the operation parameters, a third model deployment location for the first AI model that is different from the first model deployment location, and generate a script for deploying the first AI model to the third model deployment location, of the system of Section 70.

[0255] Section 78. The operation parameters include the computational cost associated with the deployment of the first AI model at the first model deployment location, and selecting the third model deployment location includes determining to move the first AI model to the third model deployment location when the computational cost associated with the deployment of the first AI model at the first model deployment location is greater than the predicted computational cost associated with deploying the first AI model at the third model deployment location, of the system of Section 77.

[0256] A non-transitory computer-readable storage medium storing executable instructions, which, when executed by one or more processors of a system, cause the system to receive, for a first artificial intelligence (AI) model used by an entity, a first request to deploy the first AI model to make it available for use in a production environment for processing input data and generating corresponding output; select, based on a model deployment engine, a first model deployment location for the first AI model, the model deployment engine being configured to select the first model deployment location from a set of one or more cloud provider environments or an on-premises environment operated by the entity; generate a script for deploying the first AI model to the first model deployment location; monitor operation parameters associated with the deployment of the first AI model at the selected model deployment location when the first AI model processes input data and generates corresponding output; update the model deployment engine based on values of the monitored operation parameters; and select, in response to a second request to deploy a second AI model, a second model deployment location for the second AI model based on the updated model deployment engine.

[0257] Article 80. The non-transitory computer-readable storage medium of Article 79, wherein the model deployment engine comprises a decision tree, a knowledge graph, or a rule engine, and selecting the first model deployment location based on the model deployment engine comprises selecting the first model deployment location based on one or more of a computational cost for deploying the first AI model at the selected model deployment location, a computational capacity of the selected model deployment location, a response time from the selected model deployment location, or a measure of accuracy of the first AI model when deployed at the selected model deployment location. US04 pending claims

[0258] A system for reducing the amount of computing resources used when accessing silo storage data spanning heterogeneous locations via an integrated metadata graph, the system comprising at least one hardware processor and, when executed by the at least one hardware processor, identifying a set of keywords associated with a request to access a set of data objects, performing natural language processing on the set of keywords to determine a set of semantically similar phrases corresponding to each keyword in the set of keywords, accessing a metadata graph to determine nodes corresponding to the set of semantically similar phrases, the metadata graph comprising (i) a set of nodes that indicate (a) metadata of internal data objects stored in a data silo and (b) location identifiers of the data silo, and (ii) edges that indicate data lineages between a first node and a second node in the set of nodes, the metadata graph being generated using a metadata data structure based on file-level and container-level metadata identifiers, accessing, using the location identifier corresponding to the determined nodes, to determine a data silo storing at least one data object of the set of data objects in order to obtain at least one data object of the set of data objects via the data silo, and generating a visual representation of at least one data object for display on a graphical user interface (GUI), the visual representation of at least one data object comprising lineage information of the at least one data object, and at least one non-transitory memory storing instructions that cause the system to perform the generating.

[0259] S82. The metadata graph extracts from a second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, where each file-level metadata identifier in the set of file-level metadata identifiers indicates the metadata of a given data object stored within a respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates the metadata of a respective data silo in the second set of data silos; generating respective sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier; generating a metadata data structure for mapping each semantically similar metadata identifier in the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; and generating a metadata graph using the generated metadata data structure. The system of S81 is generated by the above.

[0260] S83. Receiving, via a second GUI, a second user-specified query that instructs a request to generate an intended result; providing the second user-specified query to an artificial intelligence model to generate a recommendation, where the recommendation comprises (i) a second artificial intelligence model that is to be used to generate the intended result and (ii) a second set of data objects that is to be used when training the second artificial intelligence model; in response to receiving a user selection that indicates acceptance of the recommendation, further comprising instructions to (i) access a database to obtain the second artificial intelligence model and (ii) use the metadata graph to obtain the second set of data objects; training the second artificial intelligence model using the set of data objects; and applying the second artificial intelligence model to generate the intended result. The system of S81 is further provided with the above.

[0261] To obtain a set of policies that indicate usage metrics corresponding to a second set of data objects, access a governance database, and use the set of policies that indicate usage metrics corresponding to the second set of data objects to determine whether the second set of data objects is approved to be used to train a second artificial intelligence model, and use a second set of policies that indicate usage metrics corresponding to artificial intelligence model predictions to determine whether the output of the second artificial intelligence model is approved to be provided to one or more computing systems, and in response to (i) the second set of data objects being approved to be used to train a second artificial intelligence model and (ii) the output of the second artificial intelligence model being approved to be provided to one or more computing systems, further comprising instructions to apply the second artificial intelligence model to generate the intended result, the system of section 83.

[0262] A method for reducing the amount of computing resources used when accessing silo storage data spanning heterogeneous locations via an integrated metadata graph, the method comprising: identifying a set of keywords associated with a user-specified query for accessing a set of data objects; performing natural language processing on the user-specified query to determine a set of clauses corresponding to the user-specified query; accessing a metadata graph to determine nodes corresponding to the set of clauses, the metadata graph comprising: (i) a set of nodes comprising (a) metadata indicating internal data objects stored in a data silo, and (b) a location identifier of the data silo, and (ii) edges indicating data lineages of the set of nodes, the metadata graph being generated using a metadata data structure based on file-level and container-level metadata identifiers; accessing; using the location identifier corresponding to the determined nodes to determine a data silo storing at least one data object of the set of data objects in order to obtain at least one data object of the set of data objects via the data silo; and generating a representation of the at least one data object.

[0263] Section 86. The metadata graph extracts from a second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, where each file-level metadata identifier in the set of file-level metadata identifiers indicates the metadata of a given data object stored within a respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates the metadata of a respective data silo in the second set of data silos; generating respective sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier; generating a metadata data structure for mapping each semantically similar metadata identifier in the sets of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; and generating a metadata graph using the generated metadata data structure. The method of Section 85 is generated by the above steps.

[0264] Section 87. Receiving, via a second GUI, a second user-specified query that instructs a request to generate an intended result; providing the second user-specified query to an artificial intelligence model to generate a recommendation, where the recommendation comprises (i) a second artificial intelligence model that is to be used to generate the intended result and (ii) a second set of data objects that is to be used when training the second artificial intelligence model; in response to receiving a user selection that indicates acceptance of the recommendation, (i) accessing a database to obtain the second artificial intelligence model and (ii) using the metadata graph to obtain the second set of data objects; training the second artificial intelligence model using the set of data objects; and applying the second artificial intelligence model to generate the intended result. The method of Section 85 further includes the above steps.

[0265] To obtain a set of policies that indicate usage metrics corresponding to a second set of data objects, access a governance database, and use the set of policies that indicate usage metrics corresponding to the second set of data objects to determine whether the second set of data objects is approved to be used to train a second artificial intelligence model; use a second set of policies that indicate usage metrics corresponding to artificial intelligence model predictions to determine whether the output of the second artificial intelligence model is approved to be provided to one or more computing systems; and in response to (i) the second set of data objects being approved to be used to train a second artificial intelligence model and (ii) the output of the second artificial intelligence model being approved to be provided to one or more computing systems, apply the second artificial intelligence model to generate the intended result, the method of section 87 further comprising.

[0266] Determining a set of clauses corresponding to a user-specified query comprises parsing the user-specified query for a set of keywords, wherein each keyword in the set of keywords is associated with a set of data objects; for each keyword in the set of keywords associated with the set of data objects, determining a set of semantically similar clauses corresponding to each keyword in the set of keywords; and using the set of semantically similar clauses corresponding to each keyword in the set of keywords to determine a set of clauses corresponding to the user-specified query, the method of section 85 further comprising.

[0267] Determining a semantically similar clause corresponding to each keyword in a set of keywords comprises accessing a database that indicates a mapping between a first set of keywords and a second set of keywords, and in response to accessing the database, using each keyword to determine a set of semantically similar clauses corresponding to each keyword, the method of section 89 further comprising.

[0268] Section 91. The method of Section 85 further includes traversing each node in a set of nodes to identify a metadata identifier that matches at least one clause in a set of clauses, and in response to determining that the metadata identifier matches at least one clause in the set of clauses, determining a node corresponding to the set of clauses.

[0269] Section 92. The method of Section 85 further includes traversing each node in a set of nodes to identify a metadata identifier that matches at least one clause in a set of clauses, and in response to determining that the metadata identifier matches at least one clause in the set of clauses, determining a first node corresponding to the set of clauses, and in response to determining that the first node corresponds to the set of clauses, performing a second traversal of the nodes in the set of nodes using an edge that indicates a first data lineage of the first node, where the first data lineage of the first node indicates a second node that includes information that is a source of information associated with the first node, and using a location identifier corresponding to the second node to determine a second data silo that stores a second data object in a set of data objects in order to obtain the second data object in the set of data objects via the second data silo, and generating a second representation of the second data object.

[0270] Section 93. The method of Section 85, wherein a representation of at least one data object includes lineage information of the at least one data object.

[0271] When executed by one or more processors, perform natural language processing on a user-specified query to determine a set of phrases corresponding to the user-specified query, and access a metadata graph to determine nodes corresponding to the set of phrases, the metadata graph comprising: (i) a set of nodes comprising (a) metadata indicating internal data objects stored in data silos, and (b) location identifiers of the data silos, and (ii) edges indicating data lineages of the set of nodes, the metadata graph being generated using a metadata data structure based on file-level and container-level metadata identifiers; access; use the location identifier corresponding to the determined node to determine a data silo storing at least one data object of the set of data objects in order to obtain at least one data object of the set of data objects via the data silo; and generate a representation of the at least one data object. One or more non-transitory computer-readable media storing instructions that cause the above-described operations to be performed.

[0272] Section 95. The metadata graph extracts from a second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, where each file-level metadata identifier in the set of file-level metadata identifiers indicates the metadata of a given data object stored within a respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates the metadata of a respective data silo in the second set of data silos; generating respective sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier; generating a metadata data structure for mapping each semantically similar metadata identifier in the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; and generating a metadata graph using the generated metadata data structure. The medium of Section 94 is generated by the above.

[0273] Section 96. When instructions are executed by one or more processors, receiving, via a second GUI, a second user-specified query that instructs a request for generating an intended result; providing the second user-specified query to an artificial intelligence model to generate a recommendation, where the recommendation comprises (i) a second artificial intelligence model that is to be used to generate the intended result and (ii) a second set of data objects that is to be used when training the second artificial intelligence model; in response to receiving a user selection that indicates acceptance of the recommendation, further causing operations to be performed, including (i) accessing a database to obtain the second artificial intelligence model and (ii) using the metadata graph to obtain the second set of data objects; training the second artificial intelligence model using the set of data objects; and applying the second artificial intelligence model to generate the intended result. The medium of Section 94 is as described above.

[0274] When the instruction is executed by one or more processors, access the governance database to obtain a set of policies that indicate usage metrics corresponding to a second set of data objects, and use the set of policies that indicate usage metrics corresponding to the second set of data objects to determine whether the second set of data objects is approved to be used to train a second artificial intelligence model, and use a second set of policies that indicate usage metrics corresponding to artificial intelligence model predictions to determine whether the output of the second artificial intelligence model is approved to be provided to one or more computing systems, and in response to (i) the second set of data objects being approved to be used to train a second artificial intelligence model and (ii) the output of the second artificial intelligence model being approved to be provided to one or more computing systems, further perform an operation that includes applying the second artificial intelligence model to generate an intended result. The medium of section 96

[0275] Determining a set of clauses corresponding to a user-specified query further includes parsing the user-specified query for a set of keywords, where each keyword in the set of keywords is associated with a set of data objects, and for each keyword in the set of keywords associated with the set of data objects, determining a set of semantically similar clauses corresponding to each keyword in the set of keywords, and using the set of semantically similar clauses corresponding to each keyword in the set of keywords to determine a set of clauses corresponding to the user-specified query. The medium of section 94

[0276] The medium of clause 98 further includes accessing a database that indicates a mapping between a first set of keywords and a second set of keywords, and in response to accessing the database, determining, using each keyword, a set of semantically similar sentences corresponding to each keyword, for each keyword in the set of keywords.

[0277] The medium of clause 94 further includes traversing each node in a set of nodes to identify a metadata identifier that matches at least one sentence in a set of sentences, and in response to determining that the metadata identifier matches at least one sentence in the set of sentences, determining a node corresponding to the set of sentences, when accessing a metadata graph.

[0278] Conclusion Unless the context clearly requires otherwise, throughout the description and claims of this application, the words "comprise", "comprising", and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is, in the sense of "including, but not limited to". As used herein, the terms "connected", "coupled", or any variation thereof mean any direct or indirect connection or coupling between two or more elements, and the connection or coupling between elements can be physical, logical, or a combination thereof. Additionally, the words "herein", "above", "below", and words of similar import, when used in this application, refer to the application as a whole and not to any particular part of the application. Where the context permits, the words in the detailed description above using the singular or plural number may also include the plural or singular number respectively. The word "or" covers all interpretations of the word, including any one of the items in a list of two or more items, all of the items in the list, and any combination of the items in the list.

[0279] The above detailed description of the examples of the present technology is not intended to be exhaustive or to limit the present technology to the precise forms disclosed above. Specific examples of the present technology have been described above for illustrative purposes, but as will be recognized by those skilled in the art, various equivalent modifications are possible within the scope of the present technology. For example, although a process or block is presented in a given order, alternative implementations can perform a routine having steps or employ a system having blocks in a different order, and some processes or blocks can be deleted, moved, added, re-divided, combined, and / or modified to provide alternative or sub-combinations. Each of these processes or blocks can be implemented in a variety of different ways. Also, although a process or block may be shown as being performed sequentially, these processes or blocks can instead be performed or executed in parallel or at different times. Further, any specific numbers mentioned herein are merely examples, and alternative implementations can employ different values or ranges.

[0280] The teachings of the present technology provided herein can be applied to other systems that are not necessarily the systems described above. The elements and acts of the various examples described above can be combined to provide further implementations of the present technology. Some alternative implementations of the present technology can include not only additional elements to those implementations mentioned above, but also fewer elements.

[0281] These and other modifications are possible to the present technology from the perspective of the detailed description above. The above description illustrates specific examples of the present technology and describes the best mode contemplated, but even though it may appear detailed in the text, the present technology can be practiced in many ways. The details of the system are still encompassed by the technology disclosed herein, while the specific implementation forms thereof can vary considerably. As described above, the specific technical terms used when describing a particular feature or aspect of the present technology should not be taken to imply that the technical terms are redefined herein so as to be limited to any particular property, feature, or aspect of the present technology to which the technical terms are associated. Generally, the terms used in the following claims should not be construed to limit the technology to the specific examples disclosed herein, unless the detailed description section above explicitly defines such terms. Thus, the actual scope of the present technology encompasses not only the disclosed examples but also all equivalent ways of practicing or implementing the technology in the claims.

[0282] To reduce the number of claims, specific aspects of the present technology are presented below in a particular claim format, but the applicant contemplates various aspects of the technology in any number of claim formats. For example, only one aspect of the present technology is recited as a computer-readable medium claim, but other aspects may be embodied in the same way as a computer-readable medium claim or in other forms such as embodied as a means-plus-function claim. Any claim intended to be treated under 35 U.S.C. § 112(f) will begin with the word "means", but the use of the term "for" in any other context is not intended to invoke the treatment under 35 U.S.C. § 112(f). Accordingly, the applicant reserves the right to pursue additional claim formats in any such subsequent application or continuation application after filing this application to pursue such additional claims.

Description of the Reference Numerals

[0283] 100 User interface 102 User-specified query input 104 Result output 106 At least one data object 108 Data lineage information 200 Computer system and other devices 204 Input component 206 Output component 208 Processor 210 Storage 212a Application 212N Application 214a Model 214N Model 216 Network connection component 218 Persistent storage device 220 Computer-readable medium drive 300 Environment 302 Client computing device 302a Client computing device, computing device 302b Client computing device, computing device 302c Client computing device, computing device 302d Client computing device, computing device 304 Network 306 Server computing device 310 Server computing device 310a Server 310b Server 310c Server 308 Database 312 Database 312a Database 312b Database 312c Database 400 Process 500 Metadata graph 502 Node 502a node 502b node 502c node 502d node 502e node 502f node 502g node 502h node 502i node 502j node 502k node 502l node 502m node 504 edge 504a edge 504b edge 504c edge 504d edge 504e edge 504f edge 504g edge 504h edge 504i edge 504j edge 504k edge 504l edge 504m edge 504n edge 504o edge 504p edge 504q edge 600 Enlarged view of the metadata graph 602 node 602a First node, node 602b Second node, node 602c Third node, node 602d Fourth node, node 604 edge 604a First edge 604b Second edge 604c Third edge 606 File-level metadata identifier 606a File-level metadata identifier 608 Container-level metadata identifier 608a Container Level Metadata Identifier 610 Location Identifier 610a Location Identifier 700 Diagram of Artificial Intelligence Model 702 Model 704 Input 706 Output 800 Process 900 LLM Prompt 902 Prompt 1 903 Level 904 First Metadata Identifier 905 First Prompt Text 906 LLM 907 Second Prompt Text 908 First Intermediate Output 910 Prompt 2 912 Second Intermediate Output 914 Prompt 3 916 Prompt 4 918 Third Prompt Text 920 Metadata Graph 922 Prompt 5 1000 Subsystem Diagram 1002 User Interface 1004 Metadata Graph 1006 LLM 1008 Domain Ontology Component 1010 Raw Data Component 1012 Feedback Component 1014a Communication Link 1014b Communication Link 1014c Communication Link 1014d Communication Link 1014e Communication Link 1014f Communication Link 1014g Communication Link 1014h Communication Link 1014i Communication Link 1014j Communication Link 1014k communication link 1014l communication link 1014m communication link 1014n communication link 1014o communication link 1014p communication link 1015 data silo 1016 crawler 1018 parser 1020 profiler 1022 concept 1024 domain thesaurus 1026 versioning component 1028 generator 1030 result 1032 third - party input 1034 update component 1100 metadata graph 1102 node 1102a fifth node, node 1102b sixth node, node 1102c seventh node, node 1102d eighth node, node 1104a second edge 1106a file - level metadata identifier 1108a container - level metadata identifier 1110a location identifier 1112a domain - specific metadata identifier 1200 AI sandbox 1210 model reviewer assistant 1220 data processor 1230 model generator 1240 model governor 1250 model automator 1310 dataset, raw dataset 1320 data pre - processor 1322 application data, application dataset 1330 Data Analyzer 1340 Data Profile Generator 1350 Sampler 1342 Data Profile 1352 Training Sample 1354 Test Sample 1405 User Input 1410 Feature Set Generator 1420 Target Selector, Selector 1430 Model Selector 1440 Model Builder 1445 Model 1510 Large Language Model (LLM) 1520 RAG-based Process 1522 Industry Knowledge Repository 1524 Enterprise Knowledge Repository 1530 Explainability Layer 1600 Process 1620 Process 1700 Chat Interface 1705 Natural Language Input 1710 User Input 1715 Plot 1720 Input 1810 Model Deployment Engine, Engine 1812 Model Deployment Engine Update 1814 Orchestrator 1816 Data Handler 1820A First Public Cloud Provider 1820B Second Public Cloud Provider 1830A On-premises Machine 1830B On-premises Scalable Resources 1900 Process

Claims

1. A system for reducing the usage of computing resources when accessing silo storage data spanning heterogeneous locations via an integrated metadata graph, comprising: at least one hardware processor; when executed by the at least one hardware processor, receiving, in a graphical user interface (GUI), a user-specified query that instructs a request to access a set of data objects, wherein each data object in the set of data objects is stored in a respective data silo of a set of data silos in heterogeneous locations; parsing the user-specified query for a set of keywords, wherein each keyword in the set of keywords is associated with the set of data objects; performing natural language processing on the set of keywords to determine a set of semantically similar phrases corresponding to each keyword in the set of keywords; accessing a metadata graph to determine nodes corresponding to the set of semantically similar phrases, wherein the metadata graph comprises: (i) a set of nodes that indicate (a) metadata of internal data objects stored in data silos, and (b) location identifiers of the data silos, and (ii) edges that indicate data lineages between a first node and a second node in the set of nodes, and wherein the metadata graph is generated using a metadata data structure based on file-level and container-level metadata identifiers; using the location identifiers corresponding to the determined nodes to determine data silos storing at least one data object in the set of data objects in order to obtain at least one data object in the set of data objects via the data silos; and generating a visual representation of the at least one data object for display on the GUI, wherein the visual representation of the at least one data object comprises lineage information of the at least one data object. at least one non - transitory memory storing instructions to cause the system to perform, a system. **Claim 2** wherein the metadata graph retrieves from a second set of data silos: (i) a set of file - level metadata identifiers, and (ii) a set of container - level metadata identifiers, wherein each file - level metadata identifier in the set of file - level metadata identifiers indicates metadata of a given data object stored within a respective data silo, and each container - level metadata identifier in the set of container - level metadata identifiers indicates metadata of the respective data silo of the second set of data silos; retrieving; generating, respectively, a set of semantically similar metadata identifiers corresponding to each file - level and container - level metadata identifier; generating the metadata data structure for mapping each semantically similar metadata identifier in the set of semantically similar metadata identifiers to a normalized file - level metadata identifier and a normalized container - level metadata identifier; generating the metadata graph using the generated metadata data structure; The system according to claim 1, generated by. **Claim 3** receiving, via a second GUI, a second user - specified query that instructs a request to generate an intended result; providing the second user - specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising: (i) a second artificial intelligence model that is to be used to generate the intended result, and (ii) a second set of data objects that is to be used when training the second artificial intelligence model; providing; in response to receiving a user selection indicating acceptance of the recommendation: (i) accessing a database to obtain the second artificial intelligence model, and (ii) using the metadata graph to obtain the second set of data objects; training the second artificial intelligence model using the set of data objects; applying the second artificial intelligence model to generate the intended result; The system according to claim 1, further comprising said instruction for performing [

4. ] Accessing a governance database to obtain a set of policies that indicate usage metrics corresponding to said second set of data objects; Using said set of policies that indicate usage metrics corresponding to said second set of data objects to determine whether said second set of data objects is approved for use in training said second artificial intelligence model; Using a second set of policies that indicate usage metrics corresponding to artificial intelligence model predictions to determine whether the output of said second artificial intelligence model is approved for providing to one or more computing systems; Applying said second artificial intelligence model to generate said intended result in response to (i) said second set of data objects being approved for use in training said second artificial intelligence model and (ii) the output of said second artificial intelligence model being approved for providing to said one or more computing systems The system according to claim 3, further comprising instructions for performing [

5. ] A method for reducing the usage of computing resources when accessing siloed storage data spanning heterogeneous locations via an integrated metadata graph, comprising: Receiving, in a graphical user interface (GUI), a user-specified query that indicates a request to access a set of data objects, wherein each data object of said set of data objects is stored in a respective data silo of a set of data silos in heterogeneous locations; Performing natural language processing on said user-specified query to determine a set of clauses corresponding to said user-specified query; A step of accessing a metadata graph to determine a node corresponding to the set of sentences, wherein the metadata graph comprises: (i) a set of nodes comprising (a) metadata indicating internal data objects stored in a data silo, and (b) a location identifier of the data silo; and (ii) edges indicating data lineages of the set of nodes, and the metadata graph is generated using a metadata data structure based on file-level and container-level metadata identifiers. A step of determining a data silo storing at least one data object of the set of data objects using the location identifier corresponding to the determined node to obtain at least one data object of the set of data objects via the data silo. A step of generating a visual representation of the at least one data object for display on the GUI. A method comprising the above steps. **Claim 6** The metadata graph A step of extracting from a second set of data silos: (i) a set of file-level metadata identifiers, and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier of the set of file-level metadata identifiers indicates metadata of a given data object stored in a respective data silo, and each container-level metadata identifier of the set of container-level metadata identifiers indicates metadata of the respective data silo of the second set of data silos. A step of generating a set of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier respectively. A step of generating the metadata data structure for mapping each semantically similar metadata identifier of the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier. A step of generating the metadata graph using the generated metadata data structure. The method according to claim 5, generated by the above steps. **Claim 7** Receiving, via a second GUI, a second user-specified query that instructs a request to generate an intended result; Providing the second user-specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising: (i) a second artificial intelligence model that is to be used to generate the intended result; and (ii) a second set of data objects that is to be used when training the second artificial intelligence model; In response to receiving a user selection that indicates acceptance of the recommendation: (i) accessing a database to obtain the second artificial intelligence model; and (ii) using the metadata graph to obtain the second set of data objects; Training the second artificial intelligence model using the set of data objects; Applying the second artificial intelligence model to generate the intended result; The method of claim 5, further comprising: Claim 8 Accessing a governance database to obtain a set of policies that indicate usage metrics corresponding to the second set of data objects; Using the set of policies that indicate usage metrics corresponding to the second set of data objects to determine whether use of the second set of data objects to train the second artificial intelligence model is approved; Using a second set of policies that indicate usage metrics corresponding to artificial intelligence model predictions to determine whether providing the output of the second artificial intelligence model to one or more computing systems is approved; In response to: (i) use of the second set of data objects to train the second artificial intelligence model being approved; and (ii) providing the output of the second artificial intelligence model to the one or more computing systems being approved, applying the second artificial intelligence model to generate the intended result; The method of claim 7, further comprising: Claim 9 The step of determining the set of terms corresponding to the user-specified query comprises: Parsing the user-specified query with respect to a set of keywords, wherein each keyword in the set of keywords is associated with the set of data objects; For each keyword in the set of keywords associated with the set of data objects, determining a set of semantically similar sentences corresponding to each of the keywords in the set of keywords; Using the set of semantically similar sentences corresponding to each keyword in the set of keywords to determine the set of sentences corresponding to the user-specified query; The method according to claim 5, further comprising.

10. The step of determining a set of semantically similar sentences corresponding to each of the keywords in the set of keywords comprises: Accessing a database that indicates a mapping between a first set of keywords and a second set of keywords; In response to accessing the database, using each of the keywords to determine the set of semantically similar sentences corresponding to each of the keywords; The method according to claim 9, further comprising.

11. The step of accessing the metadata graph comprises: Traversing each node in the set of nodes to identify a metadata identifier that matches at least one sentence in the set of sentences; In response to determining that the metadata identifier matches at least one sentence in the set of sentences, determining the node corresponding to the set of sentences; The method according to claim 5, further comprising.

12. The step of accessing the metadata graph comprises: Traversing each node in the set of nodes to identify a metadata identifier that matches at least one sentence in the set of sentences; In response to determining that the metadata identifier matches at least one sentence in the set of sentences, determining a first node corresponding to the set of sentences; In response to determining that the first node corresponds to the set of phrases, performing a second traversal of the nodes of the set of nodes using an edge that indicates a first data lineage of the first node, wherein the first data lineage of the first node indicates a second node that comprises information that is a source of information associated with the first node; Determining a second data silo that stores a second data object of the set of data objects, using the location identifier corresponding to the second node, to obtain the second data object of the set of data objects via the second data silo; Generating a second visual representation of the second data object for display on the GUI; The method of claim 5, further comprising. Claim 13 The method of claim 5, wherein the visual representation of the at least one data object comprises lineage information of the at least one data object. Claim 14 When executed by one or more processors, Receiving, in a graphical user interface (GUI), a user-specified query that indicates a request to access a set of data objects, wherein each data object of the set of data objects is stored in a respective data silo of a set of data silos in heterogeneous locations; Performing natural language processing on the user-specified query to determine a set of phrases corresponding to the user-specified query; Accessing a metadata graph to determine nodes corresponding to the set of phrases, wherein the metadata graph comprises: (i) a set of nodes comprising: (a) metadata that indicates internal data objects stored in data silos, and (b) location identifiers of the data silos; and (ii) edges that indicate data lineages of the set of nodes, and wherein the metadata graph is generated using a metadata data structure based on file-level and container-level metadata identifiers; To obtain at least one data object of the set of data objects via the data silo, determining a data silo storing the at least one data object of the set of data objects using the location identifier corresponding to the determined node; Generating a visual representation of the at least one data object for display on the GUI; One or more non-transitory computer-readable media storing instructions that cause an operation including the above.

15. The metadata graph: Retrieving from a second set of data silos (i) a set of file-level metadata identifiers and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier of the set of file-level metadata identifiers indicates metadata of a given data object stored within a respective data silo, and each container-level metadata identifier of the set of container-level metadata identifiers indicates metadata of the respective data silo of the second set of data silos; Generating respective sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier; Generating the metadata data structure for mapping each semantically similar metadata identifier of the sets of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; Generating the metadata graph using the generated metadata data structure; The medium according to claim 14, generated by the above.

16. When the instructions are executed by the one or more processors, Receiving, via a second GUI, a second user-specified query that indicates a request for generating an intended result; Providing the second user-specified query to an artificial intelligence model to generate a recommendation, wherein the recommendation comprises: (i) a second artificial intelligence model that is to be used to generate the intended result; and (ii) a second set of data objects that is to be used when training the second artificial intelligence model In response to receiving a user selection indicating acceptance of the recommendation: (i) accessing a database to obtain the second artificial intelligence model; and (ii) using the metadata graph to obtain the second set of data objects Training the second artificial intelligence model using the set of data objects Applying the second artificial intelligence model to generate the intended result The medium of claim 14, further causing an operation including **Claim 17** When the instructions are executed by the one or more processors Accessing a governance database to obtain a set of policies indicating usage metrics corresponding to the second set of data objects Using the set of policies indicating usage metrics corresponding to the second set of data objects to determine whether the second set of data objects is approved to be used to train the second artificial intelligence model Using a second set of policies indicating usage metrics corresponding to artificial intelligence model predictions to determine whether the output of the second artificial intelligence model is approved to be provided to one or more computing systems In response to (i) the second set of data objects being approved to be used to train the second artificial intelligence model and (ii) the output of the second artificial intelligence model being approved to be provided to the one or more computing systems, applying the second artificial intelligence model to generate the intended result The medium of claim 16, further causing an operation including **Claim 18** Determining the set of phrases corresponding to the user-specified query Parsing the user-specified query with respect to a set of keywords, wherein each keyword in the set of keywords is associated with the set of data objects; For each keyword in the set of keywords associated with the set of data objects, determining a set of semantically similar sentences corresponding to each of the keywords in the set of keywords; Determining a set of sentences corresponding to the user-specified query using the set of semantically similar sentences corresponding to each keyword in the set of keywords; The medium according to claim 14, further comprising.

19. Determining semantically similar sentences corresponding to each of the keywords in the set of keywords includes: Accessing a database that indicates a mapping between a first keyword and a set of second keywords; In response to accessing the database, determining the set of semantically similar sentences corresponding to each of the keywords using each of the keywords; The medium according to claim 18, further comprising.

20. Accessing the metadata graph includes: Traversing each node in the set of nodes to identify a metadata identifier that matches at least one sentence in the set of sentences; Determining the nodes corresponding to the set of sentences in response to determining that the metadata identifier matches at least one sentence in the set of sentences; The medium according to claim 14, further comprising.

21. A system for reducing data search time when accessing silo storage data spanning heterogeneous locations by generating an integrated metadata graph via a retrieval augmentation generation (RAG) framework, the system comprising: At least one hardware processor; When executed by the at least one hardware processor, Receiving raw data comprising a set of metadata identifiers indicating (i) file-level metadata identifiers, (ii) container-level metadata identifiers, and (iii) system-level metadata identifiers, from a set of data silos; Selecting, from a set of structured LLM prompts, a first structured LLM prompt corresponding to a first metadata identifier of the set of metadata identifiers; Expanding the first structured LLM prompt with the first metadata identifier that is to be provided to an LLM communicatively coupled to a set of domain-specific ontologies, the LLM being configured to generate a first intermediate output that indicates a second set of metadata identifiers corresponding to the first metadata identifier without accessing the set of domain-specific ontologies; Expanding the first structured LLM prompt with the second set of metadata identifiers corresponding to the first metadata identifier that is to be provided to the LLM, the LLM being configured to generate a second intermediate output that indicates filtered domain-specific metadata identifiers by accessing the set of domain-specific ontologies; Generating, via the LLM, a domain-specific integrated metadata graph using (i) the first metadata identifier and (ii) the second intermediate output that indicates the filtered domain-specific metadata identifiers, the filtered domain-specific metadata identifiers being traversable identifiers and the first metadata identifier being a non-traversable identifier within the domain-specific integrated metadata graph; Performing a validation process on the domain-specific integrated metadata graph by comparing a first performance metric of the domain-specific integrated metadata graph to a second performance metric of another version of the domain-specific integrated metadata graph, and Performing an update process on the domain-specific integrated metadata graph in response to determining that the first performance metric cannot meet or exceed the second performance metric of the other version of the domain-specific integrated metadata graph At least one non-transitory memory storing instructions causing the system to perform the above; A system comprising the above.

22. The set of raw data is Performing a crawling process across the set of data silos associated with an entity to obtain the raw data comprising the set of metadata identifiers The system according to claim 21, as received thereby. **Claim 23** When the instructions are executed by the at least one hardware processor Extracting a first value from each of the set of data silos Determining a data type corresponding to each first value extracted from each of the set of data silos Generating a data profile for each of the data silos in the set of data silos that indicates the data type of the values stored in the data silo The system according to claim 21, further causing the system to perform. **Claim 24** Selecting the first structured LLM prompt corresponding to the first metadata identifier from the set of structured LLM prompts among the set of metadata identifiers Determining a data silo storing data corresponding to the first metadata identifier Fetching a first data profile corresponding to the data silo storing the data corresponding to the first metadata identifier Filtering the set of structured LLM prompts to generate a set of filtered LLM prompts using the first data profile Selecting the first structured LLM prompt corresponding to the first metadata identifier from the set of filtered LLM prompts The system according to claim 23, further comprising. **Claim 25** The verification process is Providing a first query requesting a location of a first data item to each of (i) the domain-specific integrated metadata graph and (ii) the other version of the domain-specific integrated metadata graph, wherein providing the first query causes generation of a first result indicating the location of the first data item from the domain-specific integrated metadata graph and a second result indicating the location of the first data item for the other version of the domain-specific integrated metadata graph Calculating a first sub-performance measurement criterion for the domain-specific integrated metadata graph and a second sub-performance measurement criterion for the other version of the domain-specific integrated metadata graph, wherein the first sub-performance measurement criterion and the second sub-performance measurement criterion are query-to-result performance measurement criteria, and Calculating a third sub-performance measurement criterion for the domain-specific integrated metadata graph and a fourth sub-performance measurement criterion for the other version of the domain-specific integrated metadata graph, wherein the third sub-performance measurement criterion and the fourth sub-performance measurement criterion are result accuracy measurement criteria, the first performance measurement criterion comprises the first sub-performance measurement criterion and the third sub-performance measurement criterion, and the second performance measurement criterion comprises the second sub-performance measurement criterion and the fourth sub-performance measurement criterion, and The system according to claim 21, further comprising.

26. Determining that the first performance measurement criterion cannot meet or exceed the second performance measurement criterion, based on (i) determining that the first sub-performance measurement criterion meets or exceeds the second sub-performance measurement criterion, or (ii) determining that the third sub-performance measurement criterion cannot meet or exceed the fourth sub-performance measurement criterion. The system according to claim 25.

27. Performing the update process on the domain-specific integrated metadata graph, Determining a set of inconsistencies between the domain-specific integrated metadata graph and the other version of the domain-specific integrated metadata graph, wherein the set of inconsistencies reflects the inconsistencies between (i) the nodes of the domain-specific integrated metadata graph and the other version of the domain-specific integrated metadata graph, and (ii) the edges connected to at least one node of the domain-specific integrated metadata graph and the other version of the domain-specific integrated metadata graph. Determining, Updating the domain-specific integrated metadata graph with the updated nodes and edges of the other version of the domain-specific integrated metadata graph corresponding to the set of inconsistencies, and The system according to claim 21, comprising.

28. When the command is executed by the at least one hardware processor, detecting an addition of a data silo to a computing environment associated with a first entity; in response to detecting the addition of the data silo, causing the second update process to be performed on the domain-specific integrated metadata graph; The system according to claim 21, further causing the system to perform.

29. The file-level metadata identifier indicates metadata of a given data object stored within each respective data silo of the set of data silos, the container-level metadata identifier indicates metadata of each respective data silo of the set of data silos, and the system-level metadata identifier indicates metadata of a computing system hosting each respective data silo of the set of data silos. The system according to claim 21.

30. The domain-specific integrated metadata graph comprises (i) a first node indicating (a) the filtered metadata identifier, (b) the first metadata identifier, (c) a location identifier of the data silo associated with the first metadata identifier, and (ii) at least one edge indicating a data lineage between the first node and a second node. The system according to claim 21.

31. A method for reducing data search time when accessing siloed storage data spanning heterogeneous locations by generating an integrated metadata graph via a Retrieval-Augmented Generation (RAG) framework, comprising: selecting a first LLM prompt corresponding to a first metadata identifier from a set of large language model (LLM) prompts; extending the first LLM prompt with the first metadata identifier to be provided to the LLM, the LLM being configured to generate a first intermediate output indicating a second set of metadata identifiers corresponding to the first metadata identifier. Extending the first LLM prompt with the second set of metadata identifiers corresponding to the first metadata identifier to be provided to the LLM, wherein the LLM is configured to generate a second intermediate output that indicates filtered domain-specific metadata identifiers by accessing a set of domain-specific ontologies; Generating a domain-specific integrated metadata graph via the LLM using (i) the first metadata identifier and (ii) the second intermediate output that indicates the filtered domain-specific metadata identifiers, wherein the filtered domain-specific metadata identifiers are traversable identifiers and the first metadata identifier is a non-traversable identifier within the domain-specific integrated metadata graph; Performing an update process on the domain-specific integrated metadata graph in response to determining that a first performance metric of the domain-specific integrated metadata graph fails to meet a performance measure with respect to a second performance metric of another version of the domain-specific integrated metadata graph; A method comprising.

32. Extracting a first value from each of a set of data silos; Determining a data type corresponding to each first value extracted from each of the set of data silos; Generating a data profile for each of the data silos in the set of data silos that indicates the data type of the values stored in the data silos; The method according to claim 31, further comprising.

33. The step of selecting the first LLM prompt corresponding to the first metadata identifier from the set of LLM prompts comprises: Determining a data silo storing data corresponding to the first metadata identifier; Retrieving a first data profile corresponding to the data silo storing the data corresponding to the first metadata identifier; Filtering the set of LLM prompts using the first data profile to generate a set of filtered LLM prompts; selecting, from the set of filtered LLM prompts, the first LLM prompt corresponding to the first metadata identifier of the set of metadata identifiers The method of claim 32, further comprising. **Claim 34** performing a verification process on the domain-specific integrated metadata graph, the verification process comprising providing, to each of (i) the domain-specific integrated metadata graph and (ii) the other version of the domain-specific integrated metadata graph, a first query that requests a location of a first data item, the providing of the first query causing the generation of a first result that indicates the location of the first data item from the domain-specific integrated metadata graph and a second result that indicates the location of the first data item for the other version of the domain-specific integrated metadata graph calculating a first sub-performance measurement criterion for the domain-specific integrated metadata graph and a second sub-performance measurement criterion for the other version of the domain-specific integrated metadata graph, the first sub-performance measurement criterion and the second sub-performance measurement criterion being query-to-result performance measurement criteria, and calculating a third sub-performance measurement criterion for the domain-specific integrated metadata graph and a fourth sub-performance measurement criterion for the other version of the domain-specific integrated metadata graph, the third sub-performance measurement criterion and the fourth sub-performance measurement criterion being result accuracy measurement criteria, the first performance measurement criterion comprising the first sub-performance measurement criterion and the third sub-performance measurement criterion, and the second performance measurement criterion comprising the second sub-performance measurement criterion and the fourth sub-performance measurement criterion The method of claim 31, further comprising steps, further comprising. **Claim 35** Determining that the first performance measurement criterion meets the performance measure for the second performance measurement criterion is based on (i) the first sub-performance measurement criterion meeting or exceeding the second sub-performance measurement criterion, or (ii) determining that the third sub-performance measurement criterion can neither meet nor exceed the fourth sub-performance measurement criterion, the method according to claim 34.

36. When executed by one or more processors, selecting a first LLM prompt corresponding to a first metadata identifier from a set of metadata identifiers, extending the first LLM prompt with the first metadata identifier to be provided to the LLM, wherein the LLM is configured to generate a first intermediate output that indicates a second set of metadata identifiers corresponding to the first metadata identifier, the extending, extending the first LLM prompt with the second set of metadata identifiers corresponding to the first metadata identifier to be provided to the LLM, wherein the LLM is configured to generate a second intermediate output that indicates filtered domain-specific metadata identifiers by accessing a set of domain-specific ontologies, the extending, generating a domain-specific integrated metadata graph via the LLM using (i) the first metadata identifier and (ii) the second intermediate output that indicates the filtered domain-specific metadata identifiers, wherein the filtered domain-specific metadata identifiers are traversable identifiers and the first metadata identifier is a non-traversable identifier within the domain-specific integrated metadata graph, the generating, performing an update process on the domain-specific integrated metadata graph in response to determining that the first performance measurement criterion of the domain-specific integrated metadata graph cannot meet the performance measure for a second performance measurement criterion of another version of the domain-specific integrated metadata graph and one or more non-transitory computer-readable media storing instructions that cause operations including the above to be performed.

37. When the instructions are executed by the one or more processors, extracting a first value from each of a set of data silos; determining a data type corresponding to each first value extracted from each of the set of data silos; generating a data profile for each of the data silos in the set of data silos that indicates the data type of the value stored in the data silo The medium according to claim 36, further causing an operation including the above to be performed.

38. selecting, from a set of LLM prompts, the first LLM prompt corresponding to the first metadata identifier among the set of metadata identifiers; determining a data silo storing data corresponding to the first metadata identifier; retrieving a first data profile corresponding to the data silo storing the data corresponding to the first metadata identifier; filtering the set of LLM prompts using the first data profile to generate a set of filtered LLM prompts; selecting, from the set of filtered LLM prompts, the first LLM prompt corresponding to the first metadata identifier among the set of metadata identifiers The medium according to claim 37, further including the above.

39. when the instructions are executed by the one or more processors, performing a verification process on the domain-specific integrated metadata graph, the verification process including: providing a first query requesting a location of a first data item to each of (i) the domain-specific integrated metadata graph and (ii) the other version of the domain-specific integrated metadata graph, the providing of the first query causing generation of a first result indicating the location of the first data item from the domain-specific integrated metadata graph and a second result indicating the location of the first data item for the other version of the domain-specific integrated metadata graph; Calculating a first sub-performance measurement criterion for the domain-specific integrated metadata graph and a second sub-performance measurement criterion for the other version of the domain-specific integrated metadata graph, wherein the first sub-performance measurement criterion and the second sub-performance measurement criterion are query-to-result performance measurement criteria, and Calculating a third sub-performance measurement criterion for the domain-specific integrated metadata graph and a fourth sub-performance measurement criterion for the other version of the domain-specific integrated metadata graph, wherein the third sub-performance measurement criterion and the fourth sub-performance measurement criterion are result accuracy measurement criteria, the first performance measurement criterion comprises the first sub-performance measurement criterion and the third sub-performance measurement criterion, and the second performance measurement criterion comprises the second sub-performance measurement criterion and the fourth sub-performance measurement criterion Further including performing The medium according to claim 36, further causing an operation including

40. Determining that the first performance measurement criterion cannot meet the performance metric regarding the second performance measurement criterion is based on (i) determining that the first sub-performance measurement criterion meets or exceeds the second sub-performance measurement criterion, or (ii) determining that the third sub-performance measurement criterion neither meets nor exceeds the fourth sub-performance measurement criterion. The medium according to claim 39

41. Receiving, in a computer system, a first natural language input from a user including a set of sentences and instructions for analyzing data associated with the set of sentences using an artificial intelligence (AI) model; In response to the first natural language input, accessing a metadata graph to determine nodes corresponding to the set of sentences, wherein The metadata graph comprises (i) a set of nodes comprising (a) metadata indicating internal data objects stored in a data silo and (b) a location identifier of the data silo, and (ii) edges indicating data lineages of the set of nodes; Step; Processing one or more internal data objects indicated by the determined nodes to generate a first set of application data; Applying the AI model to the first set of application data to generate one or more first outputs, wherein the one or more first outputs comprise a classification of data items in the first set of application data or one or more predictions made based on the first set of application data, a step; sending a representation of the one or more first outputs for display to the user; receiving a second natural language input from the user, wherein the second natural language input comprises instructions for modifying the first set of application data, generating, by the computer system, a second set of application data based on the received second natural language input; applying the AI model to the second set of application data to generate one or more second outputs A computer-executed method comprising:

42. The method of claim 41, wherein the step of processing the internal data object to generate the first set of application data comprises removing personally identifiable or private information from the internal data object.

43. The step of processing the internal data object to generate the first set of application data comprises applying a set of data modification operators to the internal data object to generate a set of modified data based on the internal data object, wherein the first set of application data comprises one or more modified data items from the set of modified data, a step The method of claim 41 comprising:

44. The step of processing the internal data object to generate the first set of application data comprises removing inaccurate data items from the internal data object The method of claim 41 comprising:

45. Before receiving the first natural language input, receiving a third natural language input comprising instructions for generating the AI model; accessing the metadata graph to determine a node corresponding to the third natural language input; To generate a set of training data, processing an internal data object associated with the node corresponding to the third natural language input; To generate a trained AI model, training the AI model using the set of training data, where applying the AI model to the first set of application data includes applying the trained AI model to the first set of application data, steps The method according to claim 41, further comprising. **Claim 46** The step of training the AI model includes accessing values of one or more model measurement criteria associated with each of a plurality of model types; selecting, by the computer system, a model type of the trained AI model from among the plurality of model types based on the accessed values of the one or more model measurement criteria; and training the selected model type. The method according to claim 45, comprising. **Claim 47** The step of training the AI model includes receiving a user selection for a model type of the AI model and training the selected model type. The method according to claim 45, comprising. **Claim 48** generating, by the computer system, a chat interface for display to the user further comprising wherein the first natural language input is received via the chat interface, and sending the representation of the one or more outputs for display to the user includes displaying the one or more outputs via the chat interface. The method according to claim 41, comprising. **Claim 49** The step of generating a second set of application data based on the second natural language input includes adding a data item to the first set of application data, removing a data item from the first set of application data, or applying a data modification operator to a value of a data item in the first set of application data to modify the value. The method according to claim 41, comprising. **Claim 50** one or more processors, One or more non-transitory computer-readable storage media storing executable instructions, wherein when the instructions are executed by the one or more processors, Receiving a first natural language input from a user including a set of clauses, and instructions for analyzing data associated with the set of clauses using an artificial intelligence (AI) model; Accessing a metadata graph to determine nodes corresponding to the set of clauses in response to the first natural language input; The metadata graph comprising: (i) a set of nodes comprising (a) metadata indicating internal data objects stored in a data silo, and (b) a location identifier of the data silo; and (ii) edges indicating data lineages of the set of nodes; Accessing; Processing one or more internal data objects indicated by the determined nodes to generate a first set of application data; Applying the AI model to the first set of application data to generate one or more first outputs; Transmitting a representation of the one or more first outputs for display to the user; Receiving a second natural language input from the user, the second natural language input including instructions for modifying the first set of application data; Generating a second set of application data based on the received second natural language input; and Applying the AI model to the second set of application data to generate one or more second outputs Causing the system to perform, one or more non-transitory computer-readable storage media and A system comprising.

51. The system of claim 50, wherein processing the internal data object to generate the first set of application data includes removing personally identifiable or private information from the internal data object.

52. Processing the internal data object to generate the first set of application data comprises Applying a set of data modification operators to the internal data object to generate a set of modified data based on the internal data object, wherein the first set of application data includes one or more modified data items from the set of modified data, Applying The system according to claim 50, comprising.

53. Processing the internal data object to generate the first set of application data, Removing inaccurate data items from the internal data object The system according to claim 50, comprising.

54. When the instructions are executed by the one or more processors, before receiving the first natural language input, Receiving a third natural language input including instructions for generating the AI model, Accessing the metadata graph to determine a node corresponding to the third natural language input, Processing an internal data object associated with the node corresponding to the third natural language input to generate a set of training data, Training the AI model using the set of training data to generate a trained AI model, Applying the AI model to the first set of application data, including applying the trained AI model to the first set of application data, Training The system according to claim 50, further causing the system to perform.

55. When the instructions are executed by the one or more processors, Generating a chat interface for display to the user Further causing the system to, wherein the first natural language input is received via the chat interface, Sending the representation of the one or more outputs for display to the user, including displaying the one or more outputs via the chat interface. The system according to claim 50.

56. Generating a second set of application data based on the second natural language input, Adding data items to the first set of application data, Removing data items from the first set of application data, or Applying a data modification operator to the value to modify the value of a data item in the first set of application data The system according to claim 50, comprising: **Claim 57** A non-transitory computer-readable storage medium storing executable instructions, which when executed by one or more processors of a system, Receiving a first natural language input from a user comprising a set of clauses, and instructions for analyzing data associated with the set of clauses using an artificial intelligence (AI) model; In response to the first natural language input, accessing a metadata graph to determine nodes corresponding to the set of clauses, The metadata graph comprising: (i) a set of nodes comprising (a) metadata indicating internal data objects stored in a data silo, and (b) a location identifier of the data silo; and (ii) edges indicating data lineages of the set of nodes; Accessing; Processing one or more internal data objects indicated by the determined nodes to generate a first set of application data; Applying the AI model to the first set of application data to generate one or more first outputs; Transmitting a representation of the one or more first outputs for display to the user; Receiving a second natural language input from the user, the second natural language input comprising instructions for modifying the first set of application data; Generating a second set of application data based on the received second natural language input; Applying the AI model to the second set of application data to generate one or more second outputs Causing the system to perform. A non-transitory computer-readable storage medium. **Claim 58** Processing the internal data object to generate the first set of application data is Applying a set of data modification operators to the internal data object to generate a set of modified data based on the internal data object wherein the first set of application data includes one or more modified data items from the set of modified data Applying A non-transitory computer-readable storage medium according to claim 57, comprising: **Claim 59** When the instructions are executed by the one or more processors, before receiving the first natural language input, Receiving a third natural language input including instructions for generating the AI model; Accessing the metadata graph to determine a node corresponding to the third natural language input; Processing an internal data object associated with the node corresponding to the third natural language input to generate a set of training data; Training the AI model using the set of training data to generate a trained AI model, wherein applying the AI model to the first set of application data includes applying the trained AI model to the first set of application data; Training A non-transitory computer-readable storage medium according to claim 57, further causing the system to perform: **Claim 60** When the instructions are executed by the one or more processors, Generating a chat interface for display to the user; further causing the system to perform: wherein the first natural language input is received via the chat interface; A non-transitory computer-readable storage medium according to claim 57, wherein sending the representation of the one or more outputs for display to the user includes displaying the one or more outputs via the chat interface. US03 Pending **Claim 61** Receiving, by an entity, a first request for deploying a first artificial intelligence (AI) model to make the first AI model available for use in a production environment for processing input data and generating corresponding outputs; Selecting, based on a model deployment engine, a first model deployment location for the first AI model; The model deployment engine is configured to select the first model deployment location from a set of one or more cloud provider environments or an on-premises environment operated by the entity. Step; Generating a script for deploying the first AI model to the first model deployment location; Monitoring the operating parameters associated with the deployment of the first AI model at the selected model deployment location when the first AI model processes the input data and generates the corresponding output; Updating the model deployment engine based on the values of the monitored operating parameters; Selecting a second model deployment location for the second AI model based on the updated model deployment engine in response to a second request for deploying the second AI model A computer-executed method including.

62. The model deployment engine includes a decision tree, a knowledge graph, or a rule engine, and the step of selecting the first model deployment location based on the model deployment engine includes The computing cost for deploying the first AI model at the selected model deployment location, The computing capacity of the selected model deployment location, The response time from the selected model deployment location, or A measured value of the accuracy of the first AI model when deployed at the selected model deployment location The computer-executed method according to claim 61, including the step of selecting the first model deployment location based on one or more of the above.

63. The step of selecting the first model deployment location based on the model deployment engine includes the step of selecting the first model deployment location based on a privacy policy associated with at least one of the first AI model, the input data, or the corresponding output. The computer-executed method according to claim 61.

64. The operating parameters are The computing cost used by the first AI model at the first model deployment location, The response time from the first AI model when deployed at the first model deployment location, The computing capacity available in an environment including the first model deployment location, The privacy policy of the environment including the first model deployment location, or The downtime or service interruption measurement of the environment including the first model deployment location The computer-executed method according to claim 61, including one or more of the above.

65. The model deployment engine includes a trained decision model, and the step of updating the model deployment engine includes retraining the trained decision model based on the difference between the value of the monitored operation parameter and the value of the set of operation parameters for which the model deployment engine was trained. The computer-executed method according to claim 61.

66. The second AI model is a second instance of the first AI model, and the second model deployment location for the second AI model is a location different from the first model deployment location for the first AI model. The computer-executed method according to claim 61.

67. Based on the updated model deployment engine, selecting a third model deployment location for the first AI model that is different from the first model deployment location; and Generating a script for deploying the first AI model to the third model deployment location The computer-executed method according to claim 61, further including the above.

68. Based on the model deployment engine and the operation parameter, selecting a third model deployment location for the first AI model that is different from the first model deployment location; and Generating a script for deploying the first AI model to the third model deployment location The computer-executed method according to claim 61, further including the above.

69. The operation parameter includes the computing cost associated with the deployment of the first AI model at the first model deployment location, and the step of selecting the third model deployment location includes When the computing cost associated with the deployment of the first AI model at the first model deployment location is greater than the predicted computing cost associated with deploying the first AI model at the third model deployment location, determining to move the first AI model to the third model deployment location The computer-executed method according to claim 68, comprising:

70. One or more processors; One or more non-transitory computer-readable storage media storing executable instructions, wherein when the instructions are executed by the one or more processors, Receiving, by an entity, a first request for deploying a first artificial intelligence (AI) model for use in a production environment for processing input data and generating a corresponding output so that the first AI model can be utilized; Selecting, based on a model deployment engine, a first model deployment location for the first AI model, wherein the model deployment engine is configured to select the first model deployment location from a set of one or more cloud provider environments or an on-premises environment operated by the entity; selecting; Generating a script for deploying the first AI model at the first model deployment location; Monitoring, at the selected model deployment location, the operational parameters associated with the deployment of the first AI model when the first AI model processes the input data and generates the corresponding output; Updating the model deployment engine based on the values of the monitored operational parameters, and In response to a second request for deploying a second AI model, selecting a second model deployment location for the second AI model based on the updated model deployment engine One or more non-transitory computer-readable storage media causing the system to perform; A system comprising:

71. The model deployment engine comprises a decision tree, a knowledge graph, or a rule engine, and selecting the first model deployment location based on the model deployment engine is The computational cost for deploying the first AI model at the selected model deployment location, The computational capacity of the selected model deployment location, The response time from the selected model deployment location, or A measured value of the accuracy of the first AI model when deployed at the selected model deployment location The system according to claim 70, comprising selecting the first model deployment location based on one or more of the foregoing. **Claim 72** The system according to claim 70, wherein selecting the first model deployment location based on the model deployment engine includes selecting the first model deployment location based on a privacy policy associated with at least one of the first AI model, the input data, or the corresponding output. **Claim 73** The operating parameters are The computational cost used by the first AI model at the first model deployment location, The response time from the first AI model when deployed at the first model deployment location, The computational capacity available in the environment including the first model deployment location, The privacy policy of the environment including the first model deployment location, or A measured value of the downtime or service interruption of the environment including the first model deployment location The system according to claim 70, including one or more of the foregoing. **Claim 74** The model deployment engine comprises a trained decision model, and updating the model deployment engine includes retraining the trained decision model based on the difference between the value of the monitored operating parameter and the value of the set of operating parameters for which the model deployment engine was trained. The system according to claim 70. **Claim 75** The second AI model is a second instance of the first AI model, and the second model deployment location for the second AI model is a location different from the first model deployment location for the first AI model. The system according to claim 70. **Claim 76** When the instructions are executed by the one or more processors, Based on the updated model deployment engine, selecting a third model deployment location for the first AI model that is different from the first model deployment location, generating a script for deploying the first AI model to the third model deployment location The system according to claim 70, further causing the system to perform.

77. When the instructions are executed by the one or more processors, Based on the model deployment engine and the operation parameters, selecting a third model deployment location for the first AI model that is different from the first model deployment location, generating a script for deploying the first AI model to the third model deployment location The system according to claim 70, further causing the system to perform.

78. The operation parameters include the calculation cost associated with the deployment of the first AI model at the first model deployment location, and selecting the third model deployment location is determining to move the first AI model to the third model deployment location when the calculation cost associated with the deployment of the first AI model at the first model deployment location is greater than the predicted calculation cost associated with deploying the first AI model at the third model deployment location The system according to claim 77, comprising.

79. A non-transitory computer-readable storage medium storing executable instructions, wherein when the instructions are executed by one or more processors of a system, receiving, for the first AI model used by an entity, a first request for deploying the first artificial intelligence (AI) model to make the first AI model available for use in a production environment for processing input data and generating corresponding output; selecting, based on a model deployment engine, a first model deployment location for the first AI model, wherein the model deployment engine is configured to select the first model deployment location from a set of one or more cloud provider environments or an on-premises environment operated by the entity selecting, generating a script for deploying the first AI model to the first model deployment location, monitoring operation parameters associated with the deployment of the first AI model at the selected model deployment location when the first AI model processes the input data and generates the corresponding output, updating the model deployment engine based on values of the monitored operation parameters, selecting a second model deployment location for the second AI model based on the updated model deployment engine in response to a second request for deploying the second AI model to cause the system to perform, a non-transitory computer-readable storage medium.

80. wherein the model deployment engine comprises a decision tree, a knowledge graph, or a rule engine, and selecting the first model deployment location based on the model deployment engine is a computational cost for deploying the first AI model at the selected model deployment location, a computational capacity of the selected model deployment location, a response time from the selected model deployment location, or a measured value of the accuracy of the first AI model when deployed at the selected model deployment location The non-transitory computer-readable storage medium according to claim 79, comprising selecting the first model deployment location based on one or more of. US04 pending claims

81. A system for reducing the usage of computing resources when accessing silo storage data spanning different locations via an integrated metadata graph, comprising at least one hardware processor, when executed by the at least one hardware processor, identifying a set of keywords associated with a request to access a set of data objects, performing natural language processing on the set of keywords to determine a set of semantically similar phrases corresponding to each keyword in the set of keywords To determine the nodes corresponding to the set of semantically similar sentences, access the metadata graph, where the metadata graph comprises: (i) a set of nodes indicating (a) metadata of internal data objects stored in data silos, and (b) location identifiers of the data silos; and (ii) edges indicating data lineages between a first node and a second node among the set of nodes, and the metadata graph is generated using a metadata data structure based on file-level and container-level metadata identifiers, the act of accessing To determine the data silo storing at least one data object among the set of data objects using the location identifier corresponding to the determined node, in order to obtain at least one data object among the set of data objects via the data silo, and To generate a visual representation of the at least one data object for display on a graphical user interface (GUI), where the visual representation of the at least one data object comprises lineage information of the at least one data object, the act of generating At least one non-transitory memory storing instructions causing the system to perform A system comprising Claim 82 The metadata graph comprises Retrieving from a second set of data silos: (i) a set of file-level metadata identifiers, and (ii) a set of container-level metadata identifiers, where each file-level metadata identifier in the set of file-level metadata identifiers indicates metadata of a given data object stored in a respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates metadata of the respective data silo among the second set of data silos, the act of retrieving Generating, respectively, a set of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier Generating the metadata data structure for mapping each semantically similar metadata identifier of the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier; Generating the metadata graph using the generated metadata data structure; The system according to claim 81, generated by: **Claim 83** Receiving, via a second GUI, a second user-specified query that instructs a request for generating an intended result; Providing the second user-specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising: (i) a second artificial intelligence model to be used to generate the intended result; and (ii) a second set of data objects to be used when training the second artificial intelligence model; In response to receiving a user selection indicating approval of the recommendation: (i) accessing a database to obtain the second artificial intelligence model; and (ii) using the metadata graph to obtain the second set of data objects; Training the second artificial intelligence model using the set of data objects; Applying the second artificial intelligence model to generate the intended result; The system according to claim 81, further comprising the instructions for performing: **Claim 84** Accessing a governance database to obtain a set of policies that indicate usage metrics corresponding to the second set of data objects; Determining, using the set of policies that indicate usage metrics corresponding to the second set of data objects, whether the second set of data objects is approved to be used to train the second artificial intelligence model; Determining, using a second set of policies that indicate usage metrics corresponding to artificial intelligence model predictions, whether the output of the second artificial intelligence model is approved to be provided to one or more computing systems; in response to (i) the second set of data objects being approved for use in training the second artificial intelligence model and (ii) the output of the second artificial intelligence model being approved for being provided to the one or more computing systems, applying the second artificial intelligence model to generate the intended result The system of claim 83, further comprising the instructions for doing so.

85. A method for reducing the usage of computing resources when accessing silo storage data spanning heterogeneous locations via an integrated metadata graph, comprising: identifying a set of keywords associated with a user-specified query for accessing a set of data objects; performing natural language processing on the user-specified query to determine a set of phrases corresponding to the user-specified query; accessing a metadata graph to determine nodes corresponding to the set of phrases, the metadata graph comprising: (i) a set of nodes comprising (a) metadata indicating internal data objects stored in a data silo and (b) a location identifier of the data silo, and (ii) edges indicating data lineages of the set of nodes, the metadata graph being generated using a metadata data structure based on file-level and container-level metadata identifiers; determining a data silo storing at least one data object of the set of data objects using the location identifier corresponding to the determined nodes to obtain at least one data object of the set of data objects via the data silo; generating a representation of the at least one data object and a method.

86. The metadata graph is A step of extracting, from a second set of data silos, (i) a set of file-level metadata identifiers, and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier in the set of file-level metadata identifiers indicates the metadata of a given data object stored within a respective data silo, and each container-level metadata identifier in the set of container-level metadata identifiers indicates the metadata of the respective data silo among the second set of data silos. A step of generating, respectively, a set of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier. A step of generating the metadata data structure for mapping each semantically similar metadata identifier in the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier. A step of generating the metadata graph using the generated metadata data structure. The method according to claim 85, generated by the above.

87. A step of receiving, via a second GUI, a second user-specified query that instructs a request for generating an intended result. A step of providing the second user-specified query to an artificial intelligence model to generate a recommendation, wherein the recommendation comprises (i) a second artificial intelligence model that will be used to generate the intended result, and (ii) a second set of data objects that will be used when training the second artificial intelligence model. In response to receiving a user selection indicating acceptance of the recommendation, (i) a step of accessing a database to obtain the second artificial intelligence model, and (ii) a step of obtaining the second set of data objects using the metadata graph. A step of training the second artificial intelligence model using the set of data objects. A step of applying the second artificial intelligence model to generate the intended result. The method according to claim 85, further comprising the above.

88. Accessing a governance database to obtain a set of policies that indicate a usage measure corresponding to the second set of data objects; Using the set of policies that indicate a usage measure corresponding to the second set of data objects to determine whether the second set of data objects is approved to be used to train the second artificial intelligence model; Using a second set of policies that indicate a usage measure corresponding to an artificial intelligence model prediction to determine whether the output of the second artificial intelligence model is approved to be provided to one or more computing systems; Applying the second artificial intelligence model to generate the intended result in response to (i) the second set of data objects being approved to be used to train the second artificial intelligence model and (ii) the output of the second artificial intelligence model being approved to be provided to the one or more computing systems; The method of claim 87, further comprising.

89. The step of determining the set of clauses corresponding to the user-specified query comprises: Parsing the user-specified query for a set of keywords, each keyword in the set of keywords being associated with the set of data objects; For each keyword in the set of keywords associated with the set of data objects, determining a set of semantically similar clauses corresponding to each respective keyword in the set of keywords; Using the set of semantically similar clauses corresponding to each keyword in the set of keywords to determine the set of clauses corresponding to the user-specified query. The method of claim 85, further comprising.

90. The step of determining a set of semantically similar clauses corresponding to each respective keyword in the set of keywords comprises: Accessing a database that indicates a mapping between a first keyword and a set of second keywords; In response to accessing the database, determining the set of semantically similar sentences corresponding to each keyword using each of the keywords The method according to claim 89, further comprising.

91. Accessing the metadata graph is Traversing each node of the set of nodes to identify a metadata identifier that matches at least one sentence of the set of sentences; and Determining the node corresponding to the set of sentences in response to determining that the metadata identifier matches at least one sentence of the set of sentences The method according to claim 85, further comprising.

92. Accessing the metadata graph is Traversing each node of the set of nodes to identify a metadata identifier that matches at least one sentence of the set of sentences; and Determining a first node corresponding to the set of sentences in response to determining that the metadata identifier matches at least one sentence of the set of sentences; and In response to determining that the first node corresponds to the set of sentences, performing a second traversal of the nodes of the set of nodes using an edge that indicates a first data lineage of the first node, wherein the first data lineage of the first node indicates a second node comprising information that is a source of information associated with the first node; Determining the second data silo storing the second data object of the set of data objects using the location identifier corresponding to the second node to obtain the second data object of the set of data objects via the second data silo; and Generating a second representation of the second data object The method according to claim 85, further comprising.

93. The method according to claim 85, wherein the representation of the at least one data object comprises lineage information of the at least one data object.

94. When executed by one or more processors To determine a set of phrases corresponding to the user-specified query, performing natural language processing on the user-specified query, To determine nodes corresponding to the set of phrases, accessing a metadata graph, wherein the metadata graph comprises (i) (a) metadata indicating internal data objects stored in data silos, and (b) location identifiers of the data silos, and (ii) edges indicating data lineages of the set of nodes, and the metadata graph is generated using a metadata data structure based on file-level and container-level metadata identifiers, accessing, To determine a data silo storing at least one data object of the set of data objects by using the location identifier corresponding to the determined node, for obtaining at least one data object of the set of data objects via the data silo, Generating a representation of the at least one data object One or more non-transitory computer-readable media storing instructions for causing operations including.

95. The metadata graph, Taking out from a second set of data silos (i) a set of file-level metadata identifiers, and (ii) a set of container-level metadata identifiers, wherein each file-level metadata identifier of the set of file-level metadata identifiers indicates metadata of a given data object stored in a respective data silo, and each container-level metadata identifier of the set of container-level metadata identifiers indicates metadata of the respective data silo of the second set of data silos, taking out, Generating respective sets of semantically similar metadata identifiers corresponding to each file-level and container-level metadata identifier, Generating the metadata data structure for mapping each semantically similar metadata identifier of the set of semantically similar metadata identifiers to a normalized file-level metadata identifier and a normalized container-level metadata identifier, Generating the metadata graph using the generated metadata data structure The medium according to claim 94, generated by **Claim 96** When the instructions are executed by the one or more processors Receiving, via a second GUI, a second user-specified query that instructs a request for generating an intended result Providing the second user-specified query to an artificial intelligence model to generate a recommendation, the recommendation comprising: (i) a second artificial intelligence model that is to be used to generate the intended result, and (ii) a second set of data objects that are to be used when training the second artificial intelligence model In response to receiving a user selection indicating approval of the recommendation, (i) accessing a database to obtain the second artificial intelligence model, and (ii) obtaining the second set of data objects using the metadata graph Training the second artificial intelligence model using the set of data objects Applying the second artificial intelligence model to generate the intended result The medium according to claim 94, further causing operations to be performed including **Claim 97** When the instructions are executed by the one or more processors Accessing a governance database to obtain a set of policies that indicate usage metrics corresponding to the second set of data objects Determining, using the set of policies that indicate usage metrics corresponding to the second set of data objects, whether the second set of data objects is approved to be used to train the second artificial intelligence model Determining, using a second set of policies that indicate usage metrics corresponding to artificial intelligence model predictions, whether the output of the second artificial intelligence model is approved to be provided to one or more computing systems performing an operation further including applying the second artificial intelligence model to generate the intended result in response to (i) approval that the second set of data objects is used to train the second artificial intelligence model and (ii) approval that the output of the second artificial intelligence model is provided to the one or more computing systems The medium according to claim 96, further causing an operation including the above.

98. determining the set of clauses corresponding to the user-specified query parsing the user-specified query for a set of keywords, wherein each keyword in the set of keywords is associated with the set of data objects for each keyword in the set of keywords associated with the set of data objects, determining a set of semantically similar clauses corresponding to each of the keywords in the set of keywords determining the set of clauses corresponding to the user-specified query using the set of semantically similar clauses corresponding to each keyword in the set of keywords The medium according to claim 94, further including the above.

99. determining a semantically similar clause corresponding to each of the keywords in the set of keywords accessing a database that indicates a mapping between a first keyword and a set of second keywords in response to accessing the database, determining the set of semantically similar clauses corresponding to each of the keywords using each of the keywords The medium according to claim 98, further including the above.

100. accessing the metadata graph traversing each node in the set of nodes to identify a metadata identifier that matches at least one clause in the set of clauses in response to determining that the metadata identifier matches at least one clause in the set of clauses, determining the nodes corresponding to the set of clauses The medium according to claim 94, further including the above.

Citation Information

Patent Citations

  • Prescribed navigation using topology metadata and navigation path

    JP2006190261A

  • Graphic representation of data relationships

    JP2011517352A

  • Generating, accessing, and displaying lineage metadata

    JP2022033825A

  • Systems, methods, and apparatuses for executing a graph query against a graph representing a plurality of data stores

    US20200226156A1

  • System and method for querying multiple data sources

    US20220292092A1

Cited By

  • Programs, information processing devices, methods, and systems

    JP7874919B1

  • Apparatus and method of personalized portfolio financial analysis using rag model

    KR102987683B1