Artificial intelligence assisted data management for diversified source systems
By training AI agent 158 to interact with diverse data source systems, the problem of time-consuming and complex data management tasks in existing technologies has been solved, achieving efficient and secure cross-system data management and automated operations, thus improving the efficiency and security of data analysis.
Patent Information
- Application Number
- CN202510557334.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-02-25
- Filing Date
- 2025-04-29
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies are time-consuming and complex in configuring and managing data tasks when interacting with diverse data source systems, making it difficult to automate operations across systems efficiently and securely.
By training an AI agent to automate task management for diverse data source systems, using tool 159 for machine-to-machine communication, and combining role-based access control and tool configuration layer 155, the interactive capabilities of the AI agent 158 are expanded to generate execution plans and verify user permissions.
It improves the efficiency and security of data management, enables seamless data access and automated operations across multiple systems, reduces the need for manual integration, and enhances the response time and insight capabilities of data analysis.
Smart Images

Figure CN120873032A_ABST
Abstract
Description
[0001] Related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 640,684, filed April 30, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to data management in computing systems. Background Technology
[0004] Data is typically queried to retrieve specific information or datasets from storage systems, enabling data analysis, data recovery, data mining, forensic analysis, and regulatory compliance. Data accessible to a data management platform can be stored across multiple cloud and on-premises environments in various formats and is manageable using a variety of tools. These tools can include third-party applications and orchestration tools, cloud services, etc. Interfaces for these tools can be integrated into the data management platform to extend and customize it to meet specific requirements. Summary of the Invention
[0005] Generally speaking, this describes techniques for artificial intelligence (AI)-assisted data management used in diverse data source systems. For example, an AI agent is trained to interact with a user and generate responses to queries or prompts (hereinafter, "queries"). The AI agent utilizes tools to generate responses to accomplish tasks involving multiple diverse data source systems in order to satisfy the query. Queries may include requests for data, requests for operational insights or guidance, requests for configuring one or more data source systems, or other queries. Each tool extends the AI agent's ability to intelligently access data from different data source systems, for example, by implementing additional protocols and defining requests that the AI agent is trained to utilize to act autonomously (or semi-autonomously) on behalf of the user to satisfy the user query.
[0006] The data management platform configures tools to use users' role-based access privileges. Therefore, the AI agent utilizing the tool inherits the user's privileges and is thus able to interact with data sources accessed through the tool as if it were a user directly interacting with the data source.
[0007] The technology described herein offers one or more technical advantages. Existing solutions for interacting with data source systems and applications are time-consuming to configure and use for new tasks involving these systems. The technology disclosed herein extends the ability of AI agents to interact with such data source systems and applications, allowing the AI agent to not only enhance the user's understanding and capabilities regarding data and applications distributed across multiple systems, but also to enable the user to leverage scalable AI agents to accomplish new tasks, generate new operational insights, and otherwise manage data accessible from multiple diverse systems more efficiently and intelligently. In other words, the AI agent can seamlessly coordinate actions across multiple data source systems using different tools, eliminating the need for manual integration between different systems. The AI agent can execute these actions concurrently, thereby improving response times to queries. Furthermore, user permissions can be verified before actions are executed, ensuring security and compliance while automating complex cross-system workflows. The technology can thereby improve the field of data management and data analysis by enhancing the capabilities and performance of specific machines (i.e., the computing systems implementing the AI agent).
[0008] In one example, a method includes: receiving a query by a computing system for processing by an artificial intelligence (AI) agent, the query including a request for execution of a task associated with a first data source system and a second data source system; utilizing the AI agent to generate an execution plan for the task based on the query to satisfy the query, wherein the execution plan includes actions to be performed regarding the first and second data source systems, and wherein a user associated with the first and second data source systems has permissions for each of the actions; the AI agent invoking a first tool to perform a first action regarding the first data source system in the action; and the AI agent invoking a second tool to perform a second action regarding the second data source system in the action.
[0009] In the example, a computing system includes: one or more storage devices; and a processing component capable of accessing the one or more storage devices and configured to: receive a query for processing by an artificial intelligence (AI) agent, the query including a request for execution of a task associated with a first data source system and a second data source system; utilize the AI agent to generate an execution plan for the task based on the query to satisfy the query, wherein the execution plan includes actions to be performed regarding the first data source system and the second data source system, and wherein a user associated with the first data source system and the second data source system has permissions for each of the actions; have the AI agent invoke a first tool to execute a first action regarding the first data source system in the action; and have the AI agent invoke a second tool to execute a second action regarding the second data source system in the action.
[0010] In the example, one or more computer-readable media include instructions that, when executed by a processing component, cause the processing component to: receive a query for processing by an artificial intelligence (AI) agent, the query including a request for the execution of a task associated with a first data source system and a second data source system; utilize the AI agent to generate an execution plan for the task based on the query to satisfy the query, wherein the execution plan includes actions to be performed regarding the first and second data source systems, and wherein a user associated with the first and second data source systems has permissions for each of the actions; invoke a first tool by the AI agent to perform a first action regarding the first data source system in the action; and invoke a second tool by the AI agent to perform a second action regarding the second data source system in the action.
[0011] Details of one or more embodiments of the invention are set forth in the accompanying drawings and the following description. Other features, objects, and advantages of the invention will become apparent from the description and drawings, and from the claims. Attached Figure Description
[0012] Figure 1 This is a block diagram illustrating an example system for data management according to one or more aspects of this disclosure.
[0013] Figure 2 This is a block diagram illustrating an example data management platform based on the technology disclosed herein.
[0014] Figure 3 This is a block diagram illustrating an example of a computing system implementing a data management platform according to the technology disclosed herein.
[0015] Figure 4 This is a block diagram illustrating the workflow of actions performed by an AI agent using tools according to the technology of this disclosure.
[0016] Figure 5 This is a flowchart illustrating an example operating mode of a computing system according to one or more technologies disclosed herein.
[0017] Throughout the text and accompanying figures, similar figure labels indicate similar elements. Detailed Implementation
[0018] Figure 1 This is a block diagram illustrating an example system for data management according to one or more examples of this disclosure. Figure 1In the example, system 100 includes application system 102. Application system 102 represents a collection of hardware devices, software components, and / or data repositories that can be used to implement one or more applications or services provided via network 113 to one or more mobile devices 108 and one or more client devices 109. Application system 102 may include one or more physical or virtual computing devices that execute workloads 175 for the applications or services. Workloads 175 may include one or more virtual machines, containers, Kubernetes pods each including one or more containers, bare-metal processes, and / or other types of workloads.
[0019] exist Figure 1 In the example, application system 102 includes application servers 170A to 170M (collectively referred to as "application server 170") connected via a network to a database server 173 that implements the database. Other examples of application system 102 may include one or more load balancers, web servers, network devices (such as switches or gateways), or other devices for implementing one or more applications or services and delivering one or more applications or services to mobile device 108 and client device 109. Application system 102 may include one or more file servers. One or more file servers may implement the primary file system of application system 102. (In such instances, file system 153 may be a secondary file system that provides backup, archiving, and / or other services to the primary file system. The file system mentioned herein may include a primary file system or a secondary file system, such as the primary file system of application system 102 or file system 153 operating as a primary or secondary file system.) Application system 102 may be located on-premises and / or in one or more data centers, each of which is part of a public cloud, private cloud, or hybrid cloud. The application or service may be a distributed application. Applications or services may support enterprise software, financial software, office or other productivity software, data analytics software, end-user relationship management, web services, educational software, database software, multimedia software, information technology, healthcare software, or other types of applications or services. Applications or services may be provided as a service (-aaS) as Software as a Service, Platform as a Service, Infrastructure as a Service, Data Storage as a Service (dSaaS), or other types of services.
[0020] In some examples, application system 102 may represent an enterprise system, which includes one or more workstations in the form of desktop computers, laptops, mobile devices, enterprise servers, network devices, and other hardware to support enterprise applications. Enterprise applications may include enterprise software, financial software, office or other productivity software, data analytics software, end-user relationship management, web services, educational software, database software, multimedia software, information technology, healthcare software, or other types of applications. Enterprise applications may include applications that generate queries to AI agent 158, which responds to the queries. AI agent 158 may respond to queries using services available at data source systems 160A to 160K (collectively, “data source system 160”) or using other data stored and available from data source system 160 based on backup data stored at storage system 105 of data source system 160A. Enterprise applications may be delivered as services from external cloud service providers or other providers, executed natively on application system 102, or both. In some examples, application system 102 may be considered a data source system.
[0021] exist Figure 1 In the example, system 100 includes a data source system 160A that uses storage system 105 to provide file system 153 and backup functionality for application system 102. In some cases, data source system 160A may use a separate secondary storage system (not shown) to store backup data. Data source system 160A implements a distributed file system 153 and storage architecture to facilitate application system 102's access to file system data and to facilitate data transfer between storage system 105 and application system 102 via network 111. In the case of a distributed file system, data source system 160A enables devices of application system 102 to access file system data via network 111 using communication protocols as if such file system data were stored locally (e.g., on the hard drive of application system 102's device). Example communication protocols for accessing files and objects include Server Message Block (SMB), Network File System (NFS), or Amazon Simple Storage Service (S3). File system 153 may be a primary file system or a secondary file system of application system 102.
[0022] File system manager 152 represents a collection of hardware devices and software components that implement file system 153 for data source system 160A. Examples of file system functions provided by file system manager 152 include storage space management, which includes deduplication, file naming, directory management, metadata management, partitioning, and access control. File system manager 152 executes communication protocols to facilitate application system 102's access to files and other objects stored on storage system 105 via network 111.
[0023] Data source system 160A includes storage system 105 having one or more storage devices 180A to 180N (collectively, “Storage Device 180”). Storage device 180 may represent one or more physical or virtual computing and / or storage devices that include storage media or are otherwise accessible to storage media. Such storage media may include one or more of the forms of flash drives, solid-state drives (SSDs), hard disk drives (HDDs), electrically programmable memory (EPROM) or electrically erasable programmable memory (EEPROM) and / or other types of storage media used to support data source system 160A. Different storage devices in storage device 180 may have different mixtures of storage media. Each storage device in storage device 180 may include system memory. Each storage device in storage device 180 may be a storage server, a network attached storage (NAS) device, or disk storage that may represent a computing device. Storage system 105 may include a redundant array of independent disks (RAID) system, storage-as-a-service (STaaS), network attached storage (NAS), and / or storage area network (SAN). In some examples, one or more storage devices in storage device 180 are both computing devices and storage devices that execute software for data source system 160A, such as file system manager 152 and data protection manager 154 in the example of system 100, and store objects and metadata for data source system 160A to the storage medium. In some examples, a separate computing device (not shown) executes software for data source system 160A, such as file system manager 152 and data protection manager 154 in the example of system 100. Each storage device in storage device 180 may be considered and referred to as a "storage node" or simply a "node". In some examples, storage device 180 may represent a virtual machine, cloud virtual machine, physical rack server, or computing model installed in a converged platform that runs on a supported hypervisor.
[0024] In some examples, data source system 160A runs natively on a physical system, virtually, or in the cloud. For example, data source system 160A can be deployed to a physical cluster, a virtual cluster, or a cloud-based cluster that runs in a private cloud, on-premises deployment, hybrid cloud, or public cloud deployed by a cloud service provider. In some examples of system 100, multiple instances of data source system 160A can be deployed, and file system 153 can be replicated among the various instances. In some cases, data source system 160A represents a compute cluster of a single management domain. The number of storage devices 180 can be scaled to meet performance requirements.
[0025] Data source system 160A can implement and provide multiple storage domains to one or more tenants, or isolate workloads 175 that require different data policies. A storage domain is a data policy domain that determines the policies for deduplication, compression, encryption, tiering, and other operations performed on objects stored using that domain. In this way, data source system 160A provides users with the flexibility to select a global data policy or a workload-specific data policy. Data source system 160A can support partitioning.
[0026] A view is a protocol export residing within a storage domain. Views inherit data policies from their storage domain, but additional data policies can be specified for views. Views can be exported via SMB, NFS, S3, and / or another communication protocol. Policies defining data processing and storage by the data source system 160A can be assigned at the view level. Protection policies can specify backup frequency and retention policies.
[0027] Each of Network 113 and Network 111 may be the Internet, or may include or represent any public or private communications network or other network. For example, each of Network 113 and Network 111 may be a cellular, Wi-Fi®, ZigBee®, Bluetooth®, Near Field Communication (NFC), satellite, enterprise, service provider, local area network, and / or other type of network enabling data transfer between computing systems, servers, computing devices, and / or storage devices. One or more of such devices may use any suitable communication technology to transmit and receive data, commands, control signals, and / or other information across Network 113 or Network 111. Each of Network 113 or Network 111 may include one or more network hubs, network switches, network routers, satellite dish panels, or any other network equipment. Such network devices or components may be operatively coupled to each other, thereby providing information exchange between computers, devices, or other components (e.g., between one or more client devices or systems and one or more computer / server / storage devices or systems). Figure 1Each of the devices or systems shown may be operatively coupled to network 113 and / or network 111 using one or more network links. The links coupling such devices or systems to network 113 and / or network 111 may be Ethernet, Asynchronous Transfer Mode (ATM), or other types of network connections, and such connections may be wireless and / or wired connections. Figure 1 One or more of the devices or systems shown or otherwise located on network 113 and / or network 111 may be in a remote location relative to one or more other devices or systems shown.
[0028] Application system 102 uses file system 153 provided by data source system 160A to generate objects and other data. File system manager 152 creates, manages, and stores these objects and other data in storage system 105. For this purpose, application system 102 may be alternatively referred to as the "source system," and file system 153 used by application system 102 may be alternatively referred to as the "source file system." Application system 102 may communicate directly with storage system 105 via network 111 to transfer objects for some purposes, and may also communicate indirectly with file system manager 152 via network 111 to obtain objects or metadata from storage system 105 for some purposes. File system manager 152 generates metadata and stores it in storage system 105. The collection of data stored in storage system 105 and used to implement file system 153 is referred to herein as file system data. File system data may include the aforementioned metadata and objects. Metadata may include file system objects, tables, trees, or other data structures; metadata generated to support deduplication; or metadata used to support snapshots. The stored objects can include files, virtual machines, databases, applications, pods, containers, any workload in workload 175, system images, directory information, or other types of objects used by application system 102. These can also be referred to as "backup objects." Objects of different types and objects of the same type can be deduplicated relative to each other.
[0029] Data source system 160A includes a data protection manager 154 that provides data protection operations for file system data of file system 153. In the example of system 100, data protection manager 154 backs up file system data to backup 142 stored in storage system 105. In some examples, a separate storage system (not shown) may store backup 142. The separate storage system may be deployed and managed by a cloud storage provider and is referred to as a “cloud storage system.” In some examples, the separate storage system may reside in a data center with storage system 105, on-premises, or in a private, public, or hybrid cloud. When storage system 105 is the primary storage system, the separate storage system may be considered a “backup” or “secondary” storage system of storage system 105. The separate storage system may be referred to as an “external target” of backups 142A to 142K (collectively referred to as “backup 142”). Any data source system from data source systems 160B to 160K may be a separate secondary storage system of data source system 160A.
[0030] Because storage system 105 is typically more difficult to scale up or more expensive to scale up, data source system 160A can use a secondary storage system to support secondary data protection use cases such as backup, archiving, mirroring, disaster recovery, and / or replication. Generally, a file system backup is a copy of file system 153 used to support protection of file system 153 for rapid recovery (typically due to some form of data loss within file system 153), while a file system archive (“archive”) is a copy of file system 153 used to support long-term retention and review. A “copy” of file system 153 may only include such data as is needed to restore or view the state of file system 153 at the time of backup or archiving. While the techniques disclosed herein are described with respect to retrieving backup data stored in storage system 105 or a secondary storage system, the techniques can be applied with respect to any data stored as backup data in any storage system. For example, backup data may include archived data, copied data, mirrored data, or snapshots.
[0031] Data Protection Manager 154 can back up the file system data of file system 153 at any time according to a backup policy that specifies, for example, backup periodicity and timing (daily, weekly, etc.), which file system data will be backed up, storage location, access control, etc. The backup of file system data corresponds to the state of the file system data at the time of backup. Backup 142 therefore represents time-series data of file system 153, as each backup stores a representation of file system 153 at a specific time. Because file system 153 changes over time due to the creation of new objects, modification of existing objects, and deletion of objects, backup 142 will vary. Depending on the backup policy, a backup can include a complete backup of file system 153 data, or it can include a complete backup of less than the amount of file system 153 data. For example, a given backup in backup 142 can include all objects of file system 153 or one or more selected objects of file system 153. A given backup in backup 142 can be a full backup or an incremental backup.
[0032] Backup 142 can be used to generate views or snapshots. A current view typically corresponds to a (near) real-time backup state of the file system 153. A snapshot represents the backup state of the primary storage system 105 at a specific point in time. That is, each snapshot provides the state of the data on the file system 153, which can be restored to the primary storage system 105 as needed. Similarly, snapshots can be exposed to non-production workloads, or clones of snapshots can be created if non-production workloads need to write to a snapshot without interfering with the original snapshot.
[0033] Therefore, Data Protection Manager 154 can use any backup in Backup 142 to subsequently restore the file system (or portions thereof) to its state at the time the backup was created, or, for example, use the backup to create or present a new file system (or "view") based on the backup. Data Protection Manager 154 can perform deduplication on file system data included in subsequent backups against file system data included in one or more previous backups. For example, deduplication can be performed on a second object in a second backup of file system 153 against a first object included in a first, earlier backup of file system 153.
[0034] Backup manager 154 may apply deduplication as part of a write process that writes (i.e., stores) objects of file system 153 to backup 142 in storage system 105. Additional description of an example deduplication process can be found in U.S. Patent Application No. 18 / 183,659, filed March 14, 2023, entitled “Adaptive Deduplication of DataChunks,” which is incorporated herein by reference in its entirety. Users or applications associated with application system 102 may be able to access (e.g., read or write) backup data stored in a separate storage system via data source system 160A or via data management platform 150.
[0035] Data source system 160 contains a large amount of enterprise information, but backup 142 suffers from high access latency due to being stored on slower storage media. Furthermore, in modern distributed architectures, the collection, verification, and utilization of data from workflows across an organization's data estate can be complex. Data source system 160 can operate in numerous locations, spanning private data centers, single or multiple clouds, SaaS applications hosted by other organizations, and edge locations such as repositories, Internet of Things (IoT) devices, and many other applications. Conventional data platforms can store quadrillions (or more) of data without categorizing, indexing, or tracking it. This is often referred to as "dark data," and it is typically unknown to the organization and is often unstructured and / or difficult to access. A major challenge with dark data is that it represents missed opportunities for organizations to gain insights, make informed decisions, and secure and protect their data.
[0036] Leveraging advanced backup systems, backup data can be readily analyzed and used by machine learning / artificial intelligence applications to deliver added value to users and businesses. U.S. Patent Application 18 / 618,695, filed March 27, 2024, entitled "DATA RETRIEVAL USINGEMBEDDINGS FOR DATA IN BACKUP SYSTEMS," describes retrieval-enhanced generation in which a data platform extracts data in text form from a data source, creates a semantic index on the data, and uses that index to generate insights into the data. This U.S. Patent Application is incorporated herein by reference in its entirety.
[0037] Data management platform 150 can provide centralized data management for data associated with users. For example, users can be organizations, tenants, individuals, enterprises, or their human agents. Data management platform 150 generates user interfaces for output and display via user devices (such as user devices 115 that access data management platform 150 via network 111).
[0038] Data associated with users and managed by the data management platform 150 can be distributed across multiple diverse data source systems 160. The data source systems 160 enable the data management platform 150 to access the data via a network 111. To access the data, the data management platform 150 utilizes tools 159A to 159N (collectively referred to as "tools 159"). Each data source system in the data source system 160 can represent a different type of data source, making the different data source systems diverse and accessible using different tools 159 and protocols, and providing data according to different data types and formats. For example, the data source systems 160 may each provide data in different formats, be dynamic or static, and otherwise differ in their accessibility to the data management platform 150, making them diverse.
[0039] Data source system 160 can be dynamic or static. Dynamic data source systems are those that store, provide, or otherwise obtain rapidly changing, accessible data. For example, these data source systems may include machine-generated data streams or real-time data feeds. Example dynamic data sources may include application programming interface (API) endpoints or software-as-a-service (SaaS) application endpoints (such as API 185 of cloud service 184), machine log data, message bus streams, relational databases (such as those shown in database system 182), key / value repositories, publish / subscribe service systems, etc. Static data source systems are those that store, provide, or otherwise obtain accessible data that changes or updates at a slower rate. Example static source systems include backup sources such as data source system 160A, vectorized context stores such as those described in U.S. Patent Application 18 / 618,695, archiving systems, etc.
[0040] Tool 159 may be a function invoked by AI agent 158 for accessing or managing data stored by or accessible from data source system 160. Tool 159 may be implemented as a standalone software application, which may execute directly on the data management platform 150 co-located with AI agent 158, or may execute on one or more external systems. One or more tools in tool 159 may be third-party applications specifically developed for accessing the corresponding data source system in data source system 160.
[0041] Each tool in tool 159 can be invoked by AI agent 158 for a northbound interface used for machine-to-machine communication. Each tool in tool 159 may be able to interact with a corresponding data source system in data source system 160 to execute requests received at the tool's northbound interface. In order to interact with data source system 160 to access or manage data or access data metadata, tool 159 may implement one or more communication protocols.
[0042] AI agent 158 receives input indicating a query, for example, from user device 115. The query may include text, for example. A query may be a request from data management platform 150 on behalf of the user of user device 115 to perform a task concerning data associated with the user and stored by any one or more data source systems 160. In some examples, the query may be independent of the user or user device 115, but still related to data associated with the user and stored by any one or more data source systems 160. Examples of such queries include internal system queries, queries related to scheduled or periodic tasks, or queries related to predictive analytics.
[0043] Fulfilling a task may require the data management platform 150 to perform multiple actions on behalf of the user of user device 115. For example, a query might be used to request: optimize backup 142, perform security operations, configure one or more data source systems 160, migrate data from data source system 160A to data source system 160B, generate analytical or operational insights for data stored in data source systems 160A and 160B, perform administrative tasks, etc. Queries can be natural language queries. (The security-related tasks mentioned herein should be understood as a form of data management.)
[0044] In some examples, requests for optimizing backups may involve at least one of the following: identifying overlapping backup windows across different systems to reduce redundancy and optimize resource usage; detecting and eliminating duplicate records across multiple data source systems before initiating backups; maintaining data alignment policies for data stored across multiple data source systems to ensure compliance; enabling incremental backups that track changes across data source systems rather than copying the entire dataset; and developing a unified recovery plan that includes data from multiple platforms to ensure seamless restoration.
[0045] In some examples, requests to perform security operations may involve requests for at least one of the following: aligning permissions and access rules across systems to ensure consistent security policies, aggregating log data from multiple data source systems to create a comprehensive audit trail for compliance purposes, and implementing consistent data masking policies across end-user databases and analytics platforms.
[0046] In some examples, requests for configuring one or more data source systems 160 may involve requests for at least the following: setting up a new relational database for user data storage; connecting multiple systems (such as CRM, ERP, and data warehouses); configuring the security of one or more data systems by setting up, for example, encryption protocols, implementing role-based access control (RBAC) and data masking, and configuring firewall rules and IP whitelists for secure connections; requests for aligning data structures between data source systems; performing backup and recovery configurations; and configuring performance optimizations to improve data flow efficiency.
[0047] In some cases, the requested task may be or include tasks that typically use a graphical user interface (GUI) or a command-line interface (CLI) of the data management platform 150. Figure 1 (The interface is not shown in the diagram) Available tasks. The data management platform 150 can implement the API according to the API specification, and can access and invoke the API to perform data management tasks.
[0048] AI agent 158 includes machine learning model 174 based on artificial intelligence or other machine learning techniques. For example, machine learning model 174 may include or use Word2Vec or Global Vectors for Word Representation (GloVe), recurrent neural networks (RNNs) such as Long Short-Term Memory (LSTM) or Gated Recurrent Unit (GRU) architectures, transformer models, convolutional neural networks (CNNs), graph neural networks (GNNs), autoencoders, gradient boosting machines (GBMs), deep neural networks (DNNs), or other artificial neural networks.
[0049] Machine learning model 174 can be a large language model (LLM). Machine learning model 174 can be trained on action-based outcomes to be more coordinated with actions that need to be performed in data management and security solutions. Such training may involve fine-tuning a third-party LLM to enable rapid execution of data management and security-related tasks. Machine learning model 174 can be implemented by a computing system separate from the computing system implementing data management platform 150. For example, machine learning model 174 can be provided as a service to data management platform 150 via a network. In such examples, the control plane of data management platform 150 manages communication with machine learning model 174 (see [link to documentation]). Figure 2 , Figure 4 ).
[0050] A machine learning system (separate from the data management platform 150 in some examples, but part of or executed by the data management platform 150 in others) can be used to train the machine learning model 174 of the AI agent 158. The machine learning system can be executed by a computing system. For example, the machine learning system can apply one or more of the following to train the machine learning model 174: nearest neighbor, Naive Bayes, decision tree, linear regression, support vector machine, neural network, k-means clustering, Q-learning, temporal difference, deep adversarial network, or other supervised learning algorithm, unsupervised learning algorithm, semi-supervised learning algorithm, or reinforcement learning algorithm.
[0051] AI agent 158 may also be referred to as an AI assistant, chatbot, chatbot, virtual assistant, or dialogue interface. AI agent 158 can perform query-based tasks by utilizing tool 159 to accomplish tasks involving one or more source systems 160 to satisfy queries. Performing tasks may include generating and outputting responses to the user. AI agent 158 can perform multiple tasks for multiple different queries. In some examples, AI agent 158 ingests API specifications for APIs implemented by data management platform 150 to perform operations typically available to the user via an interface. In such examples, AI agent 158 applying model 174 to a query can invoke the API of data management platform 150 to perform the requested task.
[0052] Each tool in tool 159 can extend the ability of AI agent 158 to intelligently access data in different source systems, for example by implementing additional protocols and developing requests that AI agent 158 and more specifically model 174 are trained to utilize to act autonomously (or semi-autonomously) on behalf of the user to satisfy user queries.
[0053] In some examples, the data management platform 150 configures tool 159 to use the user's role-based access privileges. Therefore, the AI agent 158 utilizing tool 159 inherits the user's privileges and is thus able to interact with the data source system 160 accessed through the tool as if it were a user interacting directly with the data source system. AI agent 158 can be extended to include additional tools 159.
[0054] Each tool in tool 159 can be configured for use by the AI agent 158 by configuring the tools to access the corresponding data source system in data source system 160 and by enabling the AI agent 158 to use the tools. Such configuration can be performed by the user and may involve the user specifying the specific tools in tool 159 that the AI agent 158 should use regarding user-related data, specifying how the AI agent 158 connects to tool 159, what types of calls tool 159 can make, and how tool 159 can authenticate and authorize data source system 160. (Regarding...) Figure 2 Provide a more detailed description of the tool 159 configuration.
[0055] Based on the query, the AI agent can select one or more tools from tools 159. The AI agent can use these tools to perform tasks associated with the query autonomously or semi-autonomously on behalf of the user. Privileged roles across the selected tools can be recorded and passed on, such that if the AI agent 158 is acting (semi-)autonomously on behalf of the user, the AI agent 158 acts as if it were the user with respect to the data source system 160 accessible by the selected tools.
[0056] As an example, consider a scenario where backup 142 includes a backup of data stored by data source system 160B. If a query request optimizes the backup of data stored by data source system 160B, AI agent 158 can choose to use tool 159A to interface with data source system 160A to obtain historical data describing backup 142, such as its scope, timing, applied strategy, and size. AI agent 158 can also choose to use tool 159B to interface with data source system 160B to obtain data describing database system 182. Based on the historical data describing previous backup 142 and the data describing database system 182, AI agent 158 can interact with data source system 160A via tool 159A to optimize backup settings for future backups of database system 182.
[0057] The user issuing the query has a role on data source system 160 that constrains the actions AI agent 158 can take with respect to the data source system, as well as the data that AI agent 158 can access and that the user can obtain in the response to the query. Continuing the example above, the user's role privileges with respect to data source system 160A determine whether and how AI agent 158 can be configured to optimize backup settings for future backups of database system 182.
[0058] If a user does not have sufficient privileges to perform an action regarding one of the data source systems in data source system 160, the AI agent 158 will not perform the action. This restriction facilitates secure access for the user.
[0059] The technology described herein can provide one or more technical advantages. Existing solutions for interacting with data source systems and applications are time-consuming to configure and use for new tasks involving data source systems. The technology disclosed herein can extend the ability of data management platform 150 and the AI agent 158 of the specific application model 174 to interact with data source system 160 and applications, so as to allow data management platform 150 not only to enhance the user's understanding and ability regarding data and applications distributed across multiple systems, but also to enable the user to leverage AI agent 158 to complete new tasks, generate new operational insights, and otherwise manage data accessible from multiple diverse systems more efficiently and intelligently.
[0060] Figure 2 This is a block diagram illustrating an example data management platform 150 according to various techniques of this disclosure. The data management platform 150 includes a control plane 220 implementing a user interface 153 and role-based access control (RBAC) 172, an AI agent 158, a tool configuration layer 155, tools 159, and a data access proxy layer 165. The control plane 220 configures the tool 159 in part based on RBAC 172, and the control plane 220 facilitates access to the data source system 160 via the data access proxy layer 165.
[0061] RBAC 172 assigns privileges or permissions to users of data management 150 based on user roles. Roles can represent different job functions or responsibilities within an organization. For example, roles can be "Manager," "Employee," "Administrator," etc. Permissions are the actions that users assigned to a role are allowed to perform within different data source systems 160. For example, permissions can include "read," "write," "delete," the ability to configure selected services or functions within a data source system, etc. RBAC 172 enhances security by ensuring that users can only access the resources and data required by their roles, thereby reducing the risk of unauthorized access and data breaches. RBAC 172 can improve compliance with regulatory requirements by providing a structured approach for access control and auditing.
[0062] User interface module 153 (“User Interface 153”) generates and outputs a user interface for display on a user device, through which data management platform 150 can, for example, receive user input, including prompts to AI agent 158, and output responses generated by AI agent 158.
[0063] Tool 159 includes functions that AI agent 158 can invoke (“call”). To complete a query-based task, AI agent 158 needs to access the appropriate tool 159 to complete the task, and AI agent 158 must be trained using descriptive information about the tool, and / or be able to access descriptive information about the tool so that AI agent 158 can select and use tool 159 to perform actions to complete the task. Tool 159 is a means for AI agent 158 to access other data sources, utilize protocols for such access, formulate calls to data source system 160, and filter the returned data.
[0064] To train AI agent 158 (and more specifically, model 174) to use tool 159, AI agent 158 can acquire and digest configuration information in the form of a specification for the tool, describing the actions the tool is capable of performing. Such specifications may include, for example, API specifications, user or management manuals, or websites. AI agent 158 can also be trained using training data generated from previous tasks performed by users of data management platform 150. Such training data may include records of user interactions via user interface 153, commands issued by control plane 220 to any data source system in data source system 160, descriptions of received data, or other data having a correlation between the desired task to be completed and the outcome of that task.
[0065] The tool configuration layer 155 configures the set of instructions (actions) that tool 159 can execute, and then configures the data structures for how those specific calls are executed. AI agent 158 interacts with tool 159 via the tool configuration layer 155, primarily using available calls to interact by extending the data source system 160. RBAC 172 can then be applied to those actions. For example, there may be an action for creating a backup job. Based on RBAC 172, more granular action privileges can be applied to the specific action of creating a backup job based on the user's query. For example, the user (and therefore AI agent 158) may be able to create a backup job involving a first set of objects in the data source system, but the user does not have permission to create a backup job involving a second set of objects in the data source system. Therefore, AI agent 158 must be trained or otherwise enabled to access the executable actions and the user's privileges / permissions in order to securely and successfully generate appropriate calls to execute the actions to complete the query-based task.
[0066] The tool configuration layer 155 enables individual users to specify which tools in tool 159 the AI agent 158 can use on their behalf, how the AI agent 158 connects to the tools, the types of calls the selected tool 159 can make to the data source system 160, and how the selected tool 159 can authenticate and authorize the data source system 160. Different configuration areas for each tool may include the following and corresponding configuration information:
[0067] 1) Tool Application / Data Source — Define the target application that the tool will interact with, such as any data source system in the data source system 160. Example applications may include workflow management applications, data management applications, SaaS applications, database management tools, data protection systems, and others.
[0068] 2) Tool Access Method – Specifies how the tool will access data from the data source system. Examples may include API, GraphQL, Open Database Connectivity (ODBC), and others.
[0069] 3) Tool Invocation Method – Specifies the scope of the invocation to the target application / data source system. Some examples of scopes are: GET, PUT, POST, DELETE, SELECT, INSERT, UPDATE, DELETE, DROP, etc. These scopes can be associated with access method protocols.
[0070] 4) Tool Authentication / Authorization – Specifies the methods and details used to authenticate the AI agent 158 against the target application / data source system. Example details may include user-provided API keys, user credentials, credential files, usernames and passwords, and authentication protocols such as OAuth, OpenID Connect, Secure Assertion Markup Language, Kerberos, or Lightweight Directory Access Protocol.
[0071] 5) Tool Description – A lengthy description of the tool’s purpose and the type, semantics, syntax, and / or description of the data it will return.
[0072] 6) Tool Name – The unique name of the tool that AI Agent 158 will reference.
[0073] Each of the above can be configured by the user through user interface 153. In some cases, the administrator / operator of data management platform 150 can use user interface 153 to define and configure tool 159 through tool configuration layer 155.
[0074] Data access proxy layer 165 enables tools 159 configured via tool configuration layer 155 to access data source system 160 accordingly. For a tool in tool 159 to connect to its configured data source system in data source system 160, the tool may need to authenticate with the data source system and check authorization regarding user access to data, which, according to various techniques of this disclosure, is obtained from RBAC 172 based on the user's role querying AI agent 158. Data access proxy layer 165 can constrain actions that can be performed on data source system 160 and the data visible from data source system 160. Because data management platform 150 is aware of any interaction with data source system 160 and has indications of user identity, data access proxy layer 165 can coordinate permissions and access levels among tools 159, data, and users accessing data source system 160. Data management platform 150 can receive indications of user identity through a login process.
[0075] To complete authentication / authorization, when registering a tool / data source system with the data management platform 150, the configuration state of the data access path can be stored in the data access broker layer 165. For example, if the service is a RESTful API endpoint, the user should utilize an access token to pre-configure the state or allow the user to pass their session access token.
[0076] Some of the tools in tool 159 can be configured to access the registered storage system (e.g., the data source system to which data protection is being applied by data management platform 150). In such cases, the authentication method / protocol used to access the data can be the same as the source registration, or it can be provided by the user via a pass-through method.
[0077] Data access proxy layer 165 can receive instructions from the user (i.e., the requester of the query) to obtain the user's role for the selected data source system 160 from control plane 220, obtain authentication details for the corresponding tool 159, and obtain or generate appropriate authentication mechanisms for the use of each tool in the corresponding tool. Because some of the tools in tool 159 can be stateless, data access proxy layer 165 can perform these operations each time a tool is invoked on behalf of the user.
[0078] Once the access and authentication methods have been delivered to one of the tools in tool 159, that tool can perform its given actions to advance the task according to the execution plan envisioned by AI agent 158.
[0079] exist Figure 2In the example, AI agent 158 is the top-level agent responsible for two main tasks. First, AI agent 158 interacts with the end user (or "requester"). AI agent 158 generates a response to a given input query from the user. Second, AI agent 158 selects one or more tools from the tools 159 available to the user to complete the given task. AI agent 158 invokes the selected one or more required tools to satisfy the input query.
[0080] In this way, the data management platform 150 provides a solution that enables the AI agent 158 to interact not only with the backup system and its data source, but also with other systems to which the backup system is connected or interacts with by utilizing information about those systems held by the data management platform 150, in order to act on behalf of the user who issued the query.
[0081] AI agent 158 can use tool 159 to interact with data source system 160, in order to interact with and optimize that interaction. This allows AI agent 158 to autonomously interact with external sources separate from the backup system, whether performing tasks to exchange data or configuration information, assisting in configuration optimization, or backing up and ensuring data security. Interacting with data source system 160 allows AI agent 158 to understand the configuration and status of those systems, which in turn enables AI agent 158 to autonomously or semi-autonomously interact with, optimize, and configure those environments on behalf of the user.
[0082] AI agent 158 can make multiple calls to data source system 160 to complete query-based tasks. In other words, AI agent 158 can perform information retrieval and action invocation in a multi-threaded manner. AI agent 158 may be able to perform inference based on real-time data feeds and simultaneously execute actions to complete tasks. AI agent 158 may be able to perform inference simultaneously across all configured tools 159, select multiple actions to complete tasks, and concurrently execute multiple actions on some data. In addition, data source system 160 can output data to data management platform 150 in various formats. In some examples, AI agent 158 does not need to maintain the pattern of the received data; AI agent 158 can convert the incoming data pattern into information useful to AI agent 158 in order to execute the next action.
[0083] Data received from any data source system in data source system 160 can be available to model 174 to drive RAG queries and other such AI / ML applications. RAG is a framework that combines a pre-trained sequence-to-sequence (seq2seq) model with a dense retrieval mechanism, allowing for the generation of more intelligent and context-relevant outputs. This allows users and applications to retrieve data safely and efficiently without compromising the integrity of the system or the data itself. RAG queries are also tailored to specific data types identified by machine learning analytics, ensuring that users and applications can quickly and easily access the desired information.
[0084] In the age of artificial intelligence, readily available, trained large language models (LLMs) have become powerful tools for generating human-like responses across a wide range of applications. However, most existing knowledge-based dialogue models rely on outdated material—individual documents relevant to the topic of the conversation—thus limiting LLMs' ability to generate diverse and knowledge-rich responses that can be more proprietary or domain-specific. To overcome this challenge, the concept of Relational Aggregators (RAGs) has been introduced. RAGs combine the strength of LLMs with the ability to retrieve information from multiple documents. RAGs not only enable LLMs to generate more knowledgeable, diverse, and relevant responses but also provide a more efficient way to fine-tune these models. By using RAGs to determine what to respond with and fine-tuning to guide how to respond, LLMs can deliver more engaging and informative conversational experiences.
[0085] AI agent 158, executing workflows to complete tasks, can use RAG to leverage data from any data source system in data source system 160 and incorporate (or enable) RAG-assisted large language models (LLMs) into 'AIReady'. Data security can be ensured via RBAC 172. By leveraging RAG on the enterprise's own datasets, users can teach language models (e.g., model 174) how to perform a given task without performing lengthy fine-tuning or initial training. Leveraging RAG provides up-to-date and relevant context for any query. The technology also enables responses at any point in time based on dynamic data in data source system 160.
[0086] Figure 3This is a block diagram illustrating examples of computing systems implementing data management platform 150 according to various techniques of this disclosure. Computing system 202 can be implemented as any suitable computing system, such as one or more server computers, workstations, mainframes, electrical appliances, cloud computing systems, and / or other computing systems capable of performing the operations and / or functions described in one or more aspects of this disclosure. In some examples, computing system 202 represents a cloud computing system, server farm, and / or server cluster (or a portion thereof) providing services to other devices or systems. In other examples, computing system 202 may represent or be implemented through one or more virtualized computing instances (e.g., virtual machines, containers) of a cloud computing system, server farm, data center, and / or server cluster.
[0087] exist Figure 3 In this example, computing system 202 may include one or more communication units 215, one or more input devices 217, one or more output devices 218, and one or more storage devices of storage system 105. Storage system 105 includes AI agent 158 and tools 159, which in this example are software modules. However, any one or more tools of tools 159 may execute on different systems. One or more of the devices, modules, storage areas, or other components of computing system 202 may be interconnected to enable inter-component communication (physical, communicative, and / or operational). In some examples, such connectivity may be provided via a communication channel (e.g., communication channel 212), which may represent one or more of a system bus, network connection, inter-process communication data structure, or any other method for transmitting data.
[0088] The computing system 202 includes processing components. Figure 2 In the examples, the processing component includes one or more processors 213 configured to implement functionality associated with computing system 202 or with one or more modules (including AI agent 158 and tool 159) shown and / or described herein and hereinafter, and to execute instructions associated therewith. The one or more processors 213 may be a processing circuitry, may be part of a processing circuitry, and / or may include a processing circuitry that performs operations according to one or more examples of this disclosure. Examples of processors 213 include microprocessors, application processors, display controllers, auxiliary processors, one or more sensor hubs, and any other hardware configured to function as a processor, processing unit, or processing device. Computing system 202 may use one or more processors 213 to perform operations according to one or more examples of this disclosure using software, hardware, firmware, or a mixture of hardware, software, and firmware residing in and / or executing at computing system 202.
[0089] One or more communication units 215 of computing system 202 can communicate with devices outside computing system 202 by transmitting and / or receiving data, and in some respects can operate as both an input device and an output device. In some examples, communication unit 215 can communicate with other devices via a network. In other examples, communication unit 215 can transmit and / or receive radio signals on a radio network (such as a cellular radio network). In other examples, communication unit 215 of computing system 202 can transmit and / or receive satellite signals on a satellite network. Examples of communication unit 215 include network interface cards (e.g., such as Ethernet cards), optical transceivers, radio frequency transceivers, GPS receivers, or any other type of device capable of transmitting and / or receiving information. Other examples of communication unit 215 may include devices capable of communicating via Bluetooth®, GPS, NFC, ZigBee® and cellular networks (e.g., 3G, 4G, 5G) and Wi-Fi® radios in mobile devices, as well as Universal Serial Bus (USB) controllers. Such communication may comply with, implement, or follow appropriate protocols, including Transmission Control Protocol / Internet Protocol (TCP / IP), Ethernet, Bluetooth®, NFC, or other technologies or protocols.
[0090] One or more input devices 217 may represent any input device of computing system 202 not otherwise described separately herein. Input devices 217 may generate, receive, and / or process input. For example, one or more input devices 217 may generate or receive input from a network, a user input device, or any other type of device for detecting input from a person or machine.
[0091] One or more output devices 218 may represent any output device of computing system 202 not otherwise described separately herein. Output devices 218 may generate, present, and / or process output. For example, one or more output devices 218 may generate, present, and / or process output of any form. Output devices 218 may include one or more USB interfaces, video and / or audio output interfaces, or any other type of device capable of generating tactile, audio, visual, video, electrical, or other outputs. Some devices may function as both input and output devices. For example, a communication device may transmit data to and receive data from other systems or devices via a network.
[0092] One or more storage devices of storage system 105 within computing system 202 may store information for processing during operation of computing system 202. The storage devices may store program instructions and / or data associated with one or more modules among those described according to one or more examples of this disclosure. One or more processors 213 and one or more storage devices may provide an operating environment or platform for such modules, which may be implemented as software, but in some examples may include any combination of hardware, firmware, and software. One or more processors 213 may execute instructions, and one or more storage devices of storage system 105 may store instructions and / or data of one or more modules. The combination of processors 213 and storage system 105 may retrieve, store, and / or execute instructions and / or data of one or more applications, modules, or software. Processors 213 and / or storage devices of storage system 105 may also be operatively coupled to one or more other software and / or hardware components, including but not limited to computing system 202 and / or one or more components of one or more devices or systems shown connected to computing system 202.
[0093] AI agent 158 can perform functions related to AI-assisted data management, as mentioned above. Figures 1 to 2 As described above, AI agent 158 can use tool 159 to interact with the data source accessed by the tool, just as if AI agent 158 were a user directly interacting with the data source.
[0094] Figure 4 This is a block diagram illustrating the workflow of actions performed by AI agent 158 using tool 159. User interface 153 receives query 402, for example, from user device 115. Query 402 may be associated with a user. Based on query 402, AI agent 158 formulates an execution plan to complete a task to satisfy query 402. AI agent 158 generates the execution plan as including a set of actions that are performed using the selected tool 159 to interact with the corresponding data source system 160 for the selected tool 159. AI agent 158 is configured / trained to have actions available via each of the tools 159 and selects one or more appropriate tools 159 for the user associated with query 402 based on permissions according to RBAC 172.
[0095] The execution plan can be dynamic; that is, rather than a series of static actions, some actions can depend on the results of previous actions. Furthermore, the AI agent 158 can modify the execution plan based on data obtained from the data source system 160 as the execution plan continues through the execution phases.
[0096] exist Figure 4In this process, AI agent 158 executes the generated execution plan to obtain data from data source systems 160B and 160K using actions performed using tools 159B and 159N, respectively. This data determines the actions to be performed using tool 159A regarding data source system 160A. AI agent 158 generates and outputs response 404 to user device 115 in response to query 402. In various examples, user device 115 may alternatively be any of devices 108, 109, or the device of application system 102.
[0097] In some examples, control plane 220 proxies communication between AI agent 158 and tool 159. For instance, AI agent 158 may provide control plane 220 with instructions for actions performed by tool 159B. Control plane 220 may then invoke tool 159B on behalf of AI agent 158 to perform the action based on the instructions. Control plane 220 may respond to AI agent 158 with any data, instructions, or other information provided by tool 159B as a result of the action being performed. In this way, AI agent 158 can be implemented by a separate system, while control plane 220 inherits the user's privileges and invokes tool 159 as if it were a user directly interacting with a data source. This can promote privacy and security in cases where a separate system is provided by a third party (e.g., a cloud provider).
[0098] Figure 5 This is a flowchart illustrating example operations according to one or more techniques disclosed herein. For example... Figure 5 As seen in the example, the operation may involve receiving a query for processing by an AI agent, the query including a request for the execution of a task associated with a first data source system and a second data source system (500). The AI agent 158 may generate an execution plan for the task based on the query, wherein the execution plan includes actions to be performed regarding the first and second data source systems, and wherein the user associated with the first and second data source systems has permissions for each action in the actions (505). Next, the AI agent 158 may invoke a first tool to execute a first action regarding the first data source system in the actions (510). Next, the AI agent 158 may invoke a second tool to execute a second action regarding the second data source system in the actions (515).
[0099] For the processes, devices, and other examples or illustrations described herein, including in any flowchart or schematic, certain operations, actions, steps, or events included in any of the techniques described herein may be performed in a different order, may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for the practice of the technique). Furthermore, in some examples, operations, actions, steps, or events may be performed simultaneously, for example, through multithreading, interrupt handling, or multiple processors, rather than sequentially. Additionally, some operations, actions, steps, or events may be performed automatically even if not explicitly identified as automatic. Moreover, some operations, actions, steps, or events described as automatically performed may alternatively not be automatically performed; rather, in some examples, such operations, actions, steps, or events may be performed in response to input or another event.
[0100] The detailed description set forth below in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to represent only configurations in which the concepts described herein can be practiced. The detailed description includes specific details for the purpose of providing a comprehensive understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts can be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.
[0101] According to one or more aspects of this disclosure, the term "or" may be interpreted as "and / or" unless otherwise specified in the context. Additionally, while phrases such as "one or more" or "at least one" may be used in some instances, they may not be used in others; those instances where such language is not used may be interpreted as having this implied meaning unless otherwise specified in the context.
[0102] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on and / or transmitted via a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium, such as a data storage medium, or a communication medium, including any medium that facilitates the transfer of a computer program from one place to another (e.g., according to a communication protocol). In this way, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0103] For example, and not as a limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store program code in the form of desired instructions or data structures and is accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient tangible storage media. As used, disks and optical discs include compact discs (CDs), laser discs, optical discs, digital universal discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0104] The instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuit systems. Therefore, the terms "processor" or "processing circuit system" as used herein can each refer to any of the foregoing structures or any other structure suitable for implementing the described techniques. Additionally, in some examples, the described functionality may be provided within dedicated hardware and / or software modules. Moreover, the techniques can be implemented entirely within one or more circuit or logic elements.
[0105] The processing component used herein may include the processing circuitry system described above. In some examples, the processing component may include at least one processor and at least one memory having computer code including a set of instructions that, when executed by the at least one processor, cause the at least one processor to perform any of the functions described herein. In some examples, the processing component may receive computer code including the set of instructions from at least one memory coupled to the processing component.
[0106] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless handsets, mobile or non-mobile computing devices, wearable or non-wearable computing devices, integrated circuits (ICs) or IC sets (e.g., chip sets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, various units may be combined within hardware units or provided by a collection of interoperable hardware units (including one or more processors as described above) combined with suitable software and / or firmware.
[0107] Additional details
[0108] This disclosure implements at least the following exemplary process. This process involves a computing system receiving a query for processing by an AI agent. The query includes a request to perform a task associated with a first data source system and a second data source system. The query can specify various types of data operations or analyses to be performed across the data source systems.
[0109] The AI agent generates an execution plan based on the received query. This execution plan defines the sequence of actions to be performed on both the first and second data source systems to satisfy the requirements specified in the query. The execution plan may include multiple discrete actions that may need to be executed in a specific order or in parallel. Users associated with the first and second data source systems have permissions for each action within the plan.
[0110] The AI agent proceeds to invoke a first tool to perform a first action on the first data source system. This first tool can be selected based on compatibility with the first data source system and the specific type of action required. The first tool then executes the specified action on the first data source system according to the execution plan.
[0111] The AI agent invokes a second tool to perform a second action concerning a second data source system. Similar to the first tool, the second tool is selected based on its compatibility with the second data source system and the specific requirements of the second action. This second tool performs the specified action on the second data source system, as defined in the execution plan.
[0112] Tools invoked by AI agents can include various software components, APIs, scripts, or other mechanisms suitable for interacting with their respective data source systems. Actions performed by these tools can include data retrieval, modification, analysis, or other operations supported by the corresponding data source system.
[0113] AI agents seamlessly coordinate operations across multiple data source systems through a unified execution plan. This plan determines the optimal sequencing and parallelization of actions to efficiently process tasks across system boundaries. This coordinated approach reduces processing overhead and streamlines data operations across different systems.
[0114] This approach enables the integration and utilization of different tools based on the specific requirements of each data source system. A consistent query interface masks the underlying complexities of tool selection and interaction. This method provides flexibility in tool selection while maintaining a unified experience for processing queries across diverse systems.
[0115] The aspects of this disclosure include the following embodiments.
[0116] Example 1. A computing system comprising: one or more storage devices; and a processing circuitry system capable of accessing the one or more storage devices and configured to: generate an execution plan for a task based on a query associated with a user using an artificial intelligence (AI) agent employing a machine learning model, wherein the execution plan includes actions to be performed with respect to a first data source system and a second data source system, and wherein the user has permissions for each of the actions; invoke a first tool by the AI agent to perform a first action in the actions related to the first data source system, wherein the AI agent is trained to use the first tool; and invoke a second tool by the AI agent to perform a second action in the actions related to the second data source system, wherein the AI agent is trained to use the second tool.
[0117] Example 2. A computing system as described in Example 1, wherein the processing circuitry is configured to: obtain configuration information for the first tool, wherein the configuration information specifies the scope of invocation of the first data source system; and have the AI agent invoke the first tool based on the configuration information.
[0118] Example 3. The computing system as described in Example 1, wherein the processing circuitry is configured to: obtain configuration information for the first tool, wherein the configuration information specifies how the first tool should access data from the first data source system; and have the AI agent invoke the first tool based on the configuration information.
[0119] Example 4. The computing system as described in Example 1, wherein the processing circuitry is configured to: obtain configuration information for the first tool, wherein the configuration information includes specifications describing actions that the first tool can perform with respect to the first data source system; and have the AI agent invoke the first tool based on the configuration information.
[0120] Example 5. A computing system as described in Example 4, wherein the processing circuitry is configured to: process the specification to obtain the action that the first tool can perform with respect to the first data source system; and generate an execution plan based on the action that the first tool can perform with respect to the first data source system and the query, comprising invoking the first tool to perform the first action.
[0121] Example 6. A computing system as described in Example 1, wherein the task includes optimizing the backup of data associated with the user and stored on the first data source system on the second data source system.
[0122] Example 7. A computing system as described in Example 1, wherein the task includes modifying secure data associated with the user and stored on the first data source system or the second data source system on the second data source system.
[0123] Example 8. A computing system as described in Example 1, wherein the first action includes obtaining dynamic data from the first data source system.
[0124] Example 9. A computing system as described in Example 1, wherein the processing circuitry is configured to: authenticate the first tool to the first data source system by a data access proxy layer so that the first tool can perform the first action.
[0125] Example 10. A computing system as described in Example 9, wherein, in order to authenticate the first tool to the first data source system, the processing circuitry is configured to authenticate the first tool to the first data source system based on the user's credentials.
[0126] Example 11. A computing system as described in Example 1, wherein the first action includes sending an application programming interface (API) call to an API implemented by the first data source system.
[0127] Example 12. The computing system as described in Example 1, wherein, in order to generate the execution plan, the processing circuitry is configured such that the AI agent selects the action of the execution plan to include the specific action based on determining that the user has permission to perform a specific action that the first tool can perform with respect to the first data source system.
[0128] Example 13. A computing system as described in Example 1, wherein the first data source system and the second data source system are diversified.
[0129] Example 14. A method comprising: generating an execution plan for a task based on a query associated with a user, using an artificial intelligence (AI) agent executed by a computing system and applying a machine learning model to satisfy the query, wherein the execution plan includes actions to be performed with respect to a first data source system and a second data source system, and wherein the user has permissions for each of the actions; having the AI agent invoke a first tool to perform a first action in the actions with respect to the first data source system, wherein the AI agent is trained to use the first tool; and having the AI agent invoke a second tool to perform a second action in the actions with respect to the second data source system, wherein the AI agent is trained to use the second tool.
[0130] Example 15. The method as described in Example 14 further includes: obtaining configuration information for the first tool, wherein the configuration information specifies the scope of invocation to the first data source system; and having the AI agent invoke the first tool based on the configuration information.
[0131] Example 16. The method of Example 14 further includes: obtaining configuration information for the first tool, wherein the configuration information specifies how the first tool should access data from the first data source system; and having the AI agent invoke the first tool based on the configuration information.
[0132] Example 17. The method of Example 14 further includes: obtaining configuration information for the first tool, wherein the configuration information includes specifications describing the actions that the first tool can perform with respect to the first data source system; and the AI agent invoking the first tool based on the configuration information.
[0133] Example 18. The method of Example 17 further includes: processing the specification to obtain the action that the first tool can perform with respect to the first data source system; and generating an execution plan based on the action that the first tool can perform with respect to the first data source system and the query, which includes invoking the first tool to perform the first action.
[0134] Example 19. The method as described in Example 14, wherein the task includes optimizing the backup of data associated with the user and stored on the first data source system on the second data source system.
[0135] Example 20. A non-transitory computer-readable medium comprising instructions, which, when executed by a processing circuitry system, cause the processing circuitry system to: generate an execution plan for a task based on a query associated with a user, using an artificial intelligence (AI) agent employing a machine learning model to satisfy the query, wherein the execution plan includes actions to be performed with respect to a first data source system and a second data source system, and wherein the user has permissions for each of the actions; invoke a first tool by the AI agent to perform a first action in the actions concerning the first data source system, wherein the AI agent is trained to use the first tool; and invoke a second tool by the AI agent to perform a second action in the actions concerning the second data source system, wherein the AI agent is trained to use the second tool.
Claims
1. A method comprising: The computing system receives queries for processing by an artificial intelligence (AI) agent, the queries including requests for the execution of tasks associated with a first data source system and a second data source system; The AI agent generates an execution plan for the task based on the query to satisfy the query, wherein the execution plan includes actions to be performed on the first data source system and the second data source system, and wherein the user associated with the first data source system and the second data source system has permissions for each of the actions; The AI agent invokes a first tool to execute the first action concerning the first data source system; and The AI agent invokes a second tool to execute the second action concerning the second data source system.
2. The method according to claim 1, further comprising: Obtain configuration information for the first tool, wherein the configuration information specifies at least one of the following: The scope of calls to the first data source system; and The method by which the first tool accesses data from the first data source system; or The configuration information includes specifications describing the actions that the first tool can perform with respect to the first data source system; and The AI agent retrieves the first tool based on the configuration information.
3. The method according to claim 2, wherein, The configuration information includes the specification, and the method further includes: Processing the specification to obtain the actions that the first tool can perform with respect to the first data source system; and Based on the actions and queries that the first tool can perform with respect to the first data source system, the execution plan is generated to include invoking the first tool to perform the first action.
4. The method according to claim 1, wherein, The task includes optimizing the backup of data associated with the user and stored on the first data source system on the second data source system.
5. The method according to claim 1, wherein, The task includes modifying security data associated with the user stored on the first or second data source system on the second data source system.
6. The method according to claim 1, wherein, The first action includes obtaining dynamic data from the first data source system.
7. The method of claim 1, further comprising: The data access proxy layer authenticates the first tool with the first data source system so that the first tool can perform the first action, and Optionally, the method includes: The first tool is authenticated to the first data source system based on the user's credentials, wherein the authentication of the first tool to the first data source system is performed.
8. The method according to claim 1, comprising: The AI agent, based on determining that the user has permission to perform a specific action that the first tool can perform with respect to the first data source system, selects the action in the execution plan to include the specific action, wherein, in order to generate the execution plan.
9. The method according to claim 1, wherein, The first data source system and the second data source system are diverse.
10. The method according to claim 1, wherein, The first tool and the second tool are invoked concurrently.
11. The method according to claim 1, wherein, The first tool is configured to interface with the first data source system but not with the second data source system, the second tool is configured to interface with the second data source system but not with the first data source system, and the method further includes: The first tool and the second tool are selected from multiple tools based on the query.
12. A computing system, comprising: One or more storage devices; and A processing component, which is capable of accessing the one or more storage devices, and is configured to: Receive queries for processing by an artificial intelligence (AI) agent, the queries including requests for the execution of tasks associated with a first data source system and a second data source system; The AI agent generates an execution plan for the task based on the query to satisfy the query, wherein the execution plan includes actions to be performed on the first data source system and the second data source system, and wherein the user associated with the first data source system and the second data source system has permissions for each of the actions; The AI agent invokes a first tool to execute the first action concerning the first data source system; and The AI agent invokes a second tool to execute the second action concerning the second data source system.
13. The computing system according to claim 12, wherein, The processing component is configured as follows: Obtain configuration information for the first tool, wherein the configuration information specifies at least one of the following: The scope of calls to the first data source system; and The method by which the first tool accesses data from the first data source system; or The configuration information includes specifications describing the actions that the first tool can perform with respect to the first data source system; and The AI agent retrieves the first tool based on the configuration information.
14. The computing system according to claim 12 or claim 13, wherein, The task includes optimizing the backup of data associated with the user and stored on the first data source system on the second data source system.
15. One or more computer-readable media, the one or more computer-readable media comprising instructions that, when executed by a processing member, cause the processing member to perform the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Adaptive deduplication of data chunks
US20240311342A1
Data retrieval using embeddings for data in backup systems
US20240370339A1