Method, system and computer program product for data mining with artificial intelligence (AI) based smart agents
AI-based smart Agents address the challenges of traditional data mining by autonomously connecting to diverse data sources and employing NLQ and machine learning, resulting in efficient and accurate data retrieval and analysis.
Patent Information
- Application Number
- PCT/US2024/057955
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-11-29
- Publication Date
- 2025-06-05
AI Technical Summary
Traditional data mining techniques struggle with the volume, velocity, and complexity of modern data, particularly in distributed systems, heterogeneous data formats, and incomplete data, leading to inefficiencies and time-consuming processes in extracting meaningful insights.
The use of artificial intelligence (AI) based smart Agents that can autonomously connect to various data sources, including databases, spreadsheets, and documents, enabling intelligent interaction and natural language querying (NLQ) through messaging modalities like WhatsApp. These Agents employ machine learning algorithms and large language models (LLMs) to enhance data mining accessibility, adaptability, accuracy, and efficiency.
The AI-based smart Agents streamline data mining processes by enabling rapid, efficient, and accurate data retrieval and analysis, reducing the need for extensive resources and time, and providing context-aware insights through intelligent querying and data processing pipelines.
Smart Images

Figure US2024057955_05062025_PF_FP_ABST
Abstract
Description
[0001] METHOD, SYSTEM AND COMPUTER PROGRAM PRODUCT FOR DATA MINING WITH ARTIFICIAL INTELLIGENCE (Al) BASED SMART AGENTS
[0002] FIELD OF THE INVENTION
[0003] The present invention relates to the field of data analytics, more particularly pertaining to Artificial Intelligence (Al) based smart Agents capable of autonomous connection to enterprise databases.
[0004] BACKGROUND OF THE INVENTION
[0005] In the digital age of today, enormous amounts of data are generated daily across various industries and applications. Data mining plays a pivotal role in extracting valuable insights from vast datasets. Extracting meaningful patterns and insights from this data is crucial for informed decision-making, solving problems, predicting trends, mitigating risks, and finding new opportunities. Data mining also includes establishing relationships, as well as finding anomalies and correlations to tackle issues, creating actionable information in the process. Traditional data mining techniques often struggle with the volume, velocity, and complexity of modern data. Obstacles include data distributed on several platforms without a centralized repository, complex data in various domains or forms (e.g., structured, unstructured, semi -structured, heterogeneous, multimedia), and incomplete data. Additionally, in terms of data visualization, in order to make the information relevant to a user, incorporation of background knowledge, effective input data, output information, and complicated data perception methods must be used. Moreover, information and metadata are often disorganized within the specific landscape of a company, thus demanding substantial resources from data scientists and programmers to answer a company-related query. This process is normally time-consuming and requires first locating a suitable company dataset, extracting correct data or information from the dataset, analyzing the data or information, putting the analysis results into context and generating meaningful insights. Therefore, there is a need for a method, system and computer program that enable users to easily and autonomously connect to company databases for targeted data mining to obtain a rapid, efficient and accurate answer to a company -related query.
[0006] SUMMARY OF THE INVENTION
[0007] The present invention provides an innovative method, computer system, and computer program product for data mining by employing artificial intelligence (Al) based smart Agents. The plurality of Al Agents is capable of autonomously connecting to company databases, spreadsheets, PDF documents, and various other formats, enabling intelligent interaction and natural language querying (NLQ) with any user through a chat session run via WhatsApp or another high-level messaging modality. Each separate Al Agent is configured to perform specific tasks and may operate simultaneously with its counterparts, thus permitting rapid completion of the task at hand. Expert Al Agents specialize in an explicit topic selected from an index of topics, in which they may analyze large datasets, identify intricate patterns, correlations, and anomalies in real-time data. The present invention further comprises an extract, transform, load (ETL) data processing pipeline for creating a data warehouse and improving company -related information retrieval over time. By leveraging machine learning algorithms and large language models (LLMs), the plurality of Al Agents enhances the accessibility, adaptability, accuracy and efficiency of data mining processes and workflow.
[0008] In one aspect of the invention, a computer-implemented method of data mining, comprising steps of: a. initiating a workflow involving a chat session between a user 10 and an artificial intelligence (Al) Assistant Agent 12, wherein the workflow is provided to the Al Assistant Agent 12 using an interface; b. receiving an input 14 via said chat session between a user device and an Al Agent system by said Al Assistant Agent 12, wherein said chat session is associated with said workflow; c. processing, by said Al Assistant Agent 12, said input 14 with historical inputs, outputs, and user interactions in the chat history 26 with the same user 10; d. determining, by said processing, that at least one self-contained NLQ 22 can be written from a portion of said chat history 26; e. writing said at least one self-contained NLQ 22; f. determining, by said processing, the topic 16 of said self-contained NLQ 22; g. selecting an Al Expert Agent 18 from a pool of Al Expert Agents 20 using an Expert selector 21; said selection based on said topic 16 of said self-contained NLQ 22; h. sending the said at least one self-contained NLQ 22 to said Al Expert Agent 18; i. determining, by said Al Expert Agent 18, a suitable data source 23 to answer said at least one self-contained NLQ 22; j. determining, by said Al Expert Agent 18, an output 24 based on said processing of said at least one self-contained NLQ 22, said output 24 supplementing at least one step in said workflow, wherein said Al Expert Agent 18 may iterate between said output 24 and executing more instructions until acquiring an output 24 comprising an answer to said self- contained NLQ 22; k. sending, by said Al Expert Agent 18, said output 24 to said Al Assistant Agent 12; and l. sending, by said Al Assistant Agent 12, said output to said user 10.
[0009] In another aspect of the invention, the method above is provided, wherein said at least one self- contained NLQ 22 may not be contained in a single input 14, but formulated by said Al Assistant Agent 12 from the full chat history 26 with said user 10.
[0010] In another aspect of the invention, the method as defined in any of above is provided, wherein said at least one self-contained NLQ 22 is formulated by said Al Assistant Agent 12 using large language models (LLMs) algorithms; said NLP algorithms identifying user intent and context to formulate precise queries.
[0011] In another aspect of the invention, the method as defined in any of above is provided, wherein said input 14 is added to a chat history 26 with said user 10; said chat history 26 stored on a database 152 using the open-source library LangChain or another high-level open-source library. In another aspect of the invention, the method as defined in any of above is provided, wherein said pool of Al Expert Agents 20 is integrated in a topic index 30; said topic index 30 used by said Al Assistant Agent 12 to select said topic 16 of said self-contained NLQ 22.
[0012] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al Expert Agent 18 has access to all relevant data in said data source 23 to answer said self- contained NLQ 22; said Al Expert Agent 18 possessing all the relevant knowledge and instructions to accurately answer queries on its topic of expertise.
[0013] In another aspect of the invention, the method as defined in any of above is provided, wherein said answer 24 of said output is a context-aware answer 28 based on said chat history 26 with said user 10.
[0014] In another aspect of the invention, the method as defined in any of above is provided, wherein said chat session is run via WhatsApp or another high-level messaging modality.
[0015] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al Assistant Agent 12 is configured to: a. receive said input 14 via said chat session between said user device and said Al Assistant system, wherein said chat session is associated with a workflow for data query between said Al Assistant Agent 12 and said user 10; b. process said input 14 using said NLP algorithms, wherein said Al Assistant Agent 12 is trained using large language models (LLMs) utilizing said database 152 comprising said chat history 26; c. determine, based on processing said input 14 and said chat history 26 using said NLP algorithms, at least one self-contained NLQ 22; d. determine the topic 16 of said self-contained NLQ 22 based on said topic index 30; e. select an Al Expert Agent 18 from a pool of Al Expert Agents 20 using an Expert selector 21; said selection based on said topic 16 of said self-contained NLQ 22 and said topic index 30; f. send said at least one self-contained NLQ 22 to said Al Expert Agent 18; g. receive an output 24 based on said at least one self-contained NLQ 22, wherein said output 24 comprising an answer 24 to the at least one self-contained NLQ 22 and supplementing at least one step in said workflow; and h. send said answer 24 to said user 10, wherein said answer 24 is context-aware answer 28 based on said chat history 26 with said user 10.
[0016] In another aspect of the invention, the method as defined in any of above is provided, wherein said context-aware answer 28 is formulated by said Al Assistant Agent 12 using said NLP algorithms; said NLP algorithms integrate user intent and context to formulate precise answers. In another aspect of the invention, the method as defined in any of above is provided, wherein said Al Assistant Agent 12 communicates said context-aware answer 28 to said user 10 via multiple communication channels comprising text, voice, and graphical interfaces.
[0017] In another aspect of the invention, the method as defined in any of above is provided, wherein performance of the plurality of Al Agents is enhanced using prompt engineering to pretrain and fine-tune said LLMs; said prompt engineering utilizing extensive datasets comprising historical inputs, outputs and user interactions in chat histories.
[0018] In another aspect of the invention, the method as defined in any of above is provided, further comprising a process wherein a client database 32 is utilized in building a data warehouse 34 by optimizing a flow of data within extract, transform, load (ETL) data processing pipeline 124, comprising steps of: a. asking for connection information 36 from a human Agent 38 by an Al connection Agent 40, and receiving an answer 42 comprising connection details of a connection method 44 providing a connection object 46; said connection method 44 deriving from a pool of connections 48; b. extracting data from distributed and heterogeneous data sources by an Al extraction Agent 50; said Al extraction Agent 50 asking for data localization 52 from a human Agent 54 via said connection method 44, and receiving an answer 56 to enable writing extraction queries 58 for data extraction process 60 providing a raw data layer 62; c. transforming data by an Al transformation Agent 64; said Al transformation Agent 64 asking for data meaning 66 from a human Agent 68, and receiving an answer 70 to enable writing transform queries 72 for transformation process 74 providing a staging data layer 76; and d. loading data to said data warehouse 34 by an Al loading Agent 78; said Al loading Agent 78 writing load queries 80 for loading process 82 of said data warehouse 34.
[0019] In another aspect of the invention, the method as defined in any of above is provided, wherein said human Agent is a user.
[0020] In another aspect of the invention, the method as defined in any of above is provided, wherein said data is tabular data.
[0021] In another aspect of the invention, the method as defined in any of above is provided, wherein said tabular data is read to memory as a Pandas Dataframe 110.
[0022] In another aspect of the invention, the method as defined in any of above is provided, wherein said data warehouse 34 is stored in AWS Redshift or Google BigQuery or another high-level data warehousing solution.
[0023] In another aspect of the invention, the method as defined in any of above is provided, wherein required data is received by: a. said Al loading Agent 84; b. said Al transformation Agent from said Al loading Agent 86; and c. said Al extraction Agent from said Al transformation Agent 88.
[0024] In another aspect of the invention, the method as defined in any of above is provided, wherein each step of said ETL data processing pipeline 124 is repeated in a loop until no new data is required for said each step.
[0025] In another aspect of the invention, the method as defined in any of above is provided, wherein said extraction queries 58, transform queries 72 and load queries 80 are written once; said queries are run in said ETL data processing pipeline 124 for updating said data warehouse 34; said queries are run on a regular basis or triggered by new data in said client database; said regular basis comprising a daily basis or other fixed timeframe. In another aspect of the invention, the method as defined in any of above is provided, wherein said answers from said human Agents in said ETL data processing pipeline 124 are added to a chat history; said chat history between said human Agent and Al connection Agent 90, Al extraction Agent 92, and Al transformation Agent 94, is stored on a database 152.
[0026] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al connection Agent 40 receives regular instruction updates 96 relating to said chat with said human Agent 38.
[0027] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al extraction Agent 50 receives regular instruction updates 98 relating to said chat with said human Agent 54; said instructions 98 based on a database schema 100; said database schema 100 deriving from said connection object 46 with said client database 32.
[0028] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al transformation Agent 64 receives regular instruction updates 102 relating to said chat with said human Agent 68; said instructions 102 based on a raw data schema 104; said raw data schema 104 deriving from said raw data layer 62.
[0029] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al loading Agent 78 receives regular instruction updates 106; said instructions 106 based on a staging schema 108; said staging schema 108 deriving from said staging data layer 76.
[0030] In another aspect of the invention, the method as defined in any of above is provided, further comprising a process wherein a client database 32 is utilized in transforming data by optimizing a flow of data within extract, load, transform (ELT) data processing pipeline 156, comprising steps of a. asking for connection information 158 from a human Agent 160 by an Al connection Agent 162 and receiving an answer 164 comprising connection details of a connection method 166 providing a Bronze data layer 168; said connection method 166 deriving from a pool of connections 170; b. fixing data from distributed and heterogeneous data sources by an Al Fixing Agent 172; said Al Fixing Agent 172 asking for data localization 174 from a human Agent 176 , and receiving an answer 178 to enable writing Fixing queries 180 for data Fixing process 182 providing a Silver data layer 184; c. modeling data by an Al modeling Agent 186; said Al modeling Agent 186 asking for data meaning 188 from a human Agent 190, and receiving an answer 192 to enable writing modeling queries 194 for modeling process 196 providing a Gold data layer 198; and d. generating metadata and findings 200 from said Gold data layer 198.
[0031] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al modeling Agent 186 is configured to write transformation code to generate said metadata and findings 200.
[0032] In another aspect of the invention, the method as defined in any of above is provided, wherein said metadata and findings 200 are presented to said user by said Al Assistant Agent 12.
[0033] In another aspect of the invention, the method as defined in any of above is provided, further comprising steps of interacting with said human Agent 176 during data curation, wherein said Al Fixing Agent 172 performs an interview with said human Agent 176 to clarify outliers, missing data, or issues with data format, and then fixes said data accordingly.
[0034] In another aspect of the invention, the method as defined in any of above is provided, further comprising steps of outlier detection, wherein said Al Fixing Agent 172 identifies outliers in said raw data of said Bronze data layer 168 and triggers a user interaction to ask said human Agent 176 how to handle said outliers.
[0035] In another aspect of the invention, the method as defined in any of above is provided, further comprising steps of answering user queries using said metadata and findings 200, wherein said Al Assistant Agent 12 utilizes said modeled data to respond to user queries by generating output comprising insights and visualizing said metadata and findings 200 in graphs or plots.
[0036] In another aspect of the invention, the method as defined in any of above is provided, further comprising steps of integrating said modeled data with dashboard tools, wherein said modeled data from said Gold data layer 198 is visualized in real-time using AWS Q Al Assistant and QuickSight or similar tools to generate reports and insights. In another aspect of the invention, the method as defined in any of above is provided, further comprising steps of generating machine learning (ML) models on modeled data, wherein said Al Modeling Agent 186 implements ML algorithms to predict or infer patterns based on the fixed data in said Silver data layer 184.
[0037] In another aspect of the invention, the method as defined in any of above is provided, wherein said data is tabular data.
[0038] In another aspect of the invention, the method as defined in any of above is provided, wherein said tabular data is read to memory as a Pandas Dataframe.
[0039] In another aspect of the invention, the method as defined in any of above is provided, wherein said metadata and findings are stored in a data warehouse.
[0040] In another aspect of the invention, the method as defined in any of above is provided, wherein said data warehouse is stored in AWS Redshift or Google BigQuery or another high-level data warehousing solution.
[0041] In another aspect of the invention, the method as defined in any of above is provided, wherein a. required data 86 is received by said Al modeling Agent 186; and b. required data 88 is received by said Al Fixing Agent 172 from said Al modeling Agent 186
[0042] In another aspect of the invention, the method as defined in any of above is provided, wherein each step of said ELT data processing pipeline 156 is repeated in a loop until no new data is required for said each step.
[0043] In another aspect of the invention, the method as defined in any of above is provided, wherein said Fixing queries 180 and modeling queries 72 and load queries 194 are written once; said queries are run in said ELT data processing pipeline 156 for updating; said queries are run on a regular basis or triggered by new data in said client database; said regular basis comprising a daily basis or other fixed timeframe.
[0044] In another aspect of the invention, the method as defined in any of above is provided, wherein said answers from said human Agents in said ELT data processing pipeline 156 are added to a chat history; said chat history between said human Agent and Al connection Agent 202, Al Fixing Agent 204, and Al modeling Agent 206, is stored on a database 152.
[0045] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al connection Agent 162 receives general instructions 208 relating to said chat with said human Agent 160.
[0046] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al Fixing Agent 172 receives general instructions 210 relating to said chat with said human Agent 176; said instructions 210 based on metadata and findings 100; said metadata and findings 100 deriving from said Bronze data layer 168.
[0047] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al modeling Agent 186 receives general instructions 212 relating to said chat with said human Agent 190; said instructions 212 based on metadata and findings 104; said metadata and findings 104 deriving from said Silver data layer 184.
[0048] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al Expert Agent 18 is configured to: a. receive said at least one self-contained NLQ 22 from said Al Assistant Agent 12; b. determine a suitable data source 23 of said at least one self-contained NLQ 22; said data source 23 loaded as Pandas Dataframe 110 in said data warehouse 34; c. determine an answer 24 based on processing of said at least one self-contained NLQ 22, comprising steps of: i. writing instructions in Python 112 or another high-level programming language; ii. executing said Python instructions 112 with Python interpreter 114 in a Python run or another high-level programming language script; said Python interpreter or another high-level programming language interpreter 114 utilizing said Dataframe 110; iii. receiving output 116 from said Python run or another high-level programming script run, wherein said output 226 comprising an answer 24 to the at least one self- contained NLQ 22 and supplementing at least one step in the workflow; and d. send said answer 24 to said Al Assistant Agent 12. In another aspect of the invention, the method as defined in any of above is provided, wherein each step of said processing of said self-contained NLQ 22 is repeated in a loop 118 until said output 116 is determined; said loop 118 writing instructions, running instructions and reading responses until the required information is collected to generate said answer 24.
[0049] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al Expert Agent 18 is further configured to receive regular instruction updates 120 relating to said processing of said self-contained NLQ 22; said regular instructions updates 120 based on a dataframe description 122; said dataframe description 122 deriving from said ETL data processing pipeline 124.
[0050] In another aspect of the invention, the method as defined in any of above is provided, further comprising a process wherein said Al Agents 126 of said ETL data processing pipeline configured to add or fix (or both) config files and queries; said process updating the current ETL data processing pipeline 124 based on a query / accuracy pair 130 related to said answer 24 by said Al Expert Agent 18; said query / accuracy pair 130 created in a process comprising steps of: a. retrieving a query from a dataset comprising queries and correct answers 132; b. sending said query 134 to said Al Expert Agent 18; c. processing said query 134 by said Al Expert Agent 18; d. determining an answer 136 to said query 134; e. determining the semantic distance 138 between said answer 136 by said Al Expert Agent 18 and the correct answer 140; f. determining an accuracy metric 142 based on said semantic distance 138; and g. pairing said query and said accuracy metric 130.
[0051] In another aspect of the invention, the method as defined in any of above is provided, wherein said query / accuracy pair 130 is evaluated; said Al Agents 126 of said ETL data processing pipeline triggered by failed answers in said query / accuracy pair 130 to add or fix (or both) config files and queries to update the current ETL data processing pipeline 124; said failed answers below a predetermined threshold of said query / accuracy pair 130. In another aspect of the invention, the method as defined in any of above is provided, wherein said Al Agents 126 of said ETL data processing pipeline may be required to have conversational interactions with a human 128 to collect client-specific knowledge to perform an update of the current ETL data processing pipeline 124; said update comprising adding missing data to the extraction, fixing transforms or other issues in the ETL data processing pipeline 124.
[0052] In another aspect of the invention, the method as defined in any of above is provided, wherein said client-specific knowledge is company-specific knowledge.
[0053] In another aspect of the invention, the method as defined in any of above is provided, wherein said user 10 is an employee of a company utilizing said method for internal use; said method enabling said employee comprehensive data access, extraction and analysis related to datasets or databases of said company.
[0054] In another aspect of the invention, the method as defined in any of above is provided, wherein said user 10 is a client of a company utilizing said method for external use; said method enabling said client to access and obtain information or data (or both) from said company.
[0055] In another aspect of the invention, the method as defined in any of above is provided, further comprising steps of analyzing historical data providing predictive insights for business decisions. In another aspect of the invention, the method as defined in any of above is provided, further comprising steps for handling large datasets within a system for answering a self-contained NLQ 22, comprising: a. receiving a self-contained NLQ 22; b. retrieving, by an Al structured query language (SQL) Agent 220, a subset of data from a database, wherein the subset is selected based on said self-contained NLQ 22 and optimized to fit within system resource constraints by writing and executing a SQL instruction 221; c. executing said SQL Instruction 221 in Teramot Data Wharehouse 34 database to generate a run output 222; d. sending said Run Output 222 to a Pandas DataFrame 224; e. determining if further processing is required for answering said self-contained NLQ 22, including generating a plot; and f. triggering, by an Al Python Agent 226, further processing if required, to process the Dataframe 224 to get the required information in the Run output 228 or generate a Plot 230 to answer said self-contained NLQ 22.
[0056] In another aspect of the invention, the method as defined in any of above is provided, wherein said run output 222 is sent to an Al supervisor Agent 236 as a SQL query and result 223.
[0057] In another aspect of the invention, the method as defined in any of above is provided, wherein said SQL query and result 223 are examined by said Al supervisor Agent 236, after which: a. if said SQL query and result 223 are enough to answer the self-contained NLQ 22, then an answer 24 is returned; or b. if said SQL query and result 223 require additional processing, further SQL task 238 is sent to said Al SQL Agent 220; or c. if said SQL query and result 223 do not require additional processing, then said run output 222 is sent to said Pandas DataFrame 224.
[0058] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al SQL Agent 220 receives general instructions 225 to perform its tasks; said instructions 225 include descriptions 227 for the available tables built by said ETL process 124.
[0059] In another aspect of the invention, the method as defined in any of above is provided, wherein said ETL process may be replaced by said ELT process 156.
[0060] In another aspect of the invention, the method as defined in any of above is provided, wherein each step of said SQL task 238 is repeated in a loop 219 until said run output 222 is interpreted by said Al SQL Agent 220 as a success for said SQL task 238; said loop 219 writing SQL instructions 221, running instructions and reading responses until the required information is collected.
[0061] In another aspect of the invention, the method as defined in any of above is provided, wherein retrieving the subset of data from the database includes performing iterative SQL queries to adjust the size of the retrieved data based on system resource constraints and ensuring that the dataset is small enough to be handled efficiently.
[0062] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al Python Agent 226 performs iterative execution of Python instructions 232 to form a Python code, interprets said Python code by a Python interpreter 234, refining said Python code based on the output 228 generated from each iteration to process the data to generate a final run output 228 or a plot 230 for answering said self-contained NLQ 22.
[0063] In another aspect of the invention, the method as defined in any of above is provided, further comprising a step of using multimodal processing to interpret both textual and graphical data, wherein said Al Python Agent 226 utilizes an image generated during processing to inform subsequent iterations of the task.
[0064] In another aspect of the invention, the method as defined in any of above is provided, wherein said Al Python Agent 226 receives general instructions 231; said instructions 231 contain table descriptions 233 deriving from said Pandas DataFrame 224.
[0065] In another aspect of the invention, the method as defined in any of above is provided, further comprising a step of managing the task flow by an Al supervisor Agent 236, wherein said Al supervisor Agent 236 assigns an SQL task 238 to said SQL Agent and a Python task 240 to said Al Python Agent based on said self-contained NLQ 22, monitors task progress, and determines whether the output generated by the SQL Agent is sufficient or if further processing is required by the Python Agent.
[0066] In another aspect of the invention, the method as defined in any of above is provided, wherein said run output 228 or plot 230 are sent to said Al supervisor Agent 236 as a Python code and result 242.
[0067] In another aspect of the invention, the method as defined in any of above is provided, wherein said Python code and result 242 are examined by said Al supervisor Agent 236, after which: a. if said Python code and result 242 require additional Python code processing, further Python task 240 is sent to said Al Python Agent 226; or b. if said Python code and result 242 do not require additional Python code processing, then an answer 24 is returned to said Al Assistant Agent 12.
[0068] In another aspect of the invention, the method as defined in any of above is provided, wherein each step of said Python task 240 is repeated in a loop 229 until said run output 228 is determined; said loop 229 writing Python instructions 232, running instructions and reading responses until the required information is collected to generate said final run output 228 or Plot 230.
[0069] In another aspect of the invention, the method as defined in any of above is provided, wherein said plot 230 is attached to the final answer 24.
[0070] In one aspect of the invention, a computer-implemented system for data mining, comprising: a. one or more processors 146; b. one or more computer-readable memories 148; c. one or more computer-readable tangible storage mediums 150; d. program instructions stored on at least one of said one or more tangible storage mediums 150 for execution by at least one of said one or more processors 146 via at least one of said one or more memories 148, wherein said computer-implemented system is capable of performing a method comprising steps of: i. beginning a workflow involving a chat session between a user 10 and an Al Assistant Agent 12, wherein said workflow is provided to said Al Assistant Agent 12 using an interface, and wherein said workflow comprising a plurality of steps in said chat session; ii. receiving an input 14 via said chat session between a user device and an Al Agent system by said Al Assistant Agent 12; iii. processing said input 14 and said chat history 26 by said Al Assistant Agent 12; iv. determining, by said processing, that at least one self-contained NLQ 22 can be written from a portion of said chat history 26; v. writing said at least one self-contained NLQ 22; vi. determining, by said processing, the topic 16 of said at least one self-contained NLQ 22; vii. selecting an Al Expert Agent 18 from a pool of Al Expert Agents 20 based on said topic 16 of said at least one self-contained NLQ 22; viii. sending said at least one self-contained NLQ 22 to said Al Expert Agent 18; ix. determining, by said Al Expert Agent 18, a suitable data source 23 of said at least one self-contained NLQ 22; x. determining, by said Al Expert Agent 18, an output 24 based on said processing of said at least one self-contained NLQ 22, wherein said output 24 supplementing at least one step in said workflow, wherein said output 24 including an answer 24 to said at least one self-contained NLQ 22; xi. sending, by said Al Expert Agent 18, said answer 24 to said Al Assistant Agent 12; and xii. sending, by said Al Assistant Agent 12, a context-aware answer 28 to said user 10; e. one or more databases 152; and f. a data warehouse 34.
[0071] In another aspect of the invention, the system above is provided, wherein said Al Agents utilize NLP algorithms to enhance NLP accuracy.
[0072] In another aspect of the invention, the system as defined in any of above is provided, wherein a personalized user wizard is assigned to said user 10; said personalized user wizard is tailored to the needs of said user 10.
[0073] In another aspect of the invention, the system as defined in any of above is provided, wherein said personalized user wizard is configured to restrict access levels based on user roles of employees within a company.
[0074] In another aspect of the invention, the system as defined in any of above is provided, further comprising a standard setup page for configuring the connections in an ETL pipeline, without the use of Al Agents. In another aspect of the invention, the system as defined in any of above is provided, wherein said user enabled comprehensive data access comprising real-time synchronization with company databases for up-to-date data extraction and analysis.
[0075] In one aspect of the invention, a computer program product for data mining comprising a computer-readable storage medium 150 having program instructions embodied therewith; said program instructions executable by said system to cause said system to: a. receive input 14 via a chat session with a user 10; b. send said input 14 to an Al Assistant Agent 12; c. process said input 14 and said chat history 26 by said Al Assistant Agent 12; d. determine, by said processing, that at least one self-contained NLQ 22 can be written from a portion of said chat history 26; e. write said at least one self-contained NLQ 22; f. determine the topic 16 of said at least one self-contained NLQ 22; g. select an Al Expert Agent 18 from a pool of Al Expert Agents 20 based on said topic 16 of said at least one self-contained NLQ 22; h. send said at least one self-contained NLQ 22 to said Al Expert Agent 18; i. determine, by said Al Expert Agent 18, a suitable data source 23 based on NLP of said at least one self-contained NLQ 22; j. determine, by said Al Expert Agent 18, an output 24 based on said NLP of said at least one self-contained NLQ 22, wherein said output 24 comprising an answer 24 to said at least one self-contained NLQ 22; k. send, by said Al Expert Agent 18, said answer 24 to said Al Assistant Agent 12; and l. send, by the Al Assistant Agent 12, a context-aware answer 28 to said user 10.
[0076] BRIEF DESCRIPTION OF THE FIGURES
[0077] The present invention will be understood and appreciated more fully from the following detailed description taken in conjunction with the figures in which:
[0078] Figure l is a simplified flowchart illustrating the steps involved in the method for data mining, including an artificial intelligence (Al) Assistant Agent and an Al Expert Agent, in accordance with an embodiment of the present disclosure.
[0079] Figure 2 is a flowchart illustrating the steps and Al Agents involved in the extract, transform, load (ETL) data processing pipeline for creating a data warehouse, in accordance with an embodiment of the present disclosure.
[0080] Figure 3 is a simplified example flowchart illustrating the steps involved in operations of an Al Expert Agent, including utilization of a Pandas Dataframe data source, in accordance with an embodiment of the present disclosure.
[0081] Figure 4 is a simplified flowchart illustrating the steps and Al Agents involved in the continuous optimization (CO) framework, in accordance with an embodiment of the present disclosure.
[0082] Figure 5 is a diagrammatic view illustrating the components of the system and computer program product thereof for data mining by employing Al based smart Agents, in accordance with an embodiment of the present disclosure.
[0083] Figure 6 is a flowchart illustrating the steps and Al Agents involved in the extract, load, transform (ELT) data processing pipeline, in accordance with an embodiment of the present disclosure.
[0084] Figure 7 is a simplified example flowchart illustrating the steps involved in operations of Al Agents handling large datasets within a system, in accordance with an embodiment of the present disclosure.
[0085] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0086] For the purposes of promoting an understanding of the principles of the invention, reference will now be made to the embodiments illustrated in the figures and specific language will be used to describe the same. While examples and features of disclosed principles are described herein, modifications, adaptations, and other implementations are possible without limitation of the scope of the disclosed embodiments. Any further applications of the principles as described herein are contemplated as would normally occur to one skilled in the art.
[0087] This disclosure employs open-ended permissive language, indicating for example, that some embodiments “may” employ, involve, or include specific features. The use of the term “may”, and other open-ended terminology is intended to indicate that although not every embodiment may employ the specific disclosed feature, at least one embodiment employs the specific disclosed feature.
[0088] The present disclosure may be embodied in various forms, including as a method, system, apparatus, or computer program, and is not limited to any one form of implementation. The method described can be implemented in a system or computer program product, which performs the steps outlined in the method. In one embodiment, the system or computer program product comprises components, modules, or elements that carry out the method, configured to execute the method steps as described.
[0089] In the following description, various working examples are provided for illustrative purposes. However, it is to be understood the present disclosure may be practiced without one or more of these details. Reference will now be made in detail to non-limiting examples of this disclosure, examples of which are illustrated in the accompanying figures. The examples are described below by referring to the figures, wherein like reference numerals refer to like elements. When similar reference numerals are shown, corresponding description(s) are not repeated, and the interested reader is referred to the previously discussed figure(s) for a description of the like element(s).
[0090] Various embodiments are described herein with reference to a method, system or computer program product. It is intended that the disclosure of one is a disclosure of all. For example, it is to be understood that disclosure of a system described herein also constitutes a disclosure of the method implemented by the system, via, for example, at least one processor. It is to be understood that this form of disclosure is for ease of discussion only, and one or more aspects of one embodiment herein may be combined with one or more aspects of other embodiments herein, within the intended scope of this disclosure.
[0091] The present invention offers a groundbreaking solution for data mining, leveraging the power of artificial intelligence (Al) based smart Agents to efficiently extract data and produce valuable insights from large datasets using large language models (LLMs). By seamlessly integrating into company systems, the plurality of Al Agents empowers users with intelligent querying capabilities, providing insightful and relevant answers. The invention enables data location and retrieval within vast company databases, spreadsheets, PDF documents, and other formats. It offers adaptability, user-friendliness, and continuous learning mechanisms which make it a valuable asset for companies seeking rapid, efficient and accurate data mining solutions.
[0092] The present disclosure involves integrating advanced Al algorithms and NLP techniques into the plurality of Al Agents. Generally, these algorithms enable the Al Agents to understand complex queries, identify relevant data sources, and generate accurate and meaningful answers.
[0093] Furthermore, the present disclosure encompasses training machine learning models by using extensive datasets comprising historical inputs, outputs, and user interactions to improve the ability of the Al Agents to provide intelligent and context-aware answers over time. More specifically, prompt engineering is used to set the state of pretrained large language models (LLMs) as a means to enhance Al performance over specific tasks, as well as fine-tune the LLMs over specific tasks with the same goal. The present disclosure may utilize pre-existing LLM services like OpenAI LLM application programming interface (API) or similar offerings from different third-party providers. Alternatively, the LLMs may be custom-built and hosted specifically for the present invention.
[0094] The present disclosure also involves creating a user-friendly interface in WhatsApp or another high-level messaging modality. The interface may provide for seamless communication between users and the Al Assistant Agent 12. Without limitation, the interface may allow users to input queries through various mediums, such as text, voice, or graphical inputs, enhancing user experience and accessibility. In certain embodiments, a personalized user wizard may be assigned to a user, wherein each personalized user wizard may be tailored to the needs of the user and configured to restrict access levels based on user roles of employees within a company. Furthermore, secure communication protocols are implemented to ensure the confidentiality and integrity of the data exchanged between the Al Agents and external systems. Data encryption and secure authentication methods are employed to protect sensitive information.
[0095] In computing, extract, transform and load (ETL) is the general process of copying data from one or more dataset sources into a destination system which represents the data differently from the source(s), or in a different context than the source(s). Data extraction of the ETL process involves extracting data from homogeneous or heterogeneous sources. Successful and efficient ETL processes often require collecting business knowledge regarding the data, data contextual information, history of database modifications, etc., in order to properly read historical data and transform it to standardized and meaningful data representations. In addition, extract, load and transform (ELT) process may also be used in the present disclosure. The ELT process can be faster since the data is loaded first, allowing for quicker access and analysis. Transformation is deferred and can be done on-demand. It is generally more scalable, especially in cloud-based environments, because it leverages the computing power of modern data warehouses. Moreover, users can apply transformations when necessary, allowing for greater flexibility in analysis and reporting. It is ideal for cloud environments with powerful computational resources, large volumes of data, or when raw data needs to be available for ad-hoc analysis. ELT retains raw data in the warehouse, enabling businesses to perform multiple transformations later as required. This is particularly advantageous for large datasets and complex queries.
[0096] The automated ETL building process of the present disclosure includes a defined set of Al Agents that will collect this type of knowledge by requesting information in the setup stage. Each separate ETL Al Agent is configured to operate simultaneously with its counterparts in performing specific tasks. In the data transformation stage, data may be processed by data cleaning and transforming them into a proper storage format / structure for the purposes of querying and analysis. Finally, the data loading step in the ETL process involves inserting data into the final target database, such as a data warehouse, data mart, data lake, or operational data store.
[0097] In the embodiments of the present disclosure, the ETL data processing pipeline may utilize Al Agents to extract data from the source systems of a company, enforce data quality and consistency standards, and conform data so that separate sources can be used together. In this context, the various Al Agents may be applied to ask for any required knowledge to build the ETL, e.g., meaning of columns, possible exceptions, reasons for missing data, etc. Lastly, the ETL Al Agents may deliver and store the data in a presentation-ready format in a data warehouse so that it may be used for downstream analysis.
[0098] In certain embodiments, the algorithms and plurality of Al Agents of the present disclosure may be customized to cater to specific industries, such as healthcare, finance, or manufacturing. Customization might involve industry-specific terminologies and domain knowledge. The present disclosure may thus be adaptable, scalable, and applicable across various domains and industries, ensuring its versatility and effectiveness in diverse contexts. Additionally, the system may scale horizontally to manage increased workloads, ensuring optimal performance even in high-demand scenarios.
[0099] In other embodiments, the system of the present disclosure may be designed to allow hybrid interactions where complex queries are initially processed by Al Agents, but in cases of ambiguity or intricate scenarios, human Experts may be integrated for further analysis and answer generation. In such cases, the human feedback from the human Expert may be used by either the Al Agents of the ETL data processing pipeline (i.e., Al extraction Agent, Al transformation Agent or Al loading Agent) in updating the ETL process to capture the new knowledge with better data representation, or by the Al Expert Agents on similar queryanswering tasks that may occur in the future. For example, if a human Expert of the company (i.e., company employee) is required to explain how to properly compute the annual STAG'S recurring revenue of the company, that knowledge may then be added to a table / column in the data warehouse after the steps of the ETL process are completed, or added to the available knowledge of an Al Expert Agent for impromptu computing from the available data. The former option, i.e., addition to the data warehouse after completion of the ETL process, is preferred when possible.
[0100] In further embodiments, the transformation stage of the ELT data processing pipeline may further include: 1) a Bronze (Raw Data) Stage: In this initial stage, raw data is received into the system. This raw data may come from various sources, each potentially using different formats, units, and structures. The key operation before transitioning from this stage is data fixing or curation. The data must be transformed into a unified format or view that can be effectively processed in subsequent stages. The transformation process that occurs between the Bronze and Silver data layers may involve: la) Data cleaning: Identifying and handling missing values, null entries, or erroneous data; lb) Data formatting: Standardizing the structure of the data, such as date formats, numeric fields, and categorical values, ensuring consistency across all data points; and 1c) Data normalization: Standardizing units of measurement and ensuring the data conforms to a common framework or specification.
[0101] The output from this transition is a curated dataset that is ready for further analysis or aggregation in the next stage of the pipeline; 2) Silver data layer (Fixed Data): Once the data is fixed and standardized after the Bronze data layer, it proceeds to the Silver data layer. At this stage, the individual data points are combined and structured into tables or datasets that are optimized for analysis. These tables may include relational data, summarizations, and other derived views that provide a more useful structure for further processing. Tasks performed between the Silver and Gold data layers may include: 2a) Data aggregation: Merging or summarizing data into more manageable tables (e.g., summing up daily sales figures into monthly aggregates); 2b) Data indexing: Structuring the data in ways that optimize its use in subsequent machine learning or business intelligence tasks; and 2c) Data linking: Connecting different data sources or tables to build relationships, such as linking customer information with transaction data. The output from this transition is a set of tables that are ready for modeling or advanced analysis, typically involving machine learning (ML) or business intelligence (BI) tasks. The system described herein incorporates a multi -Al Agent ecosystem to automate various tasks within the ELT pipeline, particularly regarding the transformation process, making the entire process more efficient and reducing the need for manual intervention. The plurality of Al Agents may be deployed at different points in the transformation process within the pipeline to automate data processing tasks and interact with the user to gather relevant information when necessary, such as the following: 1) Al Fixing Agent: The Preprocessing or Al FixingAgent is responsible for identifying data quality issues and automating the data curation or “fixing” tasks. This Agent's role is to address common data issues that hinder further processing, such as: la) Fixing formatting issues: For example, standardizing date formats (e.g., converting string data with format “YYYY-MM-DD” to native database Date format); lb) Data cleaning: Identifying missing data points and suggesting how to handle them (e.g., imputation, removal); and the Agent may interact with the user when necessary, asking clarifying questions to resolve ambiguities, such as how to handle outliers, missing values. The Agent asks the user how to handle outlier data points, ensuring that decisions are based on the user’s business requirements or domain knowledge.; and 2) Al Modeling Agent: This Agent, located between the Silver and Gold data layers, creates the aggregated and transformed data to answer specific user queries or provide insights. This Agent is tasked with: 3a) Creating New Columns: After the data has been processed, the Modeling Agent interacts with machine learning models or other analytical tools to answer specific business questions, such as forecasting sales or predicting customer churn; 3b) Machine learning model application: It applies machine learning algorithms to derive insights from the data; and 3c) Graphical output generation: The Agent generates tables suitable for graphs and plots that present the results of the analysis in a user-friendly way. This Agent ensures that the data, once processed, is presented in a manner that is actionable and useful for the end-user, often visualized through dashboards or other reporting tools.
[0102] The system in the present disclosure may also be seamlessly integrated with existing business intelligence and analytics tools, e.g., those provided by AWS, such as the Q Al Assistant and QuickSight. This integration may allow for the creation of dynamic dashboards and reports, visualizing the results of the ETL pipeline's work in real-time, as follows: 1) Q Al Assistant: This AWS tool may be used to generate insights directly from the processed data, leveraging the machine learning models applied after the Gold stage; and 2) QuickSight: The data from the Gold stage may be fed into QuickSight to create interactive visualizations and dashboards, providing real-time insights to end-users. By leveraging these existing tools and others, the system may ensure that data processing is tightly integrated with the client’s existing reporting and visualization infrastructure.
[0103] In further embodiments, the present disclosure may be implemented as a mobile application, allowing users to interact with Al Agents on-the-go. Mobile integration may include features such as voice commands and image recognition for diverse user inputs. Additionally, Al Agents of the present disclosure may communicate outputs comprising answers to company -related queries via multiple communication channels comprising text, voice, and graphical interfaces Reference is now made to Figure 1, which shows a simplified flowchart illustrating the steps and components involved in the method for conversational data analytics, including an Al Assistant Agent and an Al Expert Agent, in accordance with an embodiment of the present disclosure. The present disclosure employs a user-friendly interface, such as WhatsApp or another high-level messaging modality, facilitating communication between a user and an Al Assistant Agent. The workflow begins when the user 10 interacts with the Al Assistant Agent 12, initiating a chat session, wherein the workflow may comprise a plurality of steps. An input 14, i.e., new message, may be received via a chat session between a user device and an Al Agent system by an Al Assistant Agent 12. The new message 14 may be processed by the Al Assistant Agent 12, together with past messages in the chat history 26 with the same user 10, in order to determine and define at least one self-contained natural language query (NLQ) 22. The self-contained NLQ 22 may not necessarily be contained in a single message 14, but is rather formulated from the full chat history 26 with the user 10. Precise queries may be formulated by using NLP algorithms to identify user intent and context. The Al Assistant Agent 12 has exclusive access to the full chat history 26 with a specific user, and is enabled to ask the user 10 for missing information required for building the self-contained NLQ 22. For example, a user 10 may ask the Al Assistant Agent 12 for his "last transactions". The Al Assistant Agent 12 may understand that more specifications are required to remove ambiguities in the new message 14 and ask the user 10 for a "time interval" or a "type of transactions". The user 10 may respond to the Al Assistant Agent 12 with a message such as "last week", wherein the chat history triggers the Al Assistant Agent 12 to write the full self-contained NLQ 22 ("What are the last week transactions made by John Smith?") with its corresponding topic ("users_transactions") and wait for the Expert answer 24. In another example, a user 10 may ask the Al Assistant Agent 12 for " sales total revenue The Al Assistant Agent 12 may understand that more specifications are required, thus asking the user 10 for a "time interval". The user 10 may respond to the Al Assistant Agent 12 with a message such as "last month", wherein the chat history triggers the Al Assistant Agent 12 to write the full self-contained NLQ 22 ("What was the total revenue from sales for the last month?") with its corresponding topic ("users_sales") and wait for the Expert answer 24. Furthermore, the Al Assistant Agent 12 may determine the topic 16 of the self-contained NLQ 22 from a list of available topics in the topics index 30. Based on the topic, an Al Expert Agent 18 may be invoked from a pool of Al Expert Agents 20 using an Expert selector 21, wherein the Al Expert Agent 18 has the resources to successfully answer the self-contained NLQ 22. The Al Expert Agent 18 does not necessarily have access to the topic keyword based upon which it was selected, but it has been built to be an Expert on that topic. Then, the self-contained NLQ 22 may be sent to the Al Expert Agent 18, which does not know that the self-contained NLQ 22 originates from a user 10, such as "John Smith" or other. The Al Expert Agent 18 may then determine a suitable data source 23 to answer the self-contained NLQ 22; ergo, it has access to all relevant data, along with possessing all the relevant knowledge and instructions to accurately answer the queries on its topic of Expertise. After processing the self-contained NLQ 22, the Al Expert Agent 18 may produce an output 24, wherein the output 24 may be a text message, a message with a plot or figure, supplementing at least one step in the workflow, and includes an answer 24 to the at least one self-contained NLQ 22. At this point, the output 24 may be sent to the Al Assistant Agent 12 by the Al Expert Agent 18, which in turn may send a context-aware answer 28, based on the chat history 26, to the user 10. Lastly, the chat session may be stored in a chat history 26 on a database 152 using the open-source library LangChain or another high-level open-source library. Figure l is a flowchart illustrating the steps and Al Agents involved in optimizing the flow of data within the ETL data processing pipeline for creating a data warehouse, in accordance with an embodiment of the present disclosure. This process involves multiple Al Agents, beginning with an Al connection Agent 40 that may establish connections with client databases. The Al connection Agent 40 may interact with a human Agent 38 to obtain connection information 36, and may receive an answer 42 comprising connection details of the connection method 44 from a pool of connections 48, thus ensuring the secure creation of a connection object 46 used for subsequent data extraction. The Al extraction Agent 50 deals with the extraction of data from various sources. The Al extraction Agent 50 may communicate with a human Agent 54 to understand data localization 52 specifics, obtaining an answer 56 comprising extraction queries 58 for the data extraction process 60, to create a raw data layer 62. This layer may serve as the foundation for subsequent data transformation by an Al transformation Agent 64. The Al transformation Agent 64 may transform data into meaningful insights to comprehend the semantic meaning of the extracted data. Through asking for data meaning 66 from a human Agent 68, the Al transformation Agent 64 may receive an answer 70 comprising transform queries 72 for the transformation process 74. Subsequently, the Al transformation Agent 64 may create a staging data layer 76, optimizing data transformations. Once the data is transformed, the Al loading Agent 78 may generate load queries 80 for the loading process 82, thus ensuring efficient storage of the data warehouse 34 with the transformed and optimized data. All queries are run on a regular basis (e.g., daily basis) or triggered by new data in the client database to continuously update the data warehouse 34. The data warehouse 34 may be stored in AWS Redshift or Google BigQuery or another high-level data warehousing solution.
[0104] Figure 3 is a simplified example flowchart illustrating the steps involved in operations of an Al Expert Agent 18, in accordance with an embodiment of the present disclosure. Generally, the inputs and output of the method are: 1) Inputs: self-contained NLQ 22 and Teramot data warehouse 34, derived from ETL Process 124 or ELT process 156 applied to Client Database 32; and 2) Output: answer 24. Once selected from a pool of Al Expert Agents 20, the Al Expert Agent 18 is configured to receive one self-contained NLQ 22 from the Al Assistant Agent 12.
[0105] T1 The Al Expert Agent 18, having all the relevant knowledge and instructions to accurately answer the queries on its topic of Expertise, may determine a suitable data source 23 to answer the self- contained NLQ 22. The data source 23 is loaded as a Pandas Dataframe 110 from the data warehouse 34. Processing of the received self-contained NLQ 22 may then be performed, comprising steps of writing instructions in Python or another high-level programming language 112. The Python instructions 112, or instructions from another high-level programming language, may then be executed with Python interpreter 114, or another high-level programming language interpreter, in a Python run or other, whilst utilizing the Dataframe 110. When the Python instructions 112, or instructions from another high-level programming language, are executed, the Al Expert Agent 18 may receive an output 116, wherein the output 226 may comprise an answer 24 to the self-contained NLQ 22. The Al Expert Agent 18 may iterate between reading the execution response and generating and running more instructions in a loop 118, until it finds the answer to the self-contained NLQ 22, wherein the answer may be a plot or a figure. Once the answer 24 is acquired by the Al Expert Agent 18, it may be sent to the Al Assistant Agent 12.
[0106] Figure 4 is simplified flowchart illustrating the steps and Al Agents involved in the process of continuous optimization (CO) framework. This process provides a CO tool that leverages the capabilities of the Al framework to continuously optimize and tune the performance of the disclosed method, system and computer program product thereof. The process involves Al Agents 126 of the ETL data processing pipeline configured to add or fix (or both) config files and queries. The process is configured to update the current ETL data processing pipeline 124 based on a query / accuracy pair 130 related to an answer 24 generated by the Al Expert Agent 18. The query / accuracy pair 130 may be created in a series of steps. In the first step, a query may be retrieved from a dataset of queries and correct answers 132. In the second step, the query 134 may be sent to an Al Expert Agent 18 for processing. When an answer 136 is determined to the query in the third step, the semantic distance 138 between the answer 136 by the Al Expert Agent 18 and the correct answer 140 may be determined. Then, based on the semantic distance 138, an accuracy metric 142 may be set and paired with the query 134. In the last step, the query / accuracy pair 130 may be evaluated, thus establishing whether addition or fixing (or both) of config files and queries in the current ETL data processing pipeline 124 is necessary. In the case when the accuracy metric 142 of the answer 136 by the Al Expert Agent 18 is below a predetermined value, addition or fixing (or both) of config files and queries may be performed by the Al Agents 126 of the ETL data processing pipeline. In some cases, the Al Agents 126 of the ETL data processing pipeline may be required to have conversational interactions with a human 128 (i.e., setup user) to collect company-specific knowledge to perform an update of the current ETL data processing pipeline 124, e.g., adding missing data to the extraction, fixing transforms or other issues in the ETL data processing pipeline 124.
[0107] Figure 5 is a diagrammatic view illustrating the components of the system and computer program product thereof for mining data by employing Al based smart Agents. Without limitation, the system may include one or more processors 146, one or more computer-readable memories 148, one or more computer-readable tangible storage mediums 150, one or more databases 152, a data warehouse 34, and program instructions stored on at least one of the one or more tangible storage mediums 150 for execution by at least one of the one or more processors 146 via at least one of the one or more memories 148 to perform a method with a plurality of steps described in the present disclosure.
[0108] Figure 6 is a flowchart illustrating the steps and Al Agents involved in optimizing the flow of data within the ELT data processing pipeline, in accordance with an embodiment of the present disclosure. This process involves multiple Al Agents, beginning with an Al connection Agent 162 that may establish connections with client databases. The Al connection Agent 162 may interact with a human Agent 160 to obtain connection information 158, and may receive an answer 164 comprising connection details of the connection method 166 from a pool of connections 170, thus ensuring the secure creation of a Bronze data layer 168 used for subsequent data extraction. The Al Fixing Agent 172 deals with the fixing of data from various sources. The Al Fixing Agent 172 may communicate with a human Agent 176 to understand data localization 174 specifics, obtaining an answer 178 comprising Fixing queries 180 for the data Fixing process 182, to create a Silver data layer 184. This layer may serve as the foundation for subsequent data modeling by an Al modeling Agent 186. The Al modeling Agent 186 may model data into meaningful insights to comprehend the semantic meaning of the fixed data. Through asking for data meaning 188 from a human Agent 190, the Al modeling Agent 186 may receive an answer 192 comprising modeling queries 194 for the modeling process 196.
[0109] Subsequently, the Al modeling Agent 186 may create a Gold data layer 198, optimizing data aggregation and modeling. Once the data is modeled, metadata and findings 200 may be generated from the Gold data layer 198 and presented to the user by the Al Assistant Agent 12. All queries are run on a regular basis (e.g., daily basis) or triggered by new data in the client database.
[0110] Figure 7 illustrates the process flow in handling large datasets for answering a self-contained NLQ 22 using multiple Al Agents. The process flow depicted corresponds to another embodiment of the workflow associated with the Al Expert Agent 18 shown in Figure 3. As such, the input and output patterns of both embodiments of the method disclosed are similar, i.e., 1) Inputs: self-contained NLQ 22 and Teramot data warehouse 34, derived from ETL Process 124 or ELT process 156 applied to Client Database 32; 2) Output: answer 24. The process begins with the Al Supervisor Agent 236 receiving a self-contained NLQ 22. Then, the Al Supervisor Agent 236, acting as a manager, writes a SQL task 238 for the Al SQL Agent 220. The Al Supervisor Agent oversees the entire process, managing task assignments and ensuring the successful completion of each Agent’s task. The Al Supervisor Agent 236 decides whether the Al SQL Agent’s 220 result is sufficient for constructing the answer or if further processing by the Al Python Agent 226 is required. The Al Supervisor Agent 236 also has the responsibility of defining appropriate task descriptions for both the Al SQL 220 and Python 226 Agents, using its overall view of the workflow to direct the process effectively. The Al SQL Agent 220 retrieves only the relevant subset of data required for answering the self-contained NLQ 22, thus optimizing resource use (bandwidth, memory, time). This is a more efficient approach compared to pulling entire tables. The Al SQL Agent 220 is provided with general Agent Instructions 225, including task-specific guidelines for reducing the data size through methods such as filtering, column selection, or aggregation. It also receives Table Descriptions 227 of the data from the Teramot Data Warehouse 34 for the current topic, derived from the ETL Process 124 or ELT Process 156 applied to Client Database 32. The Al SQL Agent 220 can iterate over the SQL execution in a loop 219 until the Run Output 222 is interpreted by the Al SQL Agent 220 as a success for the SQL task, wherein the loop 219 involves writing SQL instructions 221, running instructions and reading responses until the required information is collected. The Al SQL Agent 220 may adjust its query based on the results of each iteration, ensuring that the retrieved subset fits the system’s resource constraints. As a result of the SQL query being executed successfully, the Al SQL Agent 220 downloads the required subset of the table and stores it in memory as a DataFrame 224. A brief description of the data is appended to the Run Output 222 for reference. If additional processing is needed (e.g., further data refinement or visualizations), the Al Python Agent 226 is triggered. The Al Python Agent 226 is provided with general Agent Instructions 231, including task-specific guidelines for processing the data retrieved by the Al SQL Agent 220. It receives a detailed task that may involve further data processing or generating a plot to answer the SCQ. The Agent Instructions 231 contain Table Descriptions 233 of the data from the Dataframe 224. The Al Python Agent 226 can perform iterative execution of Python instructions 232 to form a Python code, interpret the Python code by a Python interpreter 234, refine the Python code based on the output 228 generated from each iteration to process the data to generate a final run output 228 or a plot 230 for answering the self-contained NLQ 22. The Al Python Agent 226 may fix errors and refine the output 228 based on the results of each iteration in a loop 229. Additionally, the Al Python Agent 226 utilizes multimodal capabilities to process both text and images, enabling it to evaluate the visual output (e.g., generated plots) to decide if further iterations are required. If the Al Python Agent 226 successfully generates the required information or plot, the result is added to the Run Output 228 or Plot 230. The Run Output 228 or Plot 230 are then sent to the Al Supervisor Agent 236 as a Python code and Result 242, which are examined by the Al Supervisor Agent 236, after which: 1) if the Python code and result 242 require additional Python code processing, further Python task 240 is sent to the A Python Agent 226; or 2) if the Python code and result 242 do not require additional Python code processing, then an answer 24 is returned to the Al Assistant Agent 12. Alternatively, the Plot 230 may be attached to the final answer 24. An image file (if generated) is securely attached to the final message for the user’s access. These graphical representations may be useful for visualizing data insights and supporting decision-making processes. The system is capable of dynamically generating various types of plots or charts, such as: 1) Bar charts, line charts, and scatter plots for displaying trends or relationships; 2) Heat maps or histograms for understanding distributions; and 3) Correlation plots to visualize relationships between variables. These outputs may be integrated into dashboards, providing a powerful tool for end-users to interact with the results of data processing.
[0111] EXAMPLES
[0112] EXAMPLE 1 : Financial Report Generation for Corporate Employees
[0113] An employee at a multinational corporation is working with the Al Assistant to gather financial data for a quarterly review. The user asks, "Can you generate the profit margins for Q3, compared to Q2 of 2024?" The Al Assistant scans the chat history and creates an NLQ based on the user's request, identifying the topic as "quarterly profit margins." The Assistant selects an Expert Agent from a pool of financial analysis Experts who has access to the company’s internal financial data. The Al Expert Agent processes the NLQ, querying the financial database, and retrieves the profit margins for both quarters. It calculates the percentage change and sends the result back to the Assistant: "The profit margin for Q3 was 15%, a 2% increase from Q2’s 13%." This result helps the employee prepare for the upcoming presentation, providing valuable insights in real-time. The response is received within 15 seconds, making the process efficient and accessible without waiting for a manual report generation.
[0114] EXAMPLE 2: Investment Portfolio Management for Clients
[0115] A client at a wealth management firm queries the Al Assistant: "How has my tech stock portfolio been affected by recent market fluctuations, especially after the 10% drop in tech stocks this month?" The Assistant reviews the chat history and creates an NLQ focused on "tech stock portfolio performance in the context of market changes." It selects an Expert Agent in investment analysis and sends the query to the Expert. The Expert Agent retrieves the client's portfolio data and current market trends from the firm’s internal system, calculating the impact of the market downturn on the portfolio. The Agent provides a report: "Your portfolio’s tech stocks have decreased by 8% this month, but the overall portfolio is still up by 5% year-to-date due to gains in other sectors such as healthcare and energy." This output helps the client understand the effect of market conditions and guides them in making informed investment decisions. The information is delivered within 30 seconds, ensuring that the client has up-to-date insights.
[0116] EXAMPLE 3 : Retail Inventory Management for a Store Manager
[0117] A store manager wants to check inventory levels for a popular product, say "Product X," to ensure that they won’t run out of stock during an upcoming promotion. The manager types, "How are sales for Product X in the past month, and what is our current stock level?" The Al Assistant processes this request and formulates an NLQ based on the sales and inventory data. The Assistant selects an Expert Agent specializing in retail analytics. The Expert Agent queries the store’s inventory system and sales database, analyzing the data for Product X. It generates an answer: "Product X had 3,000 units sold last month, and we currently have 1,200 units in stock. You will need to restock before the promotion starts in 5 days." The response helps the manager ensure stock availability in time for the promotion. This response is delivered within 20 seconds, allowing for quick decision-making.
[0118] EXAMPLE 4: Marketing Campaign Analysis for an Advertising Agency
[0119] A digital marketing analyst at an agency requests the Al Assistant to evaluate the performance of their latest Instagram campaign. The analyst asks, "What were the engagement metrics for the last Instagram campaign, and how did it compare to the previous quarter?" The Assistant processes this request and creates an NLQ focused on "Instagram campaign performance" as the topic. The Assistant selects an Expert Agent specializing in social media analytics and sends the query. The Expert Agent queries the social media platform's API, pulling data on likes, shares, comments, and engagement rates for the current and previous campaigns. It returns an answer: "Your Instagram campaign had an engagement rate of 4%, compared to 3% in Q2, reflecting a 33% increase in interaction. The average number of comments and shares also grew by 15%." This data helps the analyst report to clients and adjust the campaign strategy accordingly. The response is received within 10 seconds, providing actionable insights that can be used in realtime for decision-making.
[0120] The description of the illustrations provided in this invention are not intended to limit or restrict the scope as claimed in any way. The aspects, examples, and details provided in this invention are considered sufficient to convey possession and enable others to make and use the best mode. Implementations should not be construed as being limited to any aspect, example, or detail provided in this invention. Regardless of whether shown and described in combination or separately, the various embodiments (both structural and methodological) are intended to be selectively included or omitted to produce an example with a particular set of embodiments. Having been provided with the description and illustration of the present invention, one skilled in the art may envision variations, modifications, and alternate examples falling within the spirit of the broader aspects of the general inventive concept embodied in this invention that do not depart from the broader scope.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method of data mining, comprising steps of: a. initiating a workflow involving a chat session between a user 10 and an artificial intelligence (Al) Assistant Agent 12, wherein the workflow is provided to the Al Assistant Agent 12 using an interface; b. receiving an input 14 via said chat session between a user device and an Al Agent system by said Al Assistant Agent 12, wherein said chat session is associated with said workflow; c. processing, by said Al Assistant Agent 12, said input 14 with historical inputs, outputs, and user interactions in the chat history 26 with the same user 10; d. determining, by said processing, that at least one self-contained NLQ 22 can be written from a portion of said chat history 26; e. writing said at least one self-contained NLQ 22; f. determining, by said processing, the topic 16 of said self-contained NLQ 22; g. selecting an Al Expert Agent 18 from a pool of Al Expert Agents 20 using an Expert selector 21; said selection based on said topic 16 of said self-contained NLQ 22; h. sending the said at least one self-contained NLQ 22 to said Al Expert Agent 18; i. determining, by said Al Expert Agent 18, a suitable data source 23 to answer said at least one self-contained NLQ 22; j. determining, by said Al Expert Agent 18, an output 24 based on said processing of said at least one self-contained NLQ 22, said output 24 supplementing at least one step in said workflow, wherein said Al Expert Agent 18 may iterate between said output 24 and executing more instructions until acquiring an output 24 comprising an answer to said self-contained NLQ 22; k. sending, by said Al Expert Agent 18, said output 24 to said Al Assistant Agent 12; and1. sending, by said Al Assistant Agent 12, said output to said user 10.
2. The method of claim 1, wherein said at least one self-contained NLQ 22 may not be contained in a single input 14, but formulated by said Al Assistant Agent 12 from the full chat history 26 with said user 10.3 The method of claim 1, wherein said at least one self-contained NLQ 22 is formulated by said Al Assistant Agent 12 using large language models (LLMs) algorithms; said NLP algorithms identifying user intent and context to formulate precise queries.4 The method of claim 1, wherein said input 14 is added to a chat history 26 with said user 10; said chat history 26 stored on a database 152 using the open-source library LangChain or another high-level open-source library.5 The method of claim 1, wherein said pool of Al Expert Agents 20 is integrated in a topic index 30; said topic index 30 used by said Al Assistant Agent 12 to select said topic 16 of said self-contained NLQ 22.6 The method of claim 1, wherein said Al Expert Agent 18 has access to all relevant data in said data source 23 to answer said self-contained NLQ 22; said Al Expert Agent 18 possessing all the relevant knowledge and instructions to accurately answer queries on its topic of expertise.7 The method of claim 1, wherein said answer 24 of said output is a context-aware answer 28 based on said chat history 26 with said user 10.8 The method of claim 1, wherein said chat session is run via WhatsApp or another high- level messaging modality.9 The method of claim 1, wherein said Al Assistant Agent 12 is configured to: a. receive said input 14 via said chat session between said user device and said Al Assistant system, wherein said chat session is associated with a workflow for data query between said Al Assistant Agent 12 and said user 10; b. process said input 14 using said NLP algorithms, wherein said Al Assistant Agent 12 is trained using large language models (LLMs) utilizing said database 152 comprising said chat history 26;c. determine, based on processing said input 14 and said chat history 26 using said NLP algorithms, at least one self-contained NLQ 22; d. determine the topic 16 of said self-contained NLQ 22 based on said topic index 30; e. select an Al Expert Agent 18 from a pool of Al Expert Agents 20 using an Expert selector 21; said selection based on said topic 16 of said self-contained NLQ 22 and said topic index 30; f. send said at least one self-contained NLQ 22 to said Al Expert Agent 18; g. receive an output 24 based on said at least one self-contained NLQ 22, wherein said output 24 comprising an answer 24 to the at least one self-contained NLQ 22 and supplementing at least one step in said workflow; and h. send said answer 24 to said user 10, wherein said answer 24 is context-aware answer 28 based on said chat history 26 with said user 10.
10. The method of claim 9, wherein said context-aware answer 28 is formulated by said Al Assistant Agent 12 using said NLP algorithms; said NLP algorithms integrate user intent and context to formulate precise answers.
11. The method of claim 1, wherein said Al Assistant Agent 12 communicates said context- aware answer 28 to said user 10 via multiple communication channels comprising text, voice, and graphical interfaces.
12. The method of claim 1, wherein performance of the plurality of Al Agents is enhanced using prompt engineering to pretrain and fine-tune said LLMs; said prompt engineering utilizing extensive datasets comprising historical inputs, outputs and user interactions in chat histories.
13. The method of claim 1, further comprising a process wherein a client database 32 is utilized in building a data warehouse 34 by optimizing a flow of data within extract, transform, load (ETL) data processing pipeline 124, comprising steps of: a. asking for connection information 36 from a human Agent 38 by an Al connection Agent 40, and receiving an answer 42 comprising connection details of aconnection method 44 providing a connection object 46; said connection method 44 deriving from a pool of connections 48; b. extracting data from distributed and heterogeneous data sources by an Al extraction Agent 50; said Al extraction Agent 50 asking for data localization 52 from a human Agent 54 via said connection method 44, and receiving an answer 56 to enable writing extraction queries 58 for data extraction process 60 providing a raw data layer 62; c. transforming data by an Al transformation Agent 64; said Al transformation Agent 64 asking for data meaning 66 from a human Agent 68, and receiving an answer 70 to enable writing transform queries 72 for transformation process 74 providing a staging data layer 76; and d. loading data to said data warehouse 34 by an Al loading Agent 78; said Al loadingAgent 78 writing load queries 80 for loading process 82 of said data warehouse 34.
14. The method of claim 13, wherein said human Agent is a user.
15. The method of claim 13, wherein said data is tabular data.
16. The method of claim 13, wherein said tabular data is read to memory as a PandasDataframe 110.
17. The method of claim 13, wherein said data warehouse 34 is stored in AWS Redshift or Google BigQuery or another high-level data warehousing solution.
18. The method of claim 13, wherein required data is received by: a. said Al loading Agent 84; b. said Al transformation Agent from said Al loading Agent 86; and c. said Al extraction Agent from said Al transformation Agent 88.
19. The method of claim 13, wherein each step of said ETL data processing pipeline 124 is repeated in a loop until no new data is required for said each step.
20. The method of claim 13, wherein said extraction queries 58, transform queries 72 and load queries 80 are written once; said queries are run in said ETL data processing pipeline 124 for updating said data warehouse 34; said queries are run on a regular basis ortriggered by new data in said client database; said regular basis comprising a daily basis or other fixed timeframe.
21. The method of claim 13, wherein said answers from said human Agents in said ETL data processing pipeline 124 are added to a chat history; said chat history between said human Agent and Al connection Agent 90, Al extraction Agent 92, and Al transformation Agent 94, is stored on a database 152.
22. The method of claim 13, wherein said Al connection Agent 40 receives regular instruction updates 96 relating to said chat with said human Agent 38.
23. The method of claim 13, wherein said Al extraction Agent 50 receives regular instruction updates 98 relating to said chat with said human Agent 54; said instructions 98 based on a database schema 100; said database schema 100 deriving from said connection object 46 with said client database 32.
24. The method of claim 13, wherein said Al transformation Agent 64 receives regular instruction updates 102 relating to said chat with said human Agent 68; said instructions 102 based on a raw data schema 104; said raw data schema 104 deriving from said raw data layer 62.
25. The method of claim 13, wherein said Al loading Agent 78 receives regular instruction updates 106; said instructions 106 based on a staging schema 108; said staging schema 108 deriving from said staging data layer 76.
26. The method of claim 1, further comprising a process wherein a client database 32 is utilized in transforming data by optimizing a flow of data within extract, load, transform (ELT) data processing pipeline 156, comprising steps of a. asking for connection information 158 from a human Agent 160 by an Al connection Agent 162 and receiving an answer 164 comprising connection details of a connection method 166 providing a Bronze data layer 168; said connection method 166 deriving from a pool of connections 170; b. fixing data from distributed and heterogeneous data sources by an Al Fixing Agent 172; said Al Fixing Agent 172 asking for data localization 174 from a human Agent176 , and receiving an answer 178 to enable writing Fixing queries 180 for data Fixing process 182 providing a Silver data layer 184; c. modeling data by an Al modeling Agent 186; said Al modeling Agent 186 asking for data meaning 188 from a human Agent 190, and receiving an answer 192 to enable writing modeling queries 194 for modeling process 196 providing a Gold data layer 198; and d. generating metadata and findings 200 from said Gold data layer 198.
27. The method of claim 26, wherein said Al modeling Agent 186 is configured to write transformation code to generate said metadata and findings 200.
28. The method of claim 27, wherein said metadata and findings 200 are presented to said user by said Al Assistant Agent 12.
29. The method of claim 26, further comprising steps of interacting with said human Agent 176 during data curation, wherein said Al Fixing Agent 172 performs an interview with said human Agent 176 to clarify outliers, missing data, or issues with data format, and then fixes said data accordingly.
30. The method of claim 26, further comprising steps of outlier detection, wherein said Al Fixing Agent 172 identifies outliers in said raw data of said Bronze data layer 168 and triggers a user interaction to ask said human Agent 176 how to handle said outliers.
31. The method of claim 26, further comprising steps of answering user queries using said metadata and findings 200, wherein said Al Assistant Agent 12 utilizes said modeled data to respond to user queries by generating output comprising insights and visualizing said metadata and findings 200 in graphs or plots.
32. The method of claim 26, further comprising steps of integrating said modeled data with dashboard tools, wherein said modeled data from said Gold data layer 198 is visualized in real-time using AWS Q Al Assistant and QuickSight or similar tools to generate reports and insights.
33. The method of claim 26, further comprising steps of generating machine learning (ML) models on modeled data, wherein said Al Modeling Agent 186 implements MLalgorithms to predict or infer patterns based on the fixed data in said Silver data layer 18434. The method of claim 26, wherein said data is tabular data.
35. The method of claim 26, wherein said tabular data is read to memory as a Pandas Dataframe.
36. The method of claim 26, wherein said metadata and findings are stored in a data warehouse.
37. The method of claim 26, wherein said data warehouse is stored in AWS Redshift or Google BigQuery or another high-level data warehousing solution.
38. The method of claim 26, wherein a. required data 86 is received by said Al modeling Agent 186; and b. required data 88 is received by said Al Fixing Agent 172 from said Al modeling Agent 186.
39. The method of claim 26, wherein each step of said ELT data processing pipeline 156 is repeated in a loop until no new data is required for said each step.
40. The method of claim 26, wherein said Fixing queries 180 and modeling queries 72 and load queries 194 are written once; said queries are run in said ELT data processing pipeline 156 for updating; said queries are run on a regular basis or triggered by new data in said client database; said regular basis comprising a daily basis or other fixed timeframe.
41. The method of claim 26, wherein said answers from said human Agents in said ELT data processing pipeline 156 are added to a chat history; said chat history between said human Agent and Al connection Agent 202, Al Fixing Agent 204, and Al modeling Agent 206, is stored on a database 152.
42. The method of claim 26, wherein said Al connection Agent 162 receives general instructions 208 relating to said chat with said human Agent 160.
43. The method of claim 26, wherein said Al Fixing Agent 172 receives general instructions 210 relating to said chat with said human Agent 176; said instructions 210 based onmetadata and findings 100; said metadata and findings 100 deriving from said Bronze data layer 168.
44. The method of claim 26, wherein said Al modeling Agent 186 receives general instructions 212 relating to said chat with said human Agent 190; said instructions 212 based on metadata and findings 104; said metadata and findings 104 deriving from said Silver data layer 184.
45. The method of claim 1, wherein said Al Expert Agent 18 is configured to: a. receive said at least one self-contained NLQ 22 from said Al Assistant Agent 12; b. determine a suitable data source 23 of said at least one self-contained NLQ 22; said data source 23 loaded as Pandas Dataframe 110 in said data warehouse 34; c. determine an answer 24 based on processing of said at least one self-contained NLQ 22, comprising steps of: i. writing instructions in Python 112 or another high-level programming language; ii. executing said Python instructions 112 with Python interpreter 114 in a Python run or another high-level programming language script; said Python interpreter or another high-level programming language interpreter 114 utilizing said Dataframe 110; iii. receiving output 116 from said Python run or another high-level programming script run, wherein said output 226 comprising an answer 24 to the at least one self-contained NLQ 22 and supplementing at least one step in the workflow; and d. send said answer 24 to said Al Assistant Agent 12.
46. The method of claim 45, wherein each step of said processing of said self-contained NLQ 22 is repeated in a loop 118 until said output 116 is determined; said loop 118 writing instructions, running instructions and reading responses until the required information is collected to generate said answer 24.
47. The method of claim 45, wherein said Al Expert Agent 18 is further configured to receive regular instruction updates 120 relating to said processing of said self-contained NLQ 22; said regular instructions updates 120 based on a dataframe description 122; said dataframe description 122 deriving from said ETL data processing pipeline 124.
48. The method of claim 13, further comprising a process wherein said Al Agents 126 of said ETL data processing pipeline configured to add or fix (or both) config files and queries; said process updating the current ETL data processing pipeline 124 based on a query / accuracy pair 130 related to said answer 24 by said Al Expert Agent 18; said query / accuracy pair 130 created in a process comprising steps of: a. retrieving a query from a dataset comprising queries and correct answers 132; b. sending said query 134 to said Al Expert Agent 18; c. processing said query 134 by said Al Expert Agent 18; d. determining an answer 136 to said query 134; e. determining the semantic distance 138 between said answer 136 by said Al Expert Agent 18 and the correct answer 140; f. determining an accuracy metric 142 based on said semantic distance 138; and g. pairing said query and said accuracy metric 130.
49. The method of claim 48, wherein said query / accuracy pair 130 is evaluated; said Al Agents 126 of said ETL data processing pipeline triggered by failed answers in said query / accuracy pair 130 to add or fix (or both) config files and queries to update the current ETL data processing pipeline 124; said failed answers below a predetermined threshold of said query / accuracy pair 130.
50. The method of claim 48, wherein said Al Agents 126 of said ETL data processing pipeline may be required to have conversational interactions with a human 128 to collect client-specific knowledge to perform an update of the current ETL data processing pipeline 124; said update comprising adding missing data to the extraction, fixing transforms or other issues in the ETL data processing pipeline 124.
51. The method of claim 50, wherein said client-specific knowledge is company-specific knowledge.
52. The method of claim 1, wherein said user 10 is an employee of a company utilizing said method for internal use; said method enabling said employee comprehensive data access, extraction and analysis related to datasets or databases of said company.
53. The method of claim 1, wherein said user 10 is a client of a company utilizing said method for external use; said method enabling said client to access and obtain information or data (or both) from said company.
54. The method of claim 1, further comprising steps of analyzing historical data providing predictive insights for business decisions.
55. The method of claim 1, further comprising steps for handling large datasets within a system for answering a self-contained NLQ 22, comprising: a. receiving a self-contained NLQ 22; b. retrieving, by an Al structured query language (SQL) Agent 220, a subset of data from a database, wherein the subset is selected based on said self-contained NLQ 22 and optimized to fit within system resource constraints by writing and executing a SQL instruction 221; c. executing said SQL Instruction 221 in Teramot Data Warehouse 34 database to generate a run output 222; d. sending said Run Output 222 to a Pandas DataFrame 224; e. determining if further processing is required for answering said self-contained NLQ 22, including generating a plot; and f. triggering, by an Al Python Agent 226, further processing if required, to process the Dataframe 224 to get the required information in the Run output 228 or generate a Plot 230 to answer said self-contained NLQ 22.
56. The method of claim 55, wherein said run output 222 is sent to an Al supervisor Agent 236 as a SQL query and result 223.
57. The method of claim 56, wherein said SQL query and result 223 are examined by said Al supervisor Agent 236, after which: a. if said SQL query and result 223 are enough to answer the self-contained NLQ 22, then an answer 24 is returned; or b. if said SQL query and result 223 require additional processing, further SQL task 238 is sent to said Al SQL Agent 220; or c. if said SQL query and result 223 do not require additional processing, then said run output 222 is sent to said Pandas DataFrame 224.
58. The method of claim 55, wherein said Al SQL Agent 220 receives general instructions 225 to perform its tasks; said instructions 225 include descriptions 227 for the available tables built by said ETL process 124.
59. The method of claim 58, wherein said ETL process may be replaced by said ELT process 15660. The method of claim 55, wherein each step of said SQL task 238 is repeated in a loop 219 until said run output 222 is interpreted by said Al SQL Agent 220 as a success for said SQL task 238; said loop 219 writing SQL instructions 221, running instructions and reading responses until the required information is collected.
61. The method of claim 55, wherein retrieving the subset of data from the database includes performing iterative SQL queries to adjust the size of the retrieved data based on system resource constraints and ensuring that the dataset is small enough to be handled efficiently.
62. The method of claim 55, wherein said Al Python Agent 226 performs iterative execution of Python instructions 232 to form a Python code, interprets said Python code by a Python interpreter 234, refining said Python code based on the output 228 generated from each iteration to process the data to generate a final run output 228 or a plot 230 for answering said self-contained NLQ 22.
63. The method of claim 55, further comprising a step of using multimodal processing to interpret both textual and graphical data, wherein said Al Python Agent 226 utilizes an image generated during processing to inform subsequent iterations of the task.
64. The method of claim 55, wherein said Al Python Agent 226 receives general instructions 231; said instructions 231 contain table descriptions 233 deriving from said Pandas DataFrame 224.
65. The method of claim 55, further comprising a step of managing the task flow by an Al supervisor Agent 236, wherein said Al supervisor Agent 236 assigns an SQL task 238 to said SQL Agent and a Python task 240 to said Al Python Agent based on said self- contained NLQ 22, monitors task progress, and determines whether the output generated by the SQL Agent is sufficient or if further processing is required by the Python Agent.
66. The method of claim 57, wherein said run output 228 or plot 230 are sent to said Al supervisor Agent 236 as a Python code and result 242.
67. The method of claim 66, wherein said Python code and result 242 are examined by said Al supervisor Agent 236, after which: a. if said Python code and result 242 require additional Python code processing, further Python task 240 is sent to said Al Python Agent 226; or b. if said Python code and result 242 do not require additional Python code processing, then an answer 24 is returned to said Al Assistant Agent 12.
68. The method of claim 65, wherein each step of said Python task 240 is repeated in a loop 229 until said run output 228 is determined; said loop 229 writing Python instructions 232, running instructions and reading responses until the required information is collected to generate said final run output 228 or Plot 230.
69. The method of claim 57, wherein said Plot 230 is attached to the final answer 24.
70. A computer-implemented system for data mining, comprising: a. one or more processors 146; b. one or more computer-readable memories 148; c. one or more computer-readable tangible storage mediums 150;d. program instructions stored on at least one of said one or more tangible storage mediums 150 for execution by at least one of said one or more processors 146 via at least one of said one or more memories 148, wherein said computer-implemented system is capable of performing a method comprising steps of: i. beginning a workflow involving a chat session between a user 10 and an Al Assistant Agent 12, wherein said workflow is provided to said Al Assistant Agent 12 using an interface, and wherein said workflow comprising a plurality of steps in said chat session; ii. receiving an input 14 via said chat session between a user device and an Al Agent system by said Al Assistant Agent 12; iii. processing said input 14 and said chat history 26 by said Al Assistant Agent 12; iv. determining, by said processing, that at least one self-contained NLQ 22 can be written from a portion of said chat history 26; v. writing said at least one self-contained NLQ 22; vi. determining, by said processing, the topic 16 of said at least one self- contained NLQ 22; vii. selecting an Al Expert Agent 18 from a pool of Al Expert Agents 20 based on said topic 16 of said at least one self-contained NLQ 22; viii. sending said at least one self-contained NLQ 22 to said Al Expert Agent 18; ix. determining, by said Al Expert Agent 18, a suitable data source 23 of said at least one self-contained NLQ 22; x. determining, by said Al Expert Agent 18, an output 24 based on said processing of said at least one self-contained NLQ 22, wherein said output 24 supplementing at least one step in said workflow, wherein said output 24 including an answer 24 to said at least one self-contained NLQ 22; xi. sending, by said Al Expert Agent 18, said answer 24 to said Al Assistant Agent 12; andXU. sending, by said Al Assistant Agent 12, a context-aware answer 28 to said user 10; e. one or more databases 152; and f. a data warehouse 34.
71. The system of claim 70, wherein said Al Agents utilize NLP algorithms to enhance NLP accuracy.
72. The system of claim 70, wherein a personalized user wizard is assigned to said user 10; said personalized user wizard is tailored to the needs of said user 10.
73. The system of claim 70, wherein said personalized user wizard is configured to restrict access levels based on user roles of employees within a company.
74. The system of claim 70, wherein said user enabled comprehensive data access comprising real-time synchronization with company databases for up-to-date data extraction and analysis.
75. A computer program product for data mining comprising a computer -readable storage medium 150 having program instructions embodied therewith; said program instructions executable by said system to cause said system to: a. receive input 14 via a chat session with a user 10; b. send said input 14 to an Al Assistant Agent 12; c. process said input 14 and said chat history 26 by said Al Assistant Agent 12; d. determine, by said processing, that at least one self-contained NLQ 22 can be written from a portion of said chat history 26; e. write said at least one self-contained NLQ 22; f. determine the topic 16 of said at least one self-contained NLQ 22; g. select an Al Expert Agent 18 from a pool of Al Expert Agents 20 based on said topic 16 of said at least one self-contained NLQ 22; h. send said at least one self-contained NLQ 22 to said Al Expert Agent 18; i. determine, by said Al Expert Agent 18, a suitable data source 23 based on NLP of said at least one self-contained NLQ 22;j. determine, by said Al Expert Agent 18, an output 24 based on said NLP of said at least one self-contained NLQ 22, wherein said output 24 comprising an answer 24 to said at least one self-contained NLQ 22; k. send, by said Al Expert Agent 18, said answer 24 to said Al Assistant Agent 12; and l. send, by the Al Assistant Agent 12, a context-aware answer 28 to said user 10.
Citation Information
Patent Citations
Methods and systems for shared language framework to maximize composability of software, translativity of information, and end-user independence
US20230252233A1
Method and system for generating intent responses through virtual agents
US20230350929A1
Conversation orchestration in interactive agents
WO2022149076A1
Cited By
Enterprise data processing method and device based on large language model and multi-agent cooperation, equipment and medium
CN122594028A