Systems and methods for adaptive ai-driven workflow optimization
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-05
- Publication Date
- 2026-08-13
Smart Images

Figure IMGF000031_0001_TABLE 
Figure IMGF000039_0001_TABLE 
Figure 00000078_0000
Abstract
Description
Docket No. FSP2504PCTSYSTEMS AND METHODS FOR ADAPTIVE AI-DRIVEN WORKFLOW OPTIMIZATIONCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. provisional patent application serial no. 63 / 868,138, filed on August 21, 2025, and of U.S. provisional patent application serial no. 63 / 754,447, filed on February 5, 2025, the contents of each of which are incorporated herein by reference in their entirety.FIELD OF INVENTION
[0002] The present disclosure relates generally to workflow optimization, and more specifically to adaptive artificial intelligence-based document creation and workflow optimization.BACKGROUND
[0003] Artificial intelligence (Al) models such as language models are promising tools for streamlining data analysis and report creation. Organizations are increasingly integrating Al solutions into their processes in order to optimize their workflows. However, the current state of the art lacks a unified framework for processing, analyzing, and transforming organizational data while maintaining semantic consistency across different presentation formats.
[0004] Current enterprise document processing systems face significant limitations in leveraging Al capabilities. Microsoft Office and other enterprise applications have attempted to integrate Al through "bolt-on" solutions like Copilots, but these additions fail to address fundamental workflow issues. They operate at character and word levels rather than conceptual levels. They maintain legacy application workflows that limit Al integration. They lack semantic understanding of document structure and content. They fail to provide seamless transformation between different document formats. They cannot reliably maintain consistency across different presentation formats.
[0005] Moreover, increased reliance on Al is hampered by the tendency of existing Al models to hallucinate while executing tasks, making the outputs generated by Al models unreliable. Hallucination refers to the phenomenon in which an Al model perceives nonexistent observation agent or objects in data, yielding an erroneous output. For example, a language model being used to generate a report may output a handful of correct statements interspersed with unsupported and / or untrue statements presented as fact. Hallucinations such as these mayDocket No. FSP2504PCTbe difficult to detect because they may appear plausible upon first glance but in reality contain false or misleading information. Reducing the incidence of hallucinations is needed in order to increase reliability of Al-based systems and make the integration of Al into organizational workflows feasible.
[0006] The current state of the art lacks a unified framework for processing, analyzing, and transforming enterprise documents while maintaining semantic consistency and enabling AI-powered automation across different output formats.BRIEF SUMMARY
[0007] A method for analyzing data is performed at a computing system. The method involves receiving a user input that includes an indication of the data to be analyzed and at least one request related to the data. An input evaluator agent accesses the data that is present in available data silos. The input evaluator agent then determines whether the request is at least partially satisfied by a cached result that was previously generated using at least one language model from a plurality of language models. If the request is not at least partially satisfied by the cached result, a response generator agent receives the user input and the data, selects a set of language models from the plurality of language models, provisions the selected language models as artificial intelligence (Al) custom agents to analyze the data in view of the user input, and analyzes the data with the Al custom agents to generate a response to the request.
[0008] A system includes one or more processors and memory that stores instructions. The instructions, when executed by the processors, cause the system to receive a user input including an indication of the data to be analyzed and at least one request related to the data, access (by an input evaluator agent) the data in available data silos, and determine (by the input evaluator agent) whether the request is at least partially satisfied by a cached result previously generated by at least one language model of a plurality of language models. If the request is not at least partially satisfied by the cached result, the system receives the user input and data at a response generator agent, selects (by the response generator agent) a set of language models from the plurality of language models, provisions the selected language models as Al custom agents to analyze the data in view of the user input, and analyzes the data using the Al custom agents to generate a response to the request.BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGSDocket No. FSP2504PCT
[0009] To easily identify the discussion of any particular element or act, the most significant digit or digits in a reference number refer to the figure number in which that element is first introduced.
[0010] FIG. 1 illustrates an AST-based data processing approach 100 in accordance with one embodiment.
[0011] FIGs. 2A-2F illustrate a routine 200a for analyzing data in accordance with one embodiment, including subroutine 200b, subroutine 200c, subroutine 200d, subroutine 200e, and subroutine 200f.
[0012] FIG. 3 illustrates an exemplary system architecture 300 in accordance with one embodiment.
[0013] FIG. 4 illustrates an exemplary system architecture 400 in accordance with one embodiment.
[0014] FIG. 5 illustrates an exemplary system architecture 500 in accordance with one embodiment.
[0015] FIG. 6 illustrates an exemplary process for adaptive document creation and workflow optimization 600 in accordance with one embodiment.
[0016] FIG. 7 illustrates an exemplary pipeline for adaptive document creation and workflow optimization 700 in accordance with one embodiment.
[0017] FIG. 8 illustrates an exemplary process flow for decision-making and operational efficiency 800 in accordance with one embodiment.
[0018] FIG. 9 illustrates an exemplary distributed processing system 900 in accordance with one embodiment.
[0019] FIG. 10 illustrates an exemplary process flow for handling user queries 1000 in accordance with one embodiment.
[0020] FIG. 11 illustrates an exemplary process flow for splicing model outputs 1100 in accordance with one embodiment.
[0021] FIG. 12 illustrates a method 1200 in accordance with one embodiment.
[0022] FIG. 13 illustrates a stages of an exemplary model evaluation process 1300 in accordance with one embodiment.
[0023] FIG. 14 illustrates an exemplary use case-specific adapter 1400 in accordance with one embodiment.Docket No. FSP2504PCT
[0024] FIG. 15 illustrates an exemplary kernel deployment at the edge 1500 in accordance with one embodiment.
[0025] FIG. 16 illustrates an exemplary kernel deployment in the cloud 1600 in accordance with one embodiment.
[0026] FIG. 17 illustrates an Al hosting edge-cloud configurations 1700 in accordance with one embodiment.
[0027] FIG. 18 illustrates an exemplary use case 1800 in accordance with one embodiment.
[0028] FIG. 19 illustrates an exemplary use case 1900 in accordance with one embodiment.
[0029] FIG. 20 illustrates an exemplary use case 2000 in accordance with one embodiment.
[0030] FIG. 21 illustrates an exemplary use case 2100 in accordance with one embodiment.
[0031] FIG. 22 illustrates an exemplary use case 2200 in accordance with one embodiment.
[0032] FIG. 23 illustrates an exemplary use case 2300 in accordance with one embodiment.
[0033] FIG. 24 illustrates an exemplary use case 2400 in accordance with one embodiment.
[0034] FIG. 25 illustrates an autonomous agent growth through natural language templates 2500 in accordance with one embodiment.
[0035] FIG. 26 illustrates an exemplary query and output 2600 in accordance with one embodiment.
[0036] FIG. 27 illustrates an exemplary query and output 2700 in accordance with one embodiment.
[0037] FIG. 28 illustrates an artificial intelligence or Al system 2800 in accordance with one embodiment.
[0038] FIG. 29 illustrates a cloud computing node 2900 in accordance with one embodiment.
[0039] FIG. 30 illustrates a cloud computing environment 3000 in accordance with one embodiment.
[0040] FIG. 31 illustrates an item 3100 in accordance with one embodiment.
[0041] FIG. 32 illustrates an exemplary tokenizer 3200 in accordance with one embodiment.
[0042] FIG. 33 illustrates a deep neural network 3300 in accordance with one embodiment.
[0043] FIG. 34A illustrates inference and / or training logic 3400a in accordance with one embodiment.Docket No. FSP2504PCT
[0044] FIG. 34B illustrates inference and / or training logic 3400b in accordance with one embodiment.
[0045] FIG. 35 illustrates a basic deep neural network 3500 in accordance with one embodiment.
[0046] FIG. 36 illustrates an artificial neuron 3600 in accordance with one embodiment.DETAILED DESCRIPTION
[0047] Described herein are systems and methods for adaptive document creation and workflow optimization leveraging an abstract semantic tree (AST) framework. The disclosed solution employs Al-driven tools to transform hierarchical argument structures such as ASTs into varied output formats, such as documents, presentations, and videos. A reliability layer ensures accuracy, timing, and cost efficiency through a novel logical language model (LM) architecture. The system dynamically orchestrates an ensemble of artificial intelligence (Al) models to meet user-defined performance constraints, enabling seamless integration into enterprise environments while maintaining guaranteed performance characteristics unavailable in conventional Al systems.
[0048] This disclosure introduces systems (FIGs. 3-5) and methods (FIGs. 6, 8, and 10-13) for adaptive document creation and workflow optimization using an AST-based approach. The disclosed system provides a novel Al-first integration layer that operates on top of existing office applications. It incorporates a small agent framework for the enterprise (SAFE) Al that processes documents using an AST architecture, enabling intelligent document transformation and generation across multiple formats while maintaining conceptual consistency.
[0049] SAFE Al is a specialized Al approach designed to run small, specialized, and compact LMs directly on user infrastructure rather than relying on massive, complex models. It aims to make enterprise-scale Al secure, easy to deploy, and cost-effective. Instead of one large, monolithic Al model, SAFE Al uses multiple smaller, specialized agents, increasing efficiency and reducing the computational overhead. This framework allows for the deployment of compact models in secure environments, addressing data privacy and security concerns for enterprises. This SAFE Al approach serves as the underlying technology for the systems and methods disclosed herein.
[0050] Key innovations include: an AST representation for documents that captures conceptual structure, a reliable logical LM architecture with guaranteed performanceDocket No. FSP2504PCTcharacteristics, a distributed processing framework that efficiently utilizes local and cloud resources, a novel caching and evaluation system that optimizes performance and reliability, and a model farm diffuser system that intelligently orchestrates multiple Al models.
[0051] These elements work together to generate and inform Al custom agents configured to perform specific tasks in parsing large amounts of aggregated data and performing analyses across that data to provide detailed and insightful responses to user queries. An Al custom agent is a modular, dynamically generated computer logic or code construct defined by the triad of skills (tool access), context (dynamic buffer), and reasoning (language models). Each individual Al custom agent is not supported by a specifically trained Al model or LM. Rather an Al custom agent is generated through the steps of context injection, skill binding, and instructional mapping. During context injection, a dynamic buffer may be populated with specific business logic and rules of engagement. During skill binding, specific skills or tools may be assigned for the agent's use, such as a SQL reader, a PDF parser, a web search tool, etc. These may allow the agent to interact with available data sources. During instructional mapping, the reasoning model may be provided with a roadmap of how to interpret data it encounters within the given business context. In this manner, a strong reasoning model may be used to create new Al custom agents on the fly at run-time by giving them the relevant skills and pointing them to data.Foundational Technologies
[0052] The disclosed solution builds upon several core technologies that would be familiar to a person having ordinary skill in the art (PHOSITA):• Language Models (LMs) and neural networks• Generative Al• Abstract Semantic Trees and semantic parsing• Distributed systems architecture• Client-server computing• Enterprise document processing systems• Modern office productivity software• Mobile and desktop computing platformsImprovement over State of the Art
[0053] The disclosed solution provides several key improvements over existing systems.Docket No. FSP2504PCT
[0054] Semantic Tree Architecture: Unlike traditional document processors that operate on flat text, this system maintains a hierarchical semantic representation that preserves meaning across transformations (FIG. 3 and FIG. 1).
[0055] Reliable Al Processing: The language model architecture provides guaranteed performance characteristics including accuracy, response time, and resource usage -capabilities not present in current Al systems (FIG. 8, FIG. 7, and FIG. 6).
[0056] Intelligent Resource Utilization: The system automatically balances processing between local devices, departmental servers, and cloud resources (FIG. 9) based on task parameters and available resources.
[0057] Predictable Performance: The disclosed solution's caching and evaluation system provides consistent response times through intelligent reuse of previous computations (FIG. 6).Technical Benefit / Effect
[0058] The disclosed solution provides several technical benefits.
[0059] Improved Processing Efficiency: By maintaining semantic structure through ASTs, the system reduces redundant processing when transforming between document formats.
[0060] Enhanced Reliability: The language model architecture provides guaranteed performance characteristics previously unavailable in Al systems.
[0061] Resource Optimization: The model farm diffuser system (FIG. 6) intelligently distributes processing across available resources while maintaining performance guarantees.
[0062] Format Independence: The semantic tree architecture enables lossless transformation between different document formats while preserving meaning and structure.Architecture
[0063] The system architecture (FIG. 3 and FIG. 9) consists of several key components.Abstract Semantic Tree (AST) Engine (FIG. 1):
[0064] Semantic Tree Parsing: The AST parser parses input documents into a structured semantic tree representation, capturing key concepts, arguments, and their relationships. This representation serves as the foundation for further processing and analysis.Docket No. FSP2504PCT
[0065] Hierarchical Structure Maintenance: It maintains the hierarchical organization of concepts and arguments, preserving contextual integrity and enabling complex reasoning and manipulation of the document content.
[0066] Output Format Transformation: The AST parser facilitates seamless transformation between various output formats, ensuring compatibility with different applications and systems while retaining semantic fidelity.Logical Language Model Framework (FIG. 6):
[0067] Input Evaluator: Input evaluator logic evaluates incoming inputs to make decisions on processing needs, such as the level of complexity and the most suitable processing resources.
[0068] Model Farm Diffuser: It orchestrates the selection and deployment of multiple Al models, optimizing for parameters like capability, latency, cost, and uptime to meet the demands of the input query.
[0069] Performance Guarantee Management System: A dedicated system ensures that performance metrics such as response time, accuracy, and uptime adhere to predefined servicelevel agreements (SLAs).
[0070] Caching and Code Generation System: Efficient caching mechanisms reduce redundant computations, while the code generation system dynamically produces optimized outputs tailored to specific use cases.Distributed Processing System (FIG. 6 and FIG. 9):
[0071] Local Processing on Client Devices: Enables computation to occur locally on client devices, reducing latency and improving responsiveness while minimizing dependency on external resources.
[0072] Departmental Server Integration: Facilitates seamless integration with departmental servers, balancing workload distribution and enhancing system reliability within localized networks.
[0073] Cloud Service Integration: Incorporates scalable cloud services to handle resourceintensive tasks, ensuring high availability and scalability for large-scale operations.
[0074] Resource Optimization and Load Balancing: Dynamically optimizes resource allocation and balances processing loads across local, departmental, and cloud infrastructures to maximize efficiency and prevent bottlenecks.Docket No. FSP2504PCTPipeline Processor (FIG. 7):
[0075] Find and Filter Stage: Identifies and filters relevant data from input sources, streamlining the pipeline to focus on high-value information.
[0076] Analyze and Process Stage: Applies analytical tools and algorithms to process filtered data, extracting insights, observation agent, and actionable results.
[0077] Recommend and Report Stage: Generates tailored recommendations or reports based on analyzed data, presenting clear and concise outputs that align with user needs and objectives.
[0078] The main characteristic of this architecture is in the Input Evaluation (FIG. 6). This takes the input from the user and dynamically decides during runtime where it should return a cached result or not. If the system finds a pre-existing generated result that satisfies the query by a (configurable) factor of at least 98% in one embodiment, it will draw the result from the cache. Otherwise it will generate a new result and store this result along with associated metadata in the cache for future use.
[0079] The system caching strategy handles two main categories of generated content: data and code. Code is generated to process and act on the data. The type of processing needed depends on the context of the incoming user query.Model Farm Diffuser Overview
[0080] Each node in the system runs multiple language models (EMs) in parallel as an ensemble, integrating a diverse collection of models, code, and other Al systems (FIG. 6 and FIG. 13). These models are characterized by varying levels of performance and resource needs:
[0081] Small Models: These run locally or with minimal resources, offering faster responses at the expense of accuracy.
[0082] Medium Models: Deployed on departmental servers, they provide greater accuracy but are slower than small models.
[0083] Large Models: These models are resource-intensive, requiring significant hardware and time, and may not always return results within a feasible timeframe.
[0084] Experimental Models: These are newly developed models that have not been fully characterized but may be assessed by the diffuser. Since they are still in development, they do not have guaranteed response times and may lack the ability to provide meaningful user input.Docket No. FSP2504PCTDynamic Frequency Control in the Diffuser
[0085] The diffuser (FIG. 6) dynamically adjusts model execution based on the initial input, with a tunable distribution to guide decision-making:
[0086] Exactly the Same: Use small models and spot-check responses 1% of the time.
[0087] Pretty Close: Utilize small and medium models, with a 1% spot-check for verification.
[0088] Complex Input: For challenging inputs, use medium and large models to ensure accuracy.
[0071] The Execution Waiter (FIG. 6) is instructed by the control plane to manage response time expectations. For example, if the guarantee is to provide a response within 5 seconds (i.e. a configurable amount of time), the system will proceed with available responses without further delay. This process may be optimized by a language model or equivalent generative Al model that weighs factors such as response time, accuracy, and cost. Business rules within the control plane — potentially guided by another LM — may define specific guidelines, such as:• "No more than a 2-second delay for easy questions."• "For very difficult questions, notify the user that the question is complex, display a timer, and wait up to 10 seconds before rechecking the response at 30 seconds."
[0089] The diffuser determines which models to call based on the difficulty of the question and may adjust the frequency of model calls to reduce costs. For example, it may vary the size of the ensemble (the "jury") to optimize performance, as larger juries tend to improve accuracy but may also increase resource consumption.Splicing
[0090] In human conversations, it is common for responses to evolve as someone reflects further on a problem. The system mimics this behavior through speculative decoding and splicing (FIG. 11). For simpler problems, speculative decoding may be sufficient, but for more complex inputs, splicing may be employed.
[0091] This process begins with small models generating an initial response. If another model subsequently provides more accurate or insightful data, the diffuser utilizes an Al (e.g., a language model) to integrate the improved response seamlessly into the existing output.
[0092] The system is trained to incorporate transitions naturally, using phrases like, “Hmmm, on further thought, it might actually be this instead.” This approach creates a conversational experience that feels organic and thoughtful.Docket No. FSP2504PCTDelay Signaling
[0093] To further emulate natural human interaction, the system may provide delay signals for complex queries. It generates an initial output to inform downstream systems, such as: "This is a challenging problem and may take some time. May I take a second or two to get back to you?"
[0094] This mechanism manages expectations while maintaining a conversational tone.Skeptic
[0095] The sceptic (FIG. 6) is another safety layer where a trained language model is used to critique the output and compare it with previously cached or trained SAFE output structured artefacts.Use-case specific adapters
[0096] The system comprises lightweight user-specific adapters interfacing with a central processing unit on a standard computing device, which hosts modular models optimized for various tasks. The core device processes user inputs, routes them to appropriate models based on task parameters, and employs caching mechanisms for efficiency. The modular models vary in complexity and resource use, enabling a scalable and efficient workflow for diverse computational tasks, with results delivered to users through the adapters.System Overview
[0097] Adapters: Intermediary layer logic (-200MB) for end-user interfaces (UI) (e.g., CEO, Developer). Serve as intermediary layers for user-specific instructions.
[0098] Core Processing Unit: A standard computing device capable of running multiple processing models. Hosts a variety of lightweight and specialized models optimized for efficiency (e.g., language models, computational routines). Includes caching mechanisms for frequently accessed data and code.
[0099] Models: Modular models with varying resource needs and processing speeds.Examples: Large-scale models optimized for text processing. Compact models designed for specialized tasks.
[0100] Workflow: User input is received via adapters, processed locally by the core device, and routed to relevant models for execution. Results are cached for future use, optimizing repeated queries.Docket No. FSP2504PCTGeneral Workflows
[0101] The system operates through several key workflows (FIG. 6 and FIG. 7).
[0102] Input Document Analysis: The workflow begins with an in-depth analysis of the document to build a contextual understanding of the existing business process workflow.
[0103] Semantic Tree Construction: Both structured and unstructured data are distilled into key concepts, grouped along semantically sensible boundaries. These concepts are then organized into a semantic tree, which includes relational links between them. This tree forms the basis of a contextual ontological graph structure.
[0104] Content Processing and Enhancement: The semantic tree is utilized to process and enhance the document content, ensuring better contextual coherence and alignment with business objectives.
[0105] Structured Output Format Generation: The final step involves generating a structured output format that accurately reflects the processed content and maintains semantic integrity.Model Farm Diffuser Workflow
[0106] Input Evaluation: The workflow begins by assessing the complexity of the input query to determine the parameters for processing.
[0107] Model Selection and Orchestration: Based on the evaluation, the system selects the most suitable language model, considering factors such as model capability, latency, cost, and uptime. The selected models are orchestrated to efficiently handle the query.
[0108] Response Aggregation and Splicing: Outputs from the chosen models are aggregated and spliced to create a cohesive and contextually accurate response.
[0109] Performance Monitoring and Adjustment: The system continuously monitors performance metrics, making adjustments as needed to optimize efficiency and maintain response quality.Resource Optimization Workflow:
[0110] Resource Availability Assessment: The workflow starts with evaluating the availability of resources, including computational power, memory, storage, and network bandwidth. This step ensures an accurate understanding of the current resource pool and identifies any potential bottlenecks or constraints.Docket No. FSP2504PCT
[0111] Processing Distribution Decisions: Based on the resource assessment, decisions are made regarding the optimal distribution of processing tasks. This involves assigning workloads to resources in a manner that maximizes efficiency, minimizes latency, and ensures balanced utilization across the system.
[0112] Cache Utilization: The workflow leverages caching strategies to reduce redundant processing and optimize data retrieval times. Frequently accessed data is stored in cache memory, minimizing the need for repeated resource-intensive operations and enhancing overall throughput.
[0113] Performance Guarantee Management: Continuous monitoring ensures that performance parameters, such as latency, throughput, and uptime, are met consistently.Adjustments are made dynamically to maintain service-level agreements (SLAs) and deliver predictable performance under varying workloads.Use Cases - Needs and Solutions
[0114] Users need to connect to multiple large data systems with a desire for faster information and data privacy. For example, customer support departments want solutions providing real-time information and sentiment scoring. The disclosed solution works for a variety of verticals, including:• Customer support: The disclosed solution may support companies wanting real-time information and consumer sentiment scoring.• Retailers and brands: The disclosed solution may support companies wanting to track demand, supply chain issues, and to forecast trends.• Government agencies: The disclosed solution may support companies wanting to get rich public data like housing, pricing, and public spending trends etc.• Banks: The disclosed solution may support companies wanting easier ways to advise customers on investment strategies based on answering their questions.
[0115] The disclosed solution provides an integration layer over numerous data silos, supporting Al that works across all a user's available data platforms. It works to improve processes that support improved customer experience. Conventional solutions may take hours and days to collage reports, spreadsheets, and web data into something meaningful, while the disclosed solution may response with recommended actions in a reliably actionable timeframe. The disclosed solution may provide all of this within a SAFE and secure Al environment,Docket No. FSP2504PCTproviding checks for hallucinations, may select among multiple models for the best and most secure one for the query, and may build small, custom agents that may run economically on a user's local system.
[0116] Enterprise Recommendation Documents: A business analyst needs to create a recommendation for a new project. The system: analyzes input documents and data, constructs semantic representation, generates multiple output formats (memo, presentation, video), and maintains consistency across all formats.
[0117] Technical Documentation Generation: An engineering team needs to create documentation for a new product. The system: processes technical specifications, generates appropriate documentation structure, creates multiple format outputs (manual, slides, web content), and ensures technical accuracy and consistency.
[0118] Marketing Content Creation: A marketing team needs to create campaign materials. The system: analyzes product information and market data, generates consistent messaging across formats, creates multiple output types (presentations, web content, video scripts), and maintains brand consistency and messaging accuracy.Abstract Semantic Trees (ASTs)
[0119] FIG. 1 illustrates an AST-based data processing approach 100 in accordance with one embodiment. The AST-based data processing approach 100 may allow the disclosed system to interact effectively and efficiently with data in the form of ASTs. The AST-based data processing approach 100 may include a purpose phase 102, an analysis phase 104, a high-level recommendation phase 106, and a detailed recommendation phase 108.
[0120] The purpose phase 102 may focus on providing a purpose or business logic layer to guide the disclosed system in its operation. Tactically, the purpose or business logic layer may be defined via natural language text that encapsulates business rules, business constraints, and identified data sources. The purpose phase 102 may provide the context that may be used to provision Al custom agents as described with respect to routine 200a and its subroutines 200b-200f. At this time, skills / tools and a reasoning model LM may also be provided to provision Al custom agents, which may then perform an analysis during an analysis phase 104.
[0121] The analysis phase 104 may focus on ingestion and analysis by the Al custom agents that underpin the operation of the disclosed systems in performing various disclosed methods and other tasks. The Al custom agents may ingest the context provided in the purpose phaseDocket No. FSP2504PCT102 along with other applicable documentary data. The Al custom agents may here parse the unstructured data provided in available data silos into ASTs and / or knowledge graphs.Knowledge graphs may be mapped to the ASTs as is described in greater detail below. The Al custom agents in this manner may efficiently produce hierarchical responses based on user input. These responses may be provided to a user via a tiered delivery system that optimizes latency. The tiered system may include the high-level recommendation phase 106 and the detailed recommendation phase 108.
[0122] During the high-level recommendation phase 106, a high-level analysis may be performed and a number of high-level recommendations may be developed based on the outcomes of the purpose phase 102 and analysis phase 104. The high-level recommendations may be generated first for quick inference and user feedback. For example, high-level recommendations 106a- 106c may be developed during the high-level recommendation phase 106. These may be presented to a user for selection in order to focus ensuing work toward a particular goal indicated by the user. In one embodiment, the AST-based data processing approach 100 may surface three courses of action it may recommend to the user based on a user query regarding inventory management. High-level recommendation 106a may present recommendation 1: stocking more of low inventory stock keeping units (SKUs). High-level recommendation 106b may present recommendation 2: increasing prices of high-demand SKUs. High-level recommendation 106c may present recommendation 3: lowering prices for and closing out obsolete SKUs.
[0123] The system may, given additional time, identify more granular or additional actions it would recommend toward accomplishing the actions of high-level recommendations 106a-106c. While the user reviews the high-level recommendations, the system may process deeper technical layers in the background as part of a detailed recommendation phase 108. The detailed analysis performed during the detailed recommendation phase 108 may provide detailed recommendations and summaries. This may allow the system to avoid information bottlenecking. In one embodiment, detailed recommendations and summaries may be surfaced once a particular high-level recommendation path is indicated, e.g., by user selection. Thus, user selection of high-level recommendation 106a may lead to the surfacing of a detailed response such as detailed recommendation and summary 108a, selection of high-level recommendation 106b may lead to the surfacing of detailed recommendation and summary 108b, and selection of high-level recommendation 106c may lead to the surfacing of detailedDocket No. FSP2504PCTrecommendation and summary 108c. While FIG. 1 illustrates an example in which three recommendations are presented using the AST-based data processing approach 100, it will be readily understood by one of skill in the art that fewer or more recommendations may be presented by the disclosed systems in various embodiments and under various use conditions or user-indicated parameters.
[0124] The concept of the abstract semantic tree may be used throughout the operation of the disclosed systems as a framework by which large amounts of data may be efficiently stored and indexed. This framework of purpose, analysis, and recommendation branches may direct the Al custom agents disclosed herein in examining the parameters of a user query to develop a purpose and identify documentary data sources, to analyze the data provided by these sources, and to present recommendations, and the data developed during the phases of the AST-based data processing approach 100 may be stored as an AST representation of purpose and analysis parent nodes and recommendation children nodes. In this manner, a record of response to a user query may be saved as an AST for rapid semantic parsing and comparison to future queries in order to provide low-latency answers out of cache, as described in more detail with respect to FIGs. 2A-2F, 6, and 10.
[0125] In one embodiment, an AST parser may parse input documents into a structured semantic tree representation, capturing key concepts, arguments, and their relationships. This representation may serve as the foundation for further processing and analysis. The AST parser may maintain the hierarchical organization of concepts and arguments comprising the AST, preserving contextual integrity and enabling complex reasoning and manipulation of the document content. The AST parser may facilitate seamless transformation between various output formats, ensuring compatibility with different applications and systems while retaining semantic fidelity.
[0126] The analysis phase 104 may follow a “thinking-oriented” reasoning loop. An Al custom agent may identify a data structure to act upon (e.g., a database schematic, comma separated value (CSV) headers, etc.) and verify if the available data satisfies the parameters set forth in the purpose phase 102. The unstructured data provided across multiple data silos may be processed into ASTs, including text, images, charts, graphs, tables, etc. All available items of unstructured data may be treated as discoverable entities.
[0127] In order to develop useful meaning from such large quantities of raw data, the system may use a dual-layered approach. First, the AST parser may develop an AST for each item ofDocket No. FSP2504PCTunstructured data. This AST may define the formal logic, hierarchy, and syntax of an unstructured data item, such as a document, file, etc. This may provide the tree-like framework for efficient reference of the document's information.
[0128] An AST may, in the second layer of this approach, be mapped to a series of knowledge graphs. A semantic graph may be generated for each primitive type (e.g., a graph for text, a graph for visual / image relationships, a graph for tables, etc.). These knowledge graphs may be interconnected either through direct graph relationships or the AST. For example, if a document contains a chart image and a paragraph describing it, the AST may provide the structural link between the image and paragraph while a knowledge graph may define the relationship between the entities in the chart and the text.
[0129] By converting unstructured documentary and file data into graph-linked ASTs, an Al custom agent accomplishes more than simply “reading” text. It navigates a structured map of concepts. This may allow for significantly more accurate cross-referencing and high-fidelity recommendations because the agent may understand the relationship between data points across different formats, not merely the proximity of these data points within a file.
[0130] To generate a response output to a user query at run-time, an Al custom agent may traverse a semantic graph (AST with knowledge graphs) created based on attributes and contents of a user query, in accordance with business rules indicated as purpose in one embodiment. The Al custom agent may use these structured relationships to synthesize a coherent answer that cites specific nodes or data points across unstructured data sources. This facilitates grounding of the response output in the purpose and validation of the response against the structural logic of the AST.
[0131] FIGs. 2A-2F illustrate a routine 200a for analyzing data in accordance with one embodiment, including subroutine 200b, subroutine 200c, subroutine 200d, subroutine 200e, and subroutine 200f, which may be performed as part of routine 200a in some embodiments, and which are described in greater detail below. In some embodiments, routine 200a and its subroutines 200b-200f may be performed alongside additional steps or subroutines such as are described with respect to the exemplary process for adaptive document creation and workflow optimization 600 of FIG. 6.
[0132] Routine 200a and its subroutines 200b-200f may be performed by variations of the systems disclosed herein, including the exemplary system architecture 300 of FIG. 3, the exemplary system architecture 400 of FIG. 4, the exemplary system architecture 500 of FIG. 5,Docket No. FSP2504PCTand agents described with respect to the exemplary process for adaptive document creation and workflow optimization 600 of FIG. 6. Such systems may be computing systems such as the personal computing device 1402, personal computing device 1714, mobile computing device 1716, and cloud computing node 2900 later described. Such systems may include one or more processors and memory storing instructions that, when executed by the one or more processors, cause the system to perform routine 200a and its subroutines 200b-200f.
[0133] FIG. 2A illustrates an example routine 200a for generating a response to a user query 1002. Although the example routine 200a and its subroutines 200b-200f depict a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of routine 200a and its subroutines 200b-200f. In other examples, different components of an example device or system that implements routine 200a and its subroutines 200b-200f may perform functions at substantially the same time or in a specific sequence.
[0134] According to some examples, the method includes receiving a user input comprising an indication of the data to be analyzed and at least one request related to the data at block 202. For example, the input evaluator agent 606 illustrated in FIG. 6 may receive a user input comprising an indication of the data to be analyzed and at least one request related to the data.
[0135] According to some examples, the method includes accessing the data in available data silos at block 204. For example, the input evaluator agent 606 illustrated in FIG. 6 may access the data in available data silos.
[0136] According to some examples, the method includes determining whether the request is at least partially satisfied by a cached result previously generated using at least one of a plurality of language models at block 206. For example, the input evaluator agent 606 illustrated in FIG. 6 may determine whether the request is at least partially satisfied by a cached result previously generated using at least one of a plurality of language models.
[0137] According to some examples, if the request is not at least partially satisfied by a cached result at decision block 208, the method includes receiving, at a response generator agent, the user input and the data at block 210.
[0138] According to some examples, the method includes selecting, by the response generator agent, a set of language models from the plurality of language models at block 212. In one embodiment, selecting the set of language models may comprise evaluating at least one ofDocket No. FSP2504PCTcomputational power, memory, storage, and network bandwidth of the computing system. In one embodiment, selecting the set of language models may comprise selecting the set of language models based on at least one of capability, latency, cost, and uptime. In one embodiment, selecting the set of language models may comprise determining a complexity of the at least one request and selecting the set of language models based on the complexity of the at least one request.
[0139] According to some examples, the method includes provisioning the set of language models as artificial intelligence (Al) custom agents to analyze the data in view of the user input at block 214. According to some examples, the method includes calling and performing subroutine 200f to provision the Al custom agents. In one embodiment, the system may store at least one Al custom agent in a cache. In one embodiment, the system may validate at least one Al custom agent. Validating the at least one Al custom agent may comprise comparing the at least one Al custom agent with a cached Al custom agent.
[0140] According to some examples, the method includes analyzing the data with the Al custom agents to generate a response to the at least one request at block 216. In one embodiment, the response may comprise code. In one embodiment, the response may be stored in cache as a cached result for future retrieval by the system. In various embodiments, routine 200a may then call either or both subroutines 200b-200c described below, in any order, to analyze data, and / or may call subroutine 200d to evaluate the resulting response. In some embodiments, routine 200a may then call and perform subroutine 200e, described below, to provide output to a user.
[0141] According to some examples, if the request is at least partially satisfied by a cached result at decision block 208, the method includes using the cached result in generating a response to the user request at block 218. In some embodiments, routine 200a may then call and perform subroutine 200e, described below, to provide output to a user.
[0142] FIG. 2B illustrates an example subroutine 200b for analyzing the data with the Al custom agents, which may be performed as part of routine 200a in some embodiments.
[0143] According to some examples, the method includes performing a high-level analysis to generate the response, at subroutine block 220. For example, the Al custom agent 308 illustrated in FIG. 3 may perform a high-level analysis to generate the response. The response may include at least one high-level recommendation.Docket No. FSP2504PCT
[0144] According to some examples, the method includes presenting the at least one high-level recommendation to a user at subroutine block 222. For example, the output agent 638 illustrated in FIG. 6 may presenting the at least one high-level recommendation to a user.
[0145] According to some examples, the method includes performing a detailed analysis to generate at least one detailed response at subroutine block 224. For example, the Al custom agent 308 illustrated in FIG. 3 may perform a detailed analysis to generate at least one detailed response. The at least one detailed response may include detailed recommendations related to the at least one high-level recommendation.
[0146] According to some examples, the method includes on condition the user requests details related to the at least one high-level recommendation, presenting the at least one detailed response to the user at subroutine block 226. For example, the output agent 638 illustrated in FIG. 6 may on condition the user requests details related to the at least one high-level recommendation, presenting the at least one detailed response to the user. FIG. 1 illustrates an AST-based data processing approach 100 which may be accomplished through the performance of subroutine 200b.
[0147] FIG. 2C illustrates an example subroutine 200c for analyzing the data with the Al custom agents, which may be performed as part of routine 200a in some embodiments.
[0148] According to some examples, the method includes generating an AST representation of unstructured data from the available data silos at subroutine block 228. For example, the AST parser 316 illustrated in FIG. 3 may generate an AST representation of unstructured data from the available data silos. .
[0149] According to some examples, the method includes generating knowledge graphs based on semantic relationships among the unstructured data at subroutine block 230. For example, the AST parser 316 illustrated in FIG. 3 may generate knowledge graphs based on semantic relationships among the unstructured data. Generating the knowledge graphs may include generating a semantic graph for each primitive type of the unstructured data, wherein the primitive type is at least one of text, visual relationships, image relationships, and tables.
[0150] According to some examples, the method includes mapping the AST to the knowledge graphs at subroutine block 232. For example, the AST parser 316 illustrated in FIG. 3 may map the AST to the knowledge graphs.
[0151] FIG. 2D illustrates an example subroutine 200d for evaluating a response generated by the Al custom agents, which may be performed as part of routine 200a in some embodiments.Docket No. FSP2504PCT
[0152] According to some examples, subroutine 200d includes determining that the response complies with business rules, business constraints, and the data indicated by the user input at subroutine block 234. For example, the methodology evaluator agent 634 illustrated in FIG. 6 may determine that the response complies with business rules, business constraints, and the data indicated by the user input.
[0153] According to some examples, subroutine 200d includes validating the response against structural logic, i.e., an AST representation knowledge graphs, or the AST mapped to the knowledge graphs at subroutine block 236. Additional validations may be performed in various embodiments, such as are described with respect to the exemplary process for adaptive document creation and workflow optimization 600 of FIG. 6.
[0154] FIG. 2E illustrates an example subroutine 200e for providing output to a user in response to user input to the system, which may be performed as part of routine 200a in some embodiments.
[0155] According to some examples, subroutine 200e includes receiving the response to the at least one request from the Al custom agents at subroutine block 238. For example, the output agent 638 illustrated in FIG. 6 may receive the response to the at least one request from the Al custom agents.
[0156] According to some examples, subroutine 200e includes generating an output comprising the response to the at least one request at subroutine block 240. For example, the output agent 638 illustrated in FIG. 6 may generate an output comprising the response to the at least one request. The output may comprise at least one of a document, a presentation, text data, an image, a video, audio data, embedded content, and web content.
[0157] FIG. 2F illustrates an example subroutine 200f for provisioning an Al custom agent from a set of language models, which may be performed as part of routine 200a in some embodiments.
[0158] According to some examples, subroutine 200f includes populating at least one Al custom agent with context based on the user input at subroutine block 242. Context may include business rules, business constraints, business logic, and data pertinent to the user input. These may in one embodiment be identified during a purpose phase as previously described.
[0159] According to some examples, subroutine 200f includes assigning at least one tool to the at least one Al custom agent, the at least one tool allowing the at least one Al custom agent to interact with the available data silos at subroutine block 244.Docket No. FSP2504PCT
[0160] According to some examples, subroutine 200f includes assigning a language model of the set of language models as a reasoning model of the at least one Al custom agent at subroutine block 246.
[0161] According to some examples, subroutine 200f includes instructing the reasoning model of the at least one Al custom agent on how to interpret the data within the populated context based on the user input at subroutine block 248.
[0162] FIG. 3 illustrates an exemplary system architecture 300 in accordance with one embodiment. The exemplary system architecture 300 may comprise an adaptive document creation and workflow optimization system 302, its operational components, and its data sources and output targets. These may include procedures 304, rules agent 306, Al custom agents 308, base models 310, a back office 312, integration agent 314, AST parser 316, data silos 318, an Observation agent 320, a Decision agent 322, an action agent 324, office productivity software data 326, the real world 328, feedback 330 data, as described below.
[0163] Procedures 304 may be provided as inputs to the adaptive document creation and workflow optimization system 302. Procedures 304 may include interactive interfaces, text documents, audio data, image data, video data, etc. Procedures 304 may inform rules agent 306, such as by indicating rules, processes, objectives, standard operating procedures (SOPs), etc. In one embodiment, the rules agent 306 may be informed during a purpose phase 102 as described with respect to the AST-based data processing approach 100 of FIG. 1.
[0164] The rules agent 306 may be provided to Al custom agents 308, which may be provided on private platforms, edge platforms, SAFE platforms, etc. The Al custom agents 308 may be dynamically created Al custom agents that are configured to support a specific business or client company “verticals” in leveraging the adaptive document creation and workflow optimization system 302. Such specialized agents may be referred to as “vertical agents” and may receive certain trainings based on their vertical. The Al custom agents 308 may also receive data from base models 310. These may be open Al models, cloud-based models, industry-standard models, etc. The Al custom agents 308 may provide data to a back office 312 element including an integration agent 314. The back office 312 may include data warehouses, enterprise applications, databases, and other data silo 318.
[0165] The integration agent 314 may be used to consolidate all of the data available in the data silo 318 into a data pipeline accessible from the back office 312. The integration agent 314 may be a high-level Al custom agent configured specifically to parse data across data silos byDocket No. FSP2504PCTdeploying many specialized sub-agents. The integration agent 314 may incorporate an AST parser 316 that indexes and summarizes unstructured data as ASTs with knowledge graphs as described with respect to the AST-based data processing approach 100 in FIG. 1. Note that any of the high-level agents of this disclosure, as well as the Al custom agents 308 disclosed herein, may include an AST parser 316 element by which they may develop and analyze data using an AST-based data processing approach 100. In one embodiment, when integration agent 314 faces a query with data sources across scattered files, it may instruct its subagent swarm including its AST parser 316 to begin the process of document ingestion, optical parsing, document primitive decomposition (extracting text, tables, figures etc.) to form one or more ASTs, then the subagent swarm may perform entity extraction and relationship building, dual graph construction, semantic linking with ASTs, then knowledge retrieval, summarization, reranking against goals, etc. Finally, the subagent swarm may output delivery of the resulting processed data to integration agent 314.
[0166] The Al custom agents 308 and integration agent 314 may provide all of the data at their disposal to a number of interfaces included in the system. These may include the observation agent 320 which is configured to recognize observation agent in the data provided, the decision agent 322 which is configured to persuade users to make recommended decisions, and the action agent 324, which is configured to perform actions either autonomously or based on user decisions. In one embodiment, the observation agent 320 may provide data on detected observation agent to the Decision agent 322 to inform its recommended decisions. The decision agent 322 may in turn provide those decisions to the action agent 324 to inform the actions it may take or may be directed by a user to take.
[0167] The interfaces provided above may output their respective products to office productivity software data 326 to facilitate the recording, communication, and action upon their outputs. Office productivity software data 326 may include Microsoft Outlook, Excel, Word, etc., image and video capture programs, Google Office tools, databases, etc.
[0168] Actions decided upon or directed toward the action agent 324 may be communicated to the real world 328. The real world 328 may include customers, operators, suppliers, etc. Real world 328 entities may provide feedback 330 to the exemplary system architecture 300. This feedback 330 may be used to adapt procedures 304 and update back office 312 data sources.
[0169] In one embodiment, an Al model based on Granit 3.1 8B may be trained as a retail sales model that calculates unit sales and revenue and provides recommendations based on userDocket No. FSP2504PCTqueries. Such a model may be run locally on a personal computing device at over 67 tps. In this manner, a powerful model may be implemented on a private system, and may be fast and economical.
[0170] FIG. 4 illustrates an exemplary system architecture 400 in accordance with one embodiment. The exemplary system architecture 400 comprises a number of elements similar to or functionally the same as elements introduced with respect to FIG. 3. These may include Rules agent 306, Al custom agents 308, base models 310, a back office 312, an Integration agent 314, data silos 318, an Observation agent 320, a Decision agent 322, an action agent 324, office productivity software data 326, and the real world 328. The adaptive document creation and workflow optimization system 402 of the exemplary system architecture 400 may receive user input 404. The adaptive document creation and workflow optimization system 402 may receive user input 404 from developers and IT professionals or end users. The adaptive document creation and workflow optimization system 402 may include a process extractor agent 406 and a Process generator agent 408. The real world 328 may present data silos 410 for use in addition to the data silos 318 of the back office 312. Data silos may include enterprise resource planning (ERP) data, customer resource management (CRM) data, supply chain management (SCM) data, human resource management (HRM) data, data lake (DL) data, data warehouse (DW) data, and other data types well understood by one of ordinary skill in the art.
[0171] Within the exemplary system architecture 400, the rules agent 306 may be an Al agent configured to determine applicable business rules, business constraints, and other context based on the user input 404 and regulatory or other controlling documents included in data silos 318 and data silos 410. The rules agent 306 may be configured to distill business rules, etc., from available data, into natural language rules that may be used to direct or inform the operation of the adaptive document creation and workflow optimization system 402 upon a user input 404 in a manner that is agnostic to model or code used to implement the disclosed solution.
[0172] Within the exemplary system architecture 400, the integration agent 314 may act as the central powerhouse for the system by integrating data from all available data silos. The integration agent 314 may use a SAFE Al platform to provide reliability, speed, and security. The integration agent 314 may provide a data integration layer that readily extends existing documentary infrastructure for a business with new processes at any time. In one embodiment, where data needed to support response to a user's query involves large amounts of visual data of a particular place, item, or event, the integration agent 314 may be directed to integrate thisDocket No. FSP2504PCTvisual data into a four-dimensional model by which the place, item, event, etc., may be explored and analyzed in all spatial dimensions and across a period of time. Visual recognition Al may be used to analyze the integrated visual model to track inventory across a warehouse in one embodiment.
[0173] In one embodiment, the Integration agent 314 may use commercially and publicly available resources such as LibraryAI, GraphAI, KernelAI, InterfaceAI, etc. GraphAI, for example, may allow the system to create a description of an agentic workflow in something more like natural language than code. This may allow a “template” of sorts for an Al custom agent to be stored economically in cache and implemented using any number of coding languages, making such agents portable across multiple implementations. For example, Al custom agent may be defined as follows:N ME : Chat Explain BriefDATABASE : <optional specific data source>TUNING :Hour roleYou are a data science assi stant expert in direct-to-consumer (DTC) retail fashion#The situation and the goalA user in DTC retail fashion has made a request about business data . They want a clear, accurate , and compelling response . They should understand the response and how it was arrived at . You will generate a response to their request based solely on input data . You may use the suppl ied metadata and background information to explain the input data you receive .Hour task<USER QUERY>MODEL : <optional des ired run-time model>
[0174] Such a definition or template for an agent may be agnostic with regard to the model it may be generated by, the data silos it references, and may even allow a number of user queries to be processed by the same operating logic used to embody the Al custom agent created toDocket No. FSP2504PCTthese specifications. Templates like these also allow a layering of abstractions and execution. Base models 310 of any foundation may be used to create the Al custom agent so defined on a local system private to an enterprise, using natural language rules. These templates may be expanded and aggregated to allow related Al custom agents to work as collections, swarms, or workflows of agents. This is further described with reference to FIG. 25.
[0175] Within the exemplary system architecture 400, the Observation agent 320 may combine available data from within a business with industry analysis. In this manner, the Observation agent 320 may detect metrics that may support insightful responses to user queries.
[0176] Within the exemplary system architecture 400, the Decision agent 322 may filter and analyze available data and business decisions based on that data, suitable to a user's needs as expressed in a user query. It may generate office productivity software files providing persuasive support for its recommended decisions. In this manner, it may support users in making important decisions for steering a business quickly.
[0177] Within the exemplary system architecture 400, the action agent 324 may monitor key performance indicators based on the industry vertical it is deployed in and automate implementation of processes that will help achieve them. It may provide recommended changes in office productivity software files.
[0178] Within the exemplary system architecture 400, the process extractor agent 406 may analyze available documents within back office 312 data silos 318 and real world 328 data silos 410 and may automatically extract business processes. In one embodiment, the process extractor agent 406 may summarize and categorize processes as associated with certain data categories or query types, making these easily accessible at run-time. Process data so distilled by the process extractor agent 406 may be available to the other interfaces of the exemplary system architecture 400 for use in completion of their specific tasks. The process extractor agent 406 may convert the documentary data it extracts into business rules that may be used to design or inform the work of Al custom agents. In this manner, a recommendation description may become reusable by the decision agent 322, and may be easily shared, revised, and re-run with different data. Such business rules may be revised by human users using natural language text.
[0179] Within the exemplary system architecture 400, the process generator agent 408 may allow the system to build workflows for integrated data in a matter of days, rather than years. The process generator agent 408 may work as a real-time filter leveraging adaptive processingDocket No. FSP2504PCTto analyze existing data on a very large scale efficiently, and may be able to generate summaries for documented processes and recommend opportunities to streamline documented business process workflows. In one embodiment, the process generator agent 408 may summarize meeting records and distill this information into action items and improved business processes.
[0180] FIG. 5 illustrates an exemplary system architecture 500 in accordance with one embodiment. The exemplary system architecture 500 may comprise an adaptive document creation and workflow optimization system 502 in communication with and receiving data from a back office 312, providing output to a real world 328, and able to generate office productivity software data 326 as one form of output, as previously described.
[0181] The adaptive document creation and workflow optimization system 502 in this streamlined configuration may comprise the Rules agent 306 provided to the Al custom agents 308, which are also in communication with base models 310. The Al custom agents 308 may interact with the Integration agent 314, process extractor agent 406, and Process generator agent 408, introduced above.
[0182] FIG. 6 illustrates an exemplary process for adaptive document creation and workflow optimization 600 in accordance with one embodiment. The exemplary process for adaptive document creation and workflow optimization 600 may be performed by the disclosed system, as illustrated in the exemplary system architecture 300 of FIG. 3, as well as the exemplary system architecture 400 and exemplary system architecture 500 of FIG. 4 and FIG. 5 respectively. Portions of the exemplary process for adaptive document creation and workflow optimization 600 may be performed as described with respect to routine 200a and its subroutines 200b-200f introduced in FIGs. 2A-2F.
[0183] User input 602 may be a user query (e.g., a request, a command, prompt, question, etc.). Computer data 604 may include data that may be used to address the user query. For example, user input 602 may be a prompt that identifies data to be used in addressing the user query (e.g., “analyze previous years’ sales data to predict sales of Product X in 2025”), and computer data 604 may include the previous years’ sales data. Computer data 604 may be received from the user who provided user input 602. For example, the user may upload computer data 604 at the same time that the user provides user input 602. Alternatively, computer data 604 may be received in response to the user input 602. For instance, the computing system may be configured to query a local storage or the cloud for computer dataDocket No. FSP2504PCT604 pertaining to user input 602. User input 602 may be provided by a computer system such as a personal computing device of the user or business. This may be a computer system / server 2902 such as is described with respect to FIG. 29 as a cloud computing node 2900, though the computer system / server 2902 may also operate independent of interaction with cloud-available resources in a cloud computing environment.
[0184] User input 602 and computer data 604 may be evaluated in order to make decisions on processing needs. For example, the level of complexity of user input 602 and the most suitable processing resources for addressing user input 602 may be evaluated. Evaluating the inputs may include determining whether user input 602 may be addressed using cached data, which may reduce redundant computations and minimize resource utilization. Cached data may include cached responses and cached code. Cached code may in one embodiment include one or more cached Al custom agents. Whether the system retrieves a cached response or cached code may depend on user input 602 and computer data 604.
[0185] The inputs to the system (e.g., user input 602 and computer data 604) may be evaluated by input evaluator agent 606. Input evaluator agent 606 may transform user input 602 into a suitable format for being provided to a set of language models. Transforming the format of the user input 602 may include parsing the user query to break the user query down into a plurality of components, such as identifiers, keywords, operators, separators, and constants. Input evaluator agent 606 may reformulate the user input 602 by organizing the plurality of components of the user input 602 into a different format. For example, input evaluator agent 606 may reformulate the plurality of components into a format that is more easily processed by a language model. Input evaluator agent 606 may then provide the re-formulated user input 602 to response generator agent 610. Alternatively, input evaluator agent 606 may provide the user input 602 to response generator agent 610 in its original format.
[0186] Input evaluator agent 606 may analyze user input 602 to determine whether an answer to the user input 602 exists in a cache 608. Input evaluator agent 606 may include one or more models (e.g., language models) for analyzing user input 602 to determine whether response to the user input 602 exists in cache 608. If a satisfactory response to the user input 602 exists in cache 608, the answer may be retrieved and provided to an output agent 638 without engaging response generator agent 610 to generate a response from scratch. Input evaluator agent 606 may analyze the semantic meaning of a user input 602 in order to determine whether it may be satisfied by a cached result. The input evaluator agent 606 may determine whether the userDocket No. FSP2504PCTinput 602 may be satisfied by a cached result by analyzing semantic relationships between the user input 602 and previously received user input 602 stored in cache 608.
[0187] If the semantic analysis indicates that the user input 602 semantically matches a previously received user input 602 or that the semantic similarity between the user input 602 and a previously received user input 602 is within an acceptable similarity range, input evaluator agent 606 may retrieve the corresponding previously generated response to the previously received user input 602 from cache 608 and provide the retrieved response to output agent 638, which may provide the response to the user. The previously generated response may be large, and cache 608 may in one embodiment store a compressed version of the previously generated response that includes enough information to reconstruct the previously generated response. For instance, cache 608 may store a hash value representing the previously generated response. Input evaluator agent 606 may retrieve and decode the hash value representing the previously generated response. In one embodiment, cache 608 may store previously generated responses in the form of ASTs such as that illustrated in FIG. 1.
[0188] In one embodiment, cache 608 may store code used for generating responses to previous user input 602. In one embodiment, the code may include one or more Al custom agents created to respond to similar user input 602. Input evaluator agent 606 may retrieve the previously generated code and provide the code to output agent 638, which may run the code and output a response to the user. In addition, cache 608 may store data such as computer data 604 used for generating responses to previous user input 602 and / or pathways to data sources used for generating these responses.
[0189] If input evaluator agent 606 determines that there is a “hit” (i.e., a cached answer or cached code may be used to address user input 602), the cached answer or cached code may be retrieved from a cache 608. The cached answer may be output by an output agent 638, or cached code may be used to process computer data 604 and the answer may be subsequently output by output agent 638. Alternatively, if there is a “miss,” i.e., a cached answer or cached code cannot be used to address user input 602, a response to user input 602 may be generated by response generator agent 610. The response generated by response generator agent 610 may be an optimized output tailored to the specific use case identified in user input 602.
[0190] Response generator agent 610 may include a code requirement checker agent 612 configured to determine whether code needs to be generated or whether an existing tool may be called (e.g., via a Model Context Protocol (MCP) server) in order to address user input 602. InDocket No. FSP2504PCTsome examples, code requirement checker agent 612 may determine that code does not need to be generated and that user input 602 may be addressed without writing code (e.g., if user input 602 is a purely mathematical question). Code requirement checker agent 612 may then provide the user input 602 to model farm diffuser agent 616 to generate a response to user input 602. Alternatively, code requirement checker agent 612 may determine that code needs to be generated, and user input 602 may be provided to code evaluator agent 614. Code evaluator agent 614 may determine whether code that may be used to address user input 602 already exists in a cache. If code evaluator agent 614 determines that such code does not exist, user input 602 may be provided to model farm diffuser agent 616 to begin the code generation process. If code evaluator agent 614 determines that code already exists in a cache, the response generator agent 610 may forego the code generation process and retrieve the cached code from cache 628. Cache 628 may be the same cache as cache 608 or may be a separate cache. In some examples, this determination may be duplicative of the determinations made by input evaluator agent 606 and may be skipped.
[0191] At model farm diffuser agent 616, the system orchestrates the selection and deployment of a set of Al models (e.g., language models 618) to address user input 602. The model farm diffuser agent 616 may select a set of language models by optimizing for parameters such as model capability, latency, cost, and uptime to meet the demands of user input 602. The set of language models 618 may be run in parallel as an ensemble in order to generate outputs based on user input 602. While FIG. 6 illustrates three language models 618, it should be understood that FIG. 6 is merely exemplary, and any number of language models may be used. Optionally, the model farm diffuser may also use input from a human (e.g., an expert) and / or other types of models or code to generate outputs based on user input 602 in addition to the set of language models 618. The model farm diffuser may include a performance guarantee management system to ensure that performance metrics such as response time, accuracy, and uptime for the selected set of language models 618 adhere to predefined service-level agreements (SLAs).
[0192] The set of language models 618 selected by model farm diffuser agent 616 may be characterized by varying levels of performance and resource needs. For example, model farm diffuser agent 616 may select one or more small models, which are models that may be run locally or with minimal resources and which may offer faster responses at the expense of accuracy. Model farm diffuser agent 616 may select one or more medium models, which areDocket No. FSP2504PCTmodels that may be deployed on departmental servers and that may provide greater accuracy but be slower than small models. Model farm diffuser agent 616 may select one or more large models, which are models that are highly accurate but resource-intensive, requiring significant hardware and time, and may not always return results within a feasible timeframe. Model farm diffuser agent 616 may select one or more experimental models, which are newly developed models that have not been fully characterized and thus do not have guaranteed response times and may lack the ability to provide meaningful input.
[0193] Model farm diffuser agent 616 may select any number of language models and any combination of sizes of language models to be included in the set of language models. The sizes of the language models and the numbers of each size selected for use in the set of language models may be based on the complexity of user input 602, the desired accuracy of the response, and / or the acceptable cost (in terms of time, money, and / or resource utilization). For example, if user input 602 is simple, model farm diffuser agent 616 may select small models. Optionally, the outputs generated by the small models may be occasionally spot-checked (e.g., 1% of the time). If user input 602 is medium difficulty, model farm diffuser agent 616 may select small and medium models. Outputs may be occasionally spot-checked (e.g., 1% of the time). If user input 602 is complex, model farm diffuser agent 616 may select medium and large models to ensure greater accuracy. The table below illustrates exemplary compositions of various sized sets of language models and procedures for running the respective sets of language models.
[0194] Model farm diffuser agent 616 may identify new models that may be selected to be included in a given set of models on a constant, periodic, or occasional basis. For example, even when a user input has not been received, model farm diffuser agent 616 may run in theDocket No. FSP2504PCTbackground to evaluate new models by identifying potential candidate models and determining whether they are suitable for use in a set of language models
[0195] User input 602 may be processed by the set of language models 618 selected by model farm diffuser agent 616. An execution waiter agent 620 may manage response time expectations. Execution waiter agent 620 may determine how to proceed based on predetermined guidelines (e.g., business rules provided by a user). For example, a user may specify that a response to user input 602 is desired within five seconds. Execution waiter agent 620 may determine that some outputs (e.g., outputs from small models) will be available within that time frame. Execution waiter agent 620 may decide to provide the available outputs to downstream processing functionalities (e.g., splicing agent 621 or consensus evaluator agent 622) in order to ensure that this stipulation is met, even though not all models have finished generating outputs. Execution waiter agent 620 may be or may include an Al model (e.g., a language model or equivalent generative Al model that weighs factors such as response time, accuracy, and cost).
[0196] Execution waiter agent 620 may also provide delay signals for complex inputs. For example, for a complex query, execution waiter agent 620 may generate a message to be provided to a user such as: "This is a challenging problem and may take some time. May I take a second or two to get back to you?" This mechanism manages expectations while maintaining a conversational tone. Execution waiter agent 620 may provide outputs from the set of language models 618 to splicing agent 621. Splicing agent 621 may aggregate the received outputs from the set of language models 618 and splice the outputs together to create a cohesive and contextually accurate response. All outputs from the set of models may be spliced together at the same time. Alternatively, if the outputs from the set of models are provided in piecemeal fashion (e.g., if some models are faster than others), some outputs may be processed and used to generate a first response to user input 602, which may then be updated as additional model outputs are received. For example, small models may be used to generate an initial response. A larger model may subsequently provide more accurate or insightful data. Splicing agent 621 may include a language model that is configured to integrate the improved response into the existing output. The updated response may incorporate transitions naturally, using phrases such as, "Hmm, on further thought, it might actually be this instead" to create a conversational experience that feels organic and thoughtful. This is further described with respect to FIG. 11.Docket No. FSP2504PCT
[0197] A consensus evaluator agent 622 may evaluate the outputs from the set of language models 618 provided by execution waiter agent 620 or a spliced response from splicing agent 621 to determine whether there is consensus between the outputs (or whether the spliced response represents a consensus answer to the user query). Determining whether there is consensus between the outputs or between different portions of a spliced response may include using a lightweight Al model (e.g., a language model having a low cost and high speed) to compare the outputs or portions. The outputs or portions may be compared to determine whether the semantic meanings of the outputs are the same or similar. The criteria for determining whether an appropriate level of consensus exists may depend on the desired accuracy of the response to user input 602. For example, if desired accuracy is relatively low, the criteria for determining whether consensus exists may be whether a majority of outputs or portions are in semantic agreement. Whether outputs or portions are in semantic agreement may be determined by comparing a semantic distance between outputs or portions to a predetermined threshold semantic distance. If desired accuracy is high (e.g., if user input 602 is a request to make a medical diagnosis), the criteria for determining whether a consensus exists may be whether all outputs are in semantic agreement.
[0198] The outputs from the set of language models may be discarded if an acceptable level of consensus is not achieved. New outputs may be generated by selecting a new set of language models (e.g., using updated model selection criteria) and running the new set of models on the same input, or by re-running the same set of language models with a re- worked input.Alternatively, consensus evaluator agent 622 may provide a message to splicing agent 621 that the outputs do not represent an acceptable level of consensus, and splicing agent 621 may incorporate one or more additional outputs from one or more additional models into an updated response, if additional outputs are available. Splicing agent 621 may then provide the updated response to consensus evaluator agent 622 to reevaluate consensus. Consensus evaluator agent 622 may provide the outputs to quality evaluator agent 624 if an acceptable level of consensus is achieved.
[0199] Quality evaluator agent 624 may determine the quality of the outputs. The quality of the outputs may be based on user input 602 (e.g., the complexity of user input 602), the quality of the set of language models 618 used, and / or the quality of the data used to generate the outputs. The quality determination may be performed by an Al model (e.g., a language model)Docket No. FSP2504PCTthat is slower and more resource-intensive but more accurate than the Al model used by consensus evaluator agent 622.
[0200] The outputs from the set of language models 618 may be discarded if an acceptable quality level is not achieved. New outputs may be generated by selecting a new set of language models (e.g., using updated model selection criteria) and running the new set of models on the same input, or by re-running the same set of language models with a re-worked input. The outputs may be provided to code checker agent 626 if an acceptable quality level is achieved.
[0201] A code checker agent 626 may determine whether a generated response includes code that should be pushed into a cache 628. If code checker agent 626 determines that the generated response includes newly generated code, the code may be pushed into the cache 628 and subsequently executed by a code executor agent 630. If code checker agent 626 determines that the generated response does not include code, code checker agent 626 may simply provide the generated response to an output processor 632.
[0202] Output processor 632 may receive the outputs from the set of language models 618, the outputs from splicing agent 621, and / or the outputs from code checker agent 626 and / or code executor agent 630 and process the outputs to generate a response to user input 602. The response may be the outputs from the set of language models and / or the code or may be based on the outputs and / or the code (e.g., a spliced response based on the outputs and / or the code).
[0203] The response to user input 602 generated by output processor 632 may undergo one or more additional checks before being output to a user. For example, methodology evaluator agent 634 may perform a holistic evaluation of the methods used to generate the response (e.g., how the models were selected, the manner in which they were executed, the manner in which the outputs generated by the models were spliced, the ground truth sources that the outputs were compared to, etc.). Methodology evaluator agent 634 may also generate an output that explains how the response was derived. The output from methodology evaluator agent 634 may include an audit log, which may be useful to end users, regulators, or other quality assurance personnel in evaluating the reliability and accuracy of a response to a user query. Alternatively or in addition, skeptic agent 636 may compare the response to a cached output or trained SAFE output structured artefacts using a trained Al model.
[0204] The response may then be output to a user by output agent 638. The response may be provided in any number of formats. For example, the response may include a document, chart, table, presentation, photo, video, audio, web content, or a combination thereof.Docket No. FSP2504PCT
[0205] FIG. 7 illustrates an exemplary pipeline for adaptive document creation and workflow optimization 700 in accordance with one embodiment. The exemplary pipeline for adaptive document creation and workflow optimization 700 may be implemented for the disclosed system, as illustrated in the exemplary system architecture 300 of FIG. 3, as well as the exemplary system architecture 400 and exemplary system architecture 500 of FIG. 4 and FIG. 5 respectively.
[0206] A user input 702 may be provided through a user chat interface 704 in one embodiment. This user input 702, along with cached base data 706, may be provided to a find and filter stage 708. This may identify and filter relevant data from the available sources, streamlining the pipeline to focus on high-value information. The analyze and process stage 710 may then apply analytical tools and algorithms to process the filtered data, extracting insights, observation agent, and actionable results. These may be sent to the recommend and report stage 712, which may generate tailored recommendations or reports based on analyzed data, presenting clear and concise outputs for export 714 that align with user needs and objectives.
[0207] The main characteristic of this architecture is in the input evaluation described with respect to the input evaluator agent 606 of FIG. 6. This takes the input from the user and dynamically decides during runtime where it should return a cached result or not. If the system finds a pre-existing generated result that satisfies the query by a (configurable) factor of at least 98% in one embodiment, it will draw the result from the cache. Otherwise it will generate a new result and store this result along with associated metadata in the cache for future use.
[0208] The user chat interface 704 and code generation 718 may be subject to configurable business rules 716, which may allow client companies to configure their own desired safeguards. When needed code generation 718 may be performed in support of the find and filter stage 708, analyze and process stage 710, and recommend and report stage 712, and antihallucination agents 720 may be in place to ensure that these stages do not introduce hallucination errors into the outputs for export 714.Workflow Diagram
[0209] FIG. 8 illustrates an exemplary process flow for decision-making and operational efficiency 800 in accordance with one embodiment. The exemplary process flow for decisionmaking and operational efficiency 800 may be performed by the disclosed system, as illustratedDocket No. FSP2504PCTin the exemplary system architecture 300 of FIG. 3, as well as the exemplary system architecture 400 and exemplary system architecture 500 of FIG. 4 and FIG. 5 respectively. The diagram outlines a process flow for decision-making and operational efficiency, emphasizing structured rules and outputs across multiple interconnected stages. Each block represents a functional component of the system, linked by rules and structured outputs, to create a streamlined feedback loop. Here's a detailed breakdown:Backoffice Pipelines
[0210] The process begins with the back office pipelines 802, which represent the foundational workflows and systems that handle operational data. These pipelines collect, preprocess, and organize raw data, providing a robust input for downstream analysis. These pipelines may collect data from the data silos 318 of the back office 312.Observe Observation agent
[0211] Data from the back office pipelines feeds into the observe observation agent 804 stage. At this point, observation agent, trends, and anomalies are detected and analyzed to gain actionable insights. The insights derived here establish the foundation for rules and decisionmaking. This may be performed by the Observation agent 320 of FIG. 3.Persuade People to Decide
[0212] Insights from the observed observation agent are channeled into the persuade people to decide 806 stage. This stage involves using data-driven arguments and contextual rules to influence decision-makers. It bridges the gap between raw data insights and human action, ensuring alignment with organizational goals. This may be performed by the Decision agent 322 of FIG. 3.Act Decisively, React Quickly
[0213] Decisions made by individuals or teams lead to the perform 808 stage. This stage focuses on translating decisions into rapid and decisive actions. The ability to act and react swiftly is essential for maintaining operational efficiency and adaptability in dynamic environments. This may be performed by the action agent 324 of FIG. 3.Interaction with the Real WorldDocket No. FSP2504PCT
[0214] The real world 328 block interacts directly with the perform 808 stage to reflect realtime outcomes and feedback. External factors, such as market conditions, customer interactions, or operational constraints, provide key feedback that informs future decisions.Integration of People and Office Productivity Software
[0215] People 810 and office productivity software data 326 are key nodes that ensure the integration of human effort and organizational processes. Rules derived from earlier stages are applied here, guiding both individual behavior and standardized office workflows.Distilling Standard Operating Procedures (SOPs)
[0216] Outputs from the people 810 and office productivity software data 326 stages are refined into SOPs at the distill the standard operating procedures 812 stage. SOPs codify best practices, ensuring consistency and efficiency in future operations. This may be performed by the Process generator agent 408 in one embodiment.Feedback Loops and Rules
[0217] Dashed lines labeled "Rules 814" indicate feedback loops within the system. These loops help enforce governance, adjust workflows, and ensure that operational guidelines remain aligned with real-world dynamics.Structured Output
[0218] Solid lines represent Structured output 816, showing how processed data and decisions are transformed into actionable results, ensuring clarity and usability.
[0219] FIG. 9 illustrates an exemplary distributed processing system 900 in accordance with one embodiment. The exemplary distributed processing system 900 comprises an end device 902 including an application 904 having application programming interface or API calls 906, code 908, and data access 910, and performing agent LM calls 912, and hardware 914such as a processor 916, a disk 918, and an Al (PC, mobile, onboard) 920 local model. The exemplary distributed processing system 900 further includes an enterprise environment 922 with a department device 924 and a corporate data center 926. The exemplary distributed processing system 900 finally includes cloud services 928 such as a web-based representational state transformer 930, interference as a service 932, GPUs as a service 934, and proprietary LMsDocket No. FSP2504PCT
[0220] The application 904 may send API calls 906 to the department device 924, which may be a Mac mini or similar computing device. The API calls 906 may also be sent to the webbased representational state transformer 930. Code 908 of the application 904 may be run on the processor 916. Data access 910 may be supported by data stored on the disk 918. Code 908 run on the processor using data from the disk 918 may provide data to the department device 924. Agent LM calls 912 by the application 904 may query the Al (PC, mobile, onboard) 920 or the proprietary LMs 936. The Al (PC, mobile, onboard) 920 may further interact with the department device 924.
[0221] The department device 924 may send all data it receives and / or generates to the corporate data center 926. The corporate data center 926 may interact with interference as a service 932, GPUs as a service 934, and proprietary LMs 936 available as cloud services 928.
[0222] FIG. 10 illustrates an exemplary process flow for handling user queries 1000 in accordance with one embodiment. The exemplary process flow for handling user queries 1000 may be performed using the exemplary system architecture 300 of FIG. 3, the exemplary system architecture 400 of FIG. 4, or the exemplary system architecture 500 of FIG. 5, and various embodiments of the solution system disclosed.
[0223] The back office 312 may provide data and a user may supply user queries 1002. An integrated agent framework 1010 may supply Al custom agents to process and respond to the user query 1002. The user query 1002 may be evaluated to determine whether the user query 1002 may be addressed using a cached result or cached code. Evaluating user query 1002 may include determining whether a user query matching user query 1002 has been previously received and answered. If the same or a similar query was previously received and answered, the evaluation may include determining which data was used to answer the previous query. The data used to answer the previous query may be the same data that would be needed to answer user query 1002 or may be different data (e.g., older data).
[0224] If user query 1002 matches a previous query, and the data used to answer the previous query is the same data that would be used to answer user query 1002, then a cached answer may be retrieved at block 1004. If user query 1002 matches a previous query but the data that would be used to answer user query 1002 is different than the data that was used to answer the previous query, then cached code may be retrieved at block 1006. The cached code may be code that was used to process old data to answer the previous query. The cached code may then be used to process the new data needed to address user query 1002. If user query 1002 does notDocket No. FSP2504PCTmatch a previous user query, then the system may forego retrieving an answer or code from the cache. At block 1008 a process (e.g., portions of the exemplary process for adaptive document creation and workflow optimization 600 of FIG. 6) may be executed in order to generate a customized response to user query 1002. This decision-making process is summarized in the table below.
[0225] The system may invoke the integrated agent framework 1010 to provide the response by developing and executing new code using the new data. The integrated agent framework 1010 may utilize open, cloud-based, or SAFE Al models. The system may provide output in forms accessible through office productivity software data 326, collectively known in some cases as the front office. This jury and caching process may provide a high reliability in addition to improved response time. Reducing entirely new Al operations reduced the chance for hallucinations to occur. In one embodiment, the determinations made in blocks 1004-1008 may be evaluated by the integrated agent framework 1010 in order to adapt and improve the responsiveness and accuracy of the integrated agent framework 1010.
[0226] FIG. 11 illustrates an exemplary process flow for splicing model outputs 1100 in accordance with one embodiment. The scaled Al kernel 1102 of the disclosed system may have access to models of varying sizes, speeds, and accuracies. Light, fast models 1104 often do not have the most comprehensive training, and so may provide the quickest but not always the best answers. Medium models 1106 may provide improved accuracy, but may take more time to respond. The best, most comprehensively trained models may be highly accurate, but may also be very large and slow models 1108.
[0227] In splicing the outputs from these various models, the disclosed system may be able to respond quickly, but may also be able to improve upon its response given time. For example, the scaled Al kernel 1102 may provide an initial response 1110 from one of its fast models 1104, but may indicate in its output that it is still working to make sure a better answer isn't available. If a better answer is determined by, for example, one of its medium models 1106, the scaled Al kernel 1102 may provide a follow up response 1112. As the medium models 1106Docket No. FSP2504PCTand slow models 1108 continue to run, additional follow up responses 1114 may continue to be output.
[0228] In order to manage user expectations and provide staged responses in an organic way that feels comfortable, conversational, and trustworthy to a human user, the disclosed system may include splicing phrases 1116 as it presents such a staged response. For example, upon providing the initial response 1110, the system may include a message such as, “Hmm, difficult question, please be patient.” This may be followed up by a swift initial response 1110, which may in one embodiment be phrased as provisional. When a later, better follow up response 1112 is available, the system may include a splicing phrase 1116 such as, “Upon further reflection,” followed by the follow up response 1112. When an even better response is determined, the system may include, “Actually, I think I have a better idea," as a splicing phrase 1116 with the follow up response 1114.
[0229] The scaled Al kernel 1102 may provide the advantages of dynamic response time 1118, dynamic accuracy level 1120, dynamic compute cost 1122, and extensibility to other metrics 1124.
[0230] The scaled Al kernel 1102 may be implemented using accurate, secure, private, responsive, economical, and reliable SAFE Al platforms in the disclosed system. It may be locally-focused, deployed in an application or a web browser interface. It may provide aggressive caching of results, data, and code for easy of reuse in the workflows previously described. Much of the functionality may be available on a local personal computing (PC) device, such as the end device 902 of FIG. 9 and the personal computing device 1402 of FIG.14, while allowing higher performance loads to be supported by resources at the enterprise or cloud level, such as the enterprise environment 922 and cloud services 928 of FIG. 9 or the CEO adapter 1414 or developer adapter 1416 of FIG. 14.
[0231] An exemplary implementation of a kernel implemented on a local stack (at the edge) is provided in FIG. 15. An implementation of a kernel in the cloud may be seen in FIG. 16.
[0232] FIG. 12 illustrates a method 1200 in accordance with one embodiment. Method 1200 may be performed by the disclosed system having exemplary system architecture 300, exemplary system architecture 400, and exemplary system architecture 500, as described with respect to FIGs. 3-5, and variations thereof as disclosed herein. In method 1200, some steps may be combined, changed, or omitted. In some embodiments, additional steps may be performed in combination with the method 1200. Accordingly, the operations illustrated (andDocket No. FSP2504PCTdescribed in greater detail below) are exemplary by nature and, as such, should not be viewed as limiting.
[0233] At step 1202, a candidate model is identified. The candidate model may be identified by a computing system (e.g., the computer system / server 2902 comprising the model selector 2826 described below or the model farm diffuser agent 616 introduced in FIG. 6). The candidate model may be a language model. The candidate model may be identified from any number of sources. For example, the candidate model may be identified by searching various Internet sources (e.g., Arxiv, GitHub, huggingface, etc.) for open-source models. Alternatively, the candidate model may be a model that is stored locally on the computing system or stored remotely (e.g., in the cloud or on another system) and accessed by the computing system. The candidate model may be a small model (e.g., a model than may run locally or with minimal resource needs), a medium model (e.g., a model that may have greater accuracy than a small model but needs greater resource utilization than a small model), a large model (e.g., a model that may have greater accuracy than a medium model but needs greater resource utilization than a medium model), or an experimental model (e.g., a model of any size that is still in development and has not yet been characterized in terms of resource utilization and / or accuracy).
[0234] At optional step 1204, the candidate model is optionally screened for safety. For example, the candidate model may be screened to verify that the code does not contain any viruses. If the candidate model contains a virus, the candidate model may be removed from consideration, and the rest of method 1200 may be skipped for that particular candidate model.
[0235] At step 1206, the candidate model is provided with a plurality of test questions having known answers. The plurality of test questions may include at least 50, 100, 200, 500, 1,000, 5,000, 10,000, or more test questions. The content and number of questions may vary based on the candidate model being screened. For example, the test questions provided to a candidate model that is a large model may be different than the test questions provided to a candidate model that is a small model. The plurality of test questions may also vary in difficulty. The test questions may be provided to the candidate model in stages. For example, a first set of test questions may be low difficulty, a second set of test questions may be medium difficulty, and the third set of test questions may be high difficulty.
[0236] At step 1208, the responses to the plurality of test questions generated by the candidate model are compared with the known answers. Comparing the responses generated by theDocket No. FSP2504PCTcandidate model with the known answers may include determining a semantic distance between the responses generated by the candidate model and the known answers. If the semantic distance is within a predetermined range for a particular response and known answer, the response may be determined to match the known answer and may be deemed satisfactory. If the semantic distance between the response and known answer is outside of the predetermined range, the response may be determined not to match the known answer and may be deemed unsatisfactory. The number of responses generated by the candidate model that match known answers may be calculated, and the candidate model may continue through the evaluation pipeline to optional step 1210 if the number of responses matching known answers is greater than or equal to a predetermined number. Alternatively, the candidate model may be directly added to a pool of “production stage” models, such as the production models 1314 of production stage 1306, introduced in FIG. 13, which are models that may be selected to constitute the set of models used to answer a given user query, without undergoing further evaluation. The candidate model may be removed from consideration, and the rest of method 1200 may be skipped if too many responses are deemed unsatisfactory (e.g., if the number of responses matching known answers is below the predetermined number of responses).
[0237] Optionally, the candidate model may be evaluated based on its response to a previously received user query. At optional step 1210, the candidate model is provided with a previously received user query. The previously received user query may be a user query that was previously addressed by the disclosed system using a set of models. The candidate model may generate a response to the previously generated user query.
[0238] At optional step 1212, the response to the previously received user query generated by the candidate model is compared with a previously generated answer to the previously received user query. Comparing the response generated by the candidate model with the previously generated answer may share one or more similarities with step 1208. Comparing the response generated by the candidate model with the previously generated answer may also include measuring the compute time, cost, and / or accuracy of the candidate model in generating the response and comparing it to the compute time, cost, and / or accuracy of the set of models used to generate the previously generated answer.
[0239] Optionally, the candidate model may be used in “shadow mode” to generate a response to a new user query. At optional step 1214, the candidate model is provided with a new user query. The new user query may be a user query that does not have a previously generatedDocket No. FSP2504PCTanswer. The candidate model may generate a response to the new user query. A set of models including models other than the candidate model may also generate a response to the new user query.
[0240] At optional step 1216, the response to the new user query generated by the candidate model may be compared with the response to the new user query generated by the set of models. Comparing the response generated by the candidate model with the response generated by the set of models may share one or more similarities with step 1208 or optional step 1212. For example, comparing the response generated by the candidate model with the response generated by the set of models may include determining a semantic distance between the responses. If the response generated by the candidate model is within a predetermined semantic distance of the response generated by the set of models, then the candidate model may be deemed to be a “production stage” model and added to the pool of models from which the set of models used to answer a given user query is selected.
[0241] The pool of production stage models may include models of various capabilities and sizes (e.g., models having about 100 million parameters to about 15 trillion parameters). For example, as shown in FIG. 13, the pool of production stage models may include a plurality of small models 1318, a plurality of medium models 1320, and a plurality of large models 1322. The pool of production stage models may include any number of models. Models may lose their production stage status at any time. For example, a given model may be retired from the production stage if a new model that has superior cost, accuracy, and / or latency has entered the production stage. Models retired from the production stage are referred to as “obsolete stage” or “veteran stage” models, such as the obsolete models 1316 of obsolete stage 1308 shown in FIG. 13. Veteran stage models may be recalled into production if needed. Veteran stage models may also be used to track model drift. Veteran stage models may be used to periodically answer newly received user queries, and their responses may be compared to the responses generated by the set of models used to answer the query. Comparing the responses may indicate whether the system has drifted over time.
[0242] Optionally, models in the production stage may be fine-tuned. The models may be fine-tuned based on user feedback on the responses generated by the system. Models may be removed from the production stage after fine-tuning and re-evaluated to determine whether the fine-tuned model should be returned to the production stage, for example by repeating method 1200.Docket No. FSP2504PCT
[0243] FIG. 13 illustrates stages of an exemplary model evaluation process 1300. The process for evaluating new models, such as the method 1200 illustrated in FIG. 12, may be conducted following the stages of an exemplary model evaluation process 1300. The stages of an exemplary model evaluation process 1300 may include a recruit stage 1302, a candidate stage 1304, a production stage 1306, and an obsolete stage 1308.
[0244] At the recruit stage 1302, the model farm diffuser system may sweep various sources to identify “recruit models 1310,” i.e., models that may be useful for answering user queries. This may be done by web search of Arxiv, by looking for new models in GitHub, huggingface, etc. Recruit models 1310 may be subjected to a screening process that involves determining if the recruit models 1310 are SAFE. For example, the recruit models 1310 may be screened to determine whether the loader code contains a virus. Each of the recruit models 1310 may then be run through a “boot camp” to test the model's capabilities against a set corpus of questions. The corpus of questions may be provided in substages of increasing difficulty. For example, a recruit model may be provided with easy questions first, so that the recruit model may be discarded if it is incapable of answering the easy questions.
[0245] At the candidate stage 1304, the initial set of recruit models 1310 has been winnowed down to a set of candidate models 1312 that have demonstrated satisfactory performance in response to the set corpus of questions in the recruit stage. The candidate models 1312 may be tested against old workloads (e.g., against previously received user queries). The candidate models 1312 may also be operated in “shadow mode” to judge their performance (i.e., the candidate models 1312 may be tested against a new user query, but their responses may not actually be used in generating the response that is output to a user). The model farm diffuser system may characterize the accuracy, cost, and response time of the candidate models 1312. Candidate models 1312 may remain in candidate stage 1304 for any amount of time. As with recruit stage 1302, models may be tested in substages throughout candidate stage 1304. For instance, candidate models 1312 may start with easy problems (e.g., by processing relatively simple previously received user queries) and eventually graduate to being operated in shadow mode and processing new queries.
[0246] Candidate models 1312 that are deemed to be qualified to use in a set of language models such as the language models 618 previously described may graduate to the production stage 1306. Production models 1314 in production stage 1306 may be selected to be used in a set of language models to answer a newly received user query. Models in production stageDocket No. FSP2504PCT1306 may have a wide variety of characteristics. For example, production models 1314 may include models ranging from 100M to 15T parameters in size so that models are available to answer queries having varying degrees of complexity. Production stage 1306 may include a plurality of small models 1318, a plurality of medium models 1320, and a plurality of large models 1322.
[0247] Models may be withdrawn from production stage 1306 when new candidate models 1312 are shown to be superior in cost, speed, and / or accuracy. Such obsolete models 1316 that have been withdrawn may be retired to obsolete stage 1308 (also referred to as the “veteran stage”). Obsolete stage 1308 may include older models. Obsolete models 1316 may be used to track model drift, so they may occasionally be brought in to process a new query to see if the system has "drifted" (e.g., to determine whether the system is making different decisions today than it was 10 years ago). Obsolete models 1316 may be recalled to production stage 1306 as needed.
[0248] FIG. 14 illustrates an exemplary use case-specific adapter 1400 in accordance with one embodiment. The exemplary use case-specific adapter 1400 may include a personal computing device 1402 accessing cached code 1404, an Al model 1406, a language model 1408, a language model 1410, and a factuality model 1412. These elements of the personal computing device 1402 may communicate to a CEO adapter 1414 and a developer adapter 1416 interfacing regarding a generated query 1418.
[0249] The personal computing device 1402 may be a commercially available desktop or laptop computer. It may be a professional-grade machine with fast processing speeds and adequate memory and storage to support the elements indicated here. In one embodiment, the cached code 1404 may be Rust, C, or Python code. The Al model 1406 may be a Gema2-8B model, which may need 9GB of RAM to be available and may need processing speeds supporting 27 transactions per second (tps). The language model 1408 may be a Llama 3.2-3B model needing 1GB of RAM and 61 tps. The language model 1410 may be a Reader-LM-1.58 model needing 27 tps. The factuality model 1412 may be a Bespoke Minicheck model needing 5GB and 36 tps. These examples are not intended to limit the present disclosure, but are provided for exemplary purposes. One of ordinary skill in the art will readily apprehend that other models having different system needs may be used.
[0250] The CEO adapter 1414 and developer adapter 1416 comprise intermediate layer adapter logic supporting flexible and dynamic interfaces for different user types. The CEODocket No. FSP2504PCTadapter 1414 and developer adapter 1416 may be 200MB adapters. The CEO adapter 1414 may allow a business executive to ask a question or pose a query, such as the generated query 1418, which may, for example, be, “Sales, step by step, ending inventory, analyze.” The developer adapter 1416 may receive the generated query 1418 in one embodiment, and a developer or programmer may use the developer adapter 1416 to direct the disclosed system to create an Al custom agent to respond to the query. These 200MB adapters may be large and expensive. Alternatively, the generated query 1418 may be addressed by the cached code 1404, Al model 1406, language model 1408, and other elements available in a much lighter weight and economical form on a personal computing device 1402.Edge Hardware Execution Environment
[0251] FIG. 15 illustrates an exemplary kernel deployment at the edge 1500 in accordance with one embodiment. FIG. 15 represents a system architecture that integrates various components for managing LMs, embeddings, user interactions, and knowledge databases, within a system core 1502 while ensuring modularity and flexibility in using both cloud and on-premise solutions. Below is a detailed breakdown of each component and its role:System core 1502 Components
[0252] Decision agent 1504.- Rust or JavaScript may be used in various embodiments. Acts as the front-end layer where users interact with the system. Built with a combination of Rust (for performance) and JavaScript (for interactivity). Facilitates communication with the user through a user chat database 1506 and ensures that data flows to / from the GraphAI system.
[0253] Natural Language Processor or NLP server 1508: Javascript may be used in various embodiments. The central orchestrator and decision-making logic of the system. Connects with multiple components like the knowledge database 1512, code sandbox 1514, embedding LM 1510, and external LMs. Handles input / output flows and logic for interpreting user input and delivering results.
[0254] Embedding LM 1510: A specialized model for generating vector embeddings of text data. Outputs embeddings are stored in a knowledge database 1512 such as a retrieval-augmented generation (RAG) Knowledge Database for retrieval-augmented generation.
[0255] Knowledge database 1512: A RAG Knowledge Database and PostgreSQL may be used in various embodiments. A knowledge base designed for fast retrieval of embeddings and related data. Supports RAG, enabling contextual responses.Docket No. FSP2504PCT
[0256] Code sandbox 1514: Ray may be used in various embodiments. A distributed execution environment for running code snippets or tasks. Likely used to dynamically evaluate or test user inputs in real-time.
[0257] Model server 1516: Ollama Server (Llama.cpp) may be used in various embodiments. Serves on-premise LMs, allowing the system to use open-source models and ensuring data privacy. Receives updates and model improvements via ollama pull from external sources.External Components
[0258] Prompt repository 1518: This may be a Git-backed Prompts Repository hidden from the user in various embodiments. A version-controlled repository that stores prompt templates and configurations. Allows dynamic updates to prompt engineering workflows without direct user interaction. Updates are fetched using git pull.
[0259] Cloud LM providers 1520: These may be Anthropic, OpenAI, etc., in various embodiments. Interfaces with external cloud-based LMs through REST APIs. Adds flexibility to use state-of-the-art language models when on-premise resources are insufficient or specific capabilities are needed.
[0260] Open source and specialized models 1522: This may be GGUF, HuggingFace, etc., in various embodiments. Supports a wide range of open-source models that may be fine-tuned for specific tasks. Includes models like GGUF, HuggingFace fine-tuned versions, and internal specialized models, offering adaptability for diverse use cases.Data Flows
[0261] User chat database 1506: This may use PostgreSQL in various embodiments. Logs and stores user interactions, enabling contextual conversations and session history. Works in tandem with GraphAI to provide responses based on prior chats.
[0262] Embedding LM 1510: Embeddings created by the embedding LM may be stored in a RAG Knowledge Database in various embodiments. Vector embeddings created by the Embedding LM are stored in the database for efficient retrieval.
[0263] NLP server 1508: This may be GraphAI in various embodiments. NLP service utilizes REST APIs to query cloud LMs and return outputs to the user. Communicates with the model server to query locally hosted LMs for tasks requiring on-premise computation. The NLP service may send code or tasks to the Code Sandbox for evaluation or execution.Docket No. FSP2504PCT
[0264] Decision agent 1504: Acts as the primary interaction point, with real-time updates and feedback to the user. This system architecture demonstrates a hybrid approach, combining cloud and local resources to optimize LM usage. It supports flexibility in deployment, security through on-premise models, and scalability with cloud integration. By modularizing components like the Decision agent, NLP Service, and Knowledge Database, it ensures robust, efficient, and user-centric operations.Cloud Stack - identical to Edge Stack Easy to Remote
[0265] FIG. 16 illustrates an exemplary kernel deployment in the cloud 1600 in accordance with one embodiment. FIG. 16 represents a modular and scalable system architecture having a system core 1602 interacting with external components and managing Al-driven workflows, including natural language processing, vector database searches, and code execution, with support for both cloud-based and on-premise solutions. The exemplary kernel deployment in the cloud 1600 may be run on cloud solutions such as AWS, Azure, GCP, or a K8s / S3 abstraction. The following describes a detailed breakdown:System core 1602 Components
[0266] Orchestration layer 1604.- Node.js or Kubernetes (K8s) may be used in various embodiments for scalability. Acts as the primary agent framework for handling user requests and delegating tasks. Connects with the Open Web UI server 1606 and user interface server 1610.
[0267] Web UI server 1606: This may be an open Web UI server, Docker or Pip, in various embodiments. Provides a user-friendly web interface for interacting with the system. Deploys via Docker for portability and uses Python (pip) for dependency management. Integrates with slashgpt.yaml to configure workflows for code, search, and database operations, as represented by data, code, and records 1612.
[0268] User interface 1608: A higher-level user interface specifically designed for the disclosed workflows. Focuses on advanced operations and managing projects within the disclosed system and related systems.
[0269] User interface server 1610: This may be K8s in various embodiments. Backend server supporting the User Interface. Deploys on Kubernetes for managing workloads at scale.Synchronizes configurations with a repository such as a Git-backed repository using git pull.Docket No. FSP2504PCT
[0270] LM server 1614.- This may be A 100 or K8s in various embodiments. Serves large language models (LMs) on high-performance GPUs (e.g., NVIDIA A100). Offers APIs (e.g., vLM API) to run inference tasks on LMs efficiently. Supports models deployed behind a firewall for privacy and security.External Connections
[0271] Existing Agent frameworks and code 1616: Supports integration with pre-built frameworks like LangChain for advanced agent-based workflows. Configured through staticFi.yaml in one embodiment for seamless interoperability.
[0272] Version controlled data repository 1618 / This may be a Git-backed repository synced to S3-Compatible Store in various embodiments. A version-controlled repository for managing system configurations and templates. May automatically synchronize updates using Git and S3-compatible storage solutions or similar.
[0273] Open source and specialized models 1620: Compatible with open-source models like GGUF and HuggingFace fine-tuned models (45K+). May includes specialized, solutionspecific models deployed behind firewalls for added security.Data Flow and Configuration Files
[0274] Configuration files may be used to configure interactions between the orchestration layer 1604 and the user interface server 1610. GraphAI.yaml files may be used for this in various embodiments. Configuration files may be used to handle workflows involving the web UI server and client requests. Slashgpt.yaml files may be used for this in various embodiments involving open web UI servers. Configuration files may be used to specify external agent frameworks and integrations. StaticFi.yaml files may specify external agent frameworks and EangChain integrations in various embodiments.
[0275] Client requests (such as code execution, searches, VectorDB queries, file handling, and other data, code, and records 1612 data) may be routed via the Open Web UI server 1606 to the EM server 1614 or other components. EM tasks may be processed on the EM Server and results may be routed back to the client.
[0276] FIG. 17 illustrates an Al hosting edge-cloud configurations 1700 in accordance with one embodiment. The edge and cloud kernel deployments described in FIG. 15 and FIG. 16 may be implemented jointly according to such a hybrid configuration. A large cloud EM 1702 may be available providing 8-16 H100 or similar high performance processors and 640- 1.2TBDocket No. FSP2504PCThigh bit rate memory. Such a large cloud LM 1702 may cost eight to sixteen million dollars. It may include a prompt cache 1704 and embeddings 1706 available by the model without roundtrip times incurred. It may have network latencies of 0.25ms, a 256-buffer bus, and context reread and contention capabilities. It may be accessible via web browsers 1708 and python code 1710 programs that include appropriate API calls. Using such a system 100% of the time may have a number of benefits, b but may be costly with regard to response time.
[0277] The presently disclosed solution may take advantage of the capabilities of such a large cloud LM 1702 while handling a number of tasks and queries on personal computing devices 1714 and mobile computing devices 1716. Personal computing devices 1714 may in one embodiment be 95% locally served by 4x3-8B small language models performant with 16GB of memory. Mobile computing devices 1716 may be 90% locally served using 4B small language models able to run with 8GB memory. In one embodiment, personal computing devices 1714 and mobile computing devices 1716 may achieve a 99% hit rate, handling almost all queries at the edge, while connecting to a commodity server 1712 in order to outsource the remaining 1% of queries to a large cloud LM 1702.
[0278] FIG. 18 illustrates an exemplary use case 1800 in accordance with one embodiment.Problem: "I Need to Integrate All My Data Silos"
[0279] For large scale projects working with a wide variety of data sources, typical cost to analyze and build an economic impact or other study may be around ~$250K+ and may take months. Thus the information may be out of date before the report is completed. Users want the reports that are taking months at high costs to be built in days, to be available for review live and on-demand, and to give me current and accurate insights.
[0280] Solution: Integration agent 314 Integrates Data silos 318
[0281] Integration agent 314 may provide integration of multiple data sources and business rules across data silos 318 such as enterprise resource planning 1804, customer relationship management 1806, and office productivity software files 1808. Integration agent 314 may dynamically create reports and provide insights and recommendations around key goal 1802a regarding economic sectors, key goal 1802b regarding job growth, and key goal 1802c regarding social insights that are generated quickly by the models described herein.
[0282] In providing this solution, the disclosed system may use a reliable SAFE Al platform. An original goal for the system may be that 75% of responses are correct on ad hoc questionsDocket No. FSP2504PCTand consistent when asked 4 times. Accuracy may be acceptable at 96% with uncorrected data, and 100% with corrected data. Consistency may be 97%, with 0% of responses being wrong if the question is vague, such as, “What do customers want in December?” As the system improves, a goal may be for the system to be 100% accurate in its responses if the question is present in training data, such that data is available and the rules are known.
[0283] FIG. 19 illustrates an exemplary use case 1900 in accordance with one embodiment.Problem: "Know Your Customer Process is Very Inefficient"
[0284] There are many existing systems to track banking customers, but the data may be in disparate silos and may be hard to integrate Having a real time and secure search capability may be crucial to improving business processes. Users want real time reporting without manual data viewing and entry across multiple systems.Solution: Process generator agent 408 Delivers Automated Real-Time Know Your Customer (KYC) Support
[0285] Process generator agent 408 may deliver the desired support around the key goal 1902a of obtaining a comprehensive customer list, key goal 1902b of merging customer data into the list from disparate sources, and key goal 1902c, or outputting citation data. To do so, process generator agent 408 may analyze data across data silos 318 such as search results 1904, internal data 1906, and office productivity software files 1908. Integration of internal, external providers, and real-time web search, simplified user interface with batch mode available, and automatic report and recommendation generation (tuned by organization's rules) may be supported by this solution.
[0286] FIG. 20 illustrates an exemplary use case 2000 in accordance with one embodiment.Problem: "We Need Live Stocking Recommendations to Drive EBIT"
[0287] Supply chain executives are seeing unprecedented disruption. Up to 30% SKU's are obsolete every quarter. Gathering the data to make decisions may take weeks of copy and paste without including time to think or act. Users need reports for review daily that have key recommendations, which may be queried, enabling immediate action, with instant, companywide communication.Solution: Decision agent 322 Generates RecommendationsDocket No. FSP2504PCT
[0288] Decision agent 322 may deliver desired outputs around the key goal 2002a of reducing stockout danger, key goal 2002b of minimizing silent big hits, and key goal 2002c of identifying weeding out underperformers. Decision agent 322 may do so by analyzing data from data silos 318 such as enterprise resource planning 2004, customer relationship management 2006, and email 2008. In this manner, decision agent 322 may convert memos into a structured documents. Decision agent 322 may integrate business rules from the company's supply chain experts. Decision agent 322 may thus provide automatic report and recommendation generation that may export to email.
[0289] FIG. 21 illustrates an exemplary use case 2100 in accordance with one embodiment.Problem: "Customer Satisfaction Surveys Are Too Late"
[0290] Customer experience managers don't want surveys, they want customer sentiments monitored 100% of the time, with actionable opportunities. Very few customers reach out when unhappy. They just leave, thus businesses never get a touchpoint. A new way is needed to propose phrases to chatbots or humans to use with customers, to improve the experience for unhappy customers based on actual sentiments at the time, live and on demand. Users know higher customer satisfaction leads to higher lifetime value with lower customer cost, but it's hard to manage it live.Solution: Process generator agent 408 Drives Business with Continuous Scoring
[0291] Process generator agent 408 may provide continuous scoring, leveraging available data to meet key goal 2102a of predicting customer sentiment, key goal 2102b of providing advice, and key goal 2102c of updating sentiment. Process generator agent 408 may do so using data from data silos 318 such as customer resource management 2104, call data 2106, and email 2108.
[0292] Process generator agent 408 may extract data from transcripts and assign a default sentiment to new customers. Process generator agent 408 may analyze transcript for sentiment and update sentiment measures in the database. In addition, process generator agent 408 may generate a call list for appointments to call with an estimated time of arrival and issues. A weather event may cause broken "promises". Process generator agent 408 may predict! vely send text with pre-customer tuned fixes and predict new sentiment. An installer may make preinstall calls and analyze sentiment. The installer may perform an install and note any issues thatDocket No. FSP2504PCTcame up. Installer arrival and departure times may be logged. Post-install follow-up call transcripts may be analyzed for sentiment, which may be updated in a customer databaseSolution: Process generator agent 408 Integrated Customer Experience Process
[0293] In one embodiment, process generator agent 408 may perform initial customer call analysis and predict sentiments on a 1-10 scale During conversations, sentiment may be reevaluated. When sentiment drop below 2, then "make goods" may be offered, such as gift cards or free shipping. Customer response may be dynamically analyzed in real time through live transcription and audio analysis, and new sentiments may be generated.
[0294] FIG. 22 illustrates an exemplary use case 2200 in accordance with one embodiment.Problem: "How Do We Automate Manual Sustainability Reporting?"
[0295] Sustainability mandates are coming in rapidly in the European Union and soon globally. Huge manual labor may be needed as a result to copy and paste pertinent sustainability data from Excel, emails, and existing enterprise systems into mandated reports. There is currently no time to think, let alone evaluate the information. It takes users days to complete such tasks, and a solution is needed that automates the needed reports, providing them in a matter of hours and leaving enough time to evaluate the data.Solution: Process generator agent 408 Builds Sustainability
[0296] Process generator agent 408 may build sustainability reports quickly by leveraging available data to meet key goal 2202a of matching a desired input template, key goal 2202b of transforming or translating existing data into other languages as needed, and key goal 2202c of generating multiple output formats using the same input data. Process generator agent 408 may do so using data from data silos 318 such as spreadsheet data 2204, enterprise resource planning 2206, and email 2208. In this manner, the process generator agent 408 may examine a report data library of existing spreadsheets, CSV files, and emails, and, following business rules for transforming the data into the mandated reports, generate these reports automatically and privately, according to a desired template and in a desired language.
[0297] FIG. 23 illustrates an exemplary use case 2300 in accordance with one embodiment.Problem: "Hard to Achieve Your Sustainability Objectives"Docket No. FSP2504PCT
[0298] For sustainability executives, the data is important, but "what does it mean" and "what are its effects" are the primary points of interest. Instead of writing Excel macros or waiting for IT users want to generate reports and use those reports quickly to understand an effective business strategy. In interacting with sustainability data, users want to learn what to do next and what is missing in their business.Solution: Decision agent 322 builds Sustainability Strategy
[0299] The decision agent 322 may build a sustainability strategy for the user, leveraging available data to meet key goal 2302a of understanding the sustainability of materials and the certifications needed, key goal 2302b of understanding the sustainability of packaging types, and key goal 2302c of recognizing lagging categories. Decision agent 322 may do so using data from data silo 318 such as output reports 2304, input data 2306, and industry comparisons 2308. In this manner, the decision agent 322 may integrate with existing data merged from multiple outputs. It may provide an automated dictionary with explanations, ad hoc tables, and certification percentages for integration into recommendations provided by decision agent 322.
[0300] FIG. 24 illustrates an exemplary use case 2400 in accordance with one embodiment.Problem: "Are We Missing Great Investments?"
[0301] Tracking investments is notoriously hard in the venture capital community. Integration across data silos is vital. Under conventional workflows, much manual processing of information is needed, and the businesses are changing all the time.Solution: Process generator agent 408 Investment Memo Process
[0302] The process generator agent 408 may develop and automate an investment memo process for a user, leveraging available data to meet key goal 2402a of ingesting data on existing deals, key goal 2402b of scoring known deals, and key goal 2402c of generating recommended next steps, process generator agent 408 may do so using data from data silos 318 such as the internet 2404, a deal database 2406, and a market activity database such as Crunchbase 2408. In this manner, process generator agent 408 may integrate data from internal and external providers with real-time web searches. A specialized Al model tuned for investments may be utilized, along with easy-to-change investment parameters determined through a rules agent, optional searches by an observation agent, and copy from PowerPoint. InDocket No. FSP2504PCTthis manner, the solution may provide rapid evaluation of thousands of deals within a single afternoon.
[0303] FIG. 25 illustrates an autonomous agent growth through natural language templates 2500 in accordance with one embodiment. Generally speaking, increasingly sophisticated Al agents are created through increasingly complex defining structures such as detailed schematic designs that are encoded in cumbersome programmatic code blocks.
[0304] A basic flow design 2502 may provide rules to design a fairly straightforward Al agent. It may be implemented using a one shot (lx) solution, such as Qwen2.S, GPT-4, and similar platforms. More complex tasks may be represented by a branched flow design 2504, which may be used to encode the more capable Al agents needed to complete them. Thee may be suitable for implementation with a Search (lOx) solution, QWC, ol-mini, and similar models.
[0305] Direct acyclic graphs 2506 may be the next step, defining a number of agent states and decision and action paths leading between them. These may allow a number of models to be aggregated, such as may be seen for DAG (10-100x) solutions, such as Llamaindex, Langchain, etc.
[0306] The disclosed solution, however, supports the swift dynamic creation of numerous agents in an Al custom agent team designs 2508 through the interaction of the interfaces disclosed herein and the distillation of data into templates that economically define numerous flexible and reliable agents that may be trained, improved, and aggregated to perform increasingly complex tasks. In one embodiment, an autonomous team of lOO-lOOOx agents may be developed effectively, analogous to solutions provided by AutoGPT, Swarm, Devin, Skynet, and similar platforms.
[0307] FIG. 26 illustrates an exemplary query and output 2600 in accordance with one embodiment. A number of input photos 2602 may be provided of people wearing items clothing. A request 2604 may be provided by a user indicating that the user likes the fashions in the input photos 2602 and would like recommendations for similar clothing at their favorite shop, the inventory data 2606 of which is available to the adaptive document creation and workflow optimization system 302 online. Alternatively, a seamstress or textile craftsperson may be looking for inspiration for new designs. The adaptive document creation and workflow optimization system 302 may analyze the input photo 2602 and return the output fashion suggestion images 2608 appropriate to what the user has requested.Docket No. FSP2504PCT
[0308] In one embodiment, the user may be a retail clothing company, and may wish to provide a targeted ad to a known social media user. The company may provide the input photos 2602 including the target person 2610 and auxiliary persons 2612 and 2614, which it has obtained from the social media accounts of the target person 2610. This user may also provide its inventory data 2606, and include in the request that the adaptive document creation and workflow optimization system 302 develop a collage image of items available at locations near the target person 2610 for use in a targeted ad.
[0309] FIG. 27 illustrates an exemplary query and output 2700 in accordance with one embodiment. In one embodiment, a request 2702 may be made of the adaptive document creation and workflow optimization system 302 in the form of a user query as follows. “I need to do a presentation on the Seattle development for the Arts Council. Use the format from the Portland one. Update the video with local artists. Include sustainability analysis.” The user may provide image data 2704 and text copy 2708. The “Portland one” may be provided as a link to a desired web content layout sample 2706, or may be discoverable or known by the adaptive document creation and workflow optimization system 302 from previous queries. The adaptive document creation and workflow optimization system 302 may return as its output the code to build output web content 2710 fulfilling the stipulations of the request 2702.
[0310] FIG. 28 illustrates an Al system 2800 in accordance with one embodiment. The Al system 2800 may comprise a sensor data 2802, a text data 2804, a signal classifier 2810, a tokenizer 2818, a prompt composer 2822, an Al model 2824, an Al output 2828, a database 2812, a user 2806, and a model selector 2826. The Al system 2800 may in one embodiment be implemented on a single computing device such as the computer system / server 2902 described in greater detail below. In other embodiments, the elements of the Al system 2800 may be distributed across cloud computing nodes 2900 interconnected in a cloud computing environment 3000 as described with respect to FIG. 30.
[0311] Input to the Al system 2800 may be in the form of sensor data 2802, text data 2804, user 2806, and other forms of data, as will be readily understood by one of ordinary skill in the art. The sensor data 2802 may be provided by sensors providing digital or analog readings. These sensors may be connected in a computing environment and may thus provide output sensor data 2802 for storage in various forms of computer memory and for processing by computer processors. Text data 2804 may be provided as input from various sources, including human users 2806 operating a text entry peripheral attached to a computational device, such asDocket No. FSP2504PCTa keyboard, a microphone with its audio output processes by speech-to-text algorithms, etc. Text data 2804 may also include historical data stored in various memory structures. Users 2806 may also provide input to the Al system 2800 in the form of user selection data 2808 through a computational input / output (I / O) interface such as a mouse, keyboard, microphone, touchscreen, etc. This description is not intended to be limiting, and additional sources of sensor data 2802, text data 2804, and user selection data 2808 may readily occur to one of ordinary skill in the art.
[0312] Sensor data 2802 may be sent to a signal classifier 2810 in order to analyze and interpret the data provided by sensors. In one embodiment, the signal classifier 2810 may act as a specialized tokenizer trained or otherwise programed to recognize and interpret sensor signals and tokenize sensor data 2802 such that the phenomena detected by the sensors and recorded in the sensor data 2802 may be communicated in a manner interpretable by other elements of a computational system, such as the components of the Al system 2800 described here. A database 2812 may include stored data 2814 which may be made available to the signal classifier 2810 in order to classify current text sensor data 2802 quickly based on known classification so similar signals based on prior knowledge, programmed relationships, or Al system 2800 learnings. In one embodiment, the signal classifier 2810 may provide such signal classifications 2816 directly to the prompt composer 2822. In another embodiment, signal classifications 2816 may be sent for further tokenization by the tokenizer 2818. text data 2804 and user selection data 2808 from the user 2806 may also be sent to the tokenizer 2818. The tokenizer 2818 may develop input tokens 2820 from the various streams of input data as is described for the exemplary tokenizer 3200 illustrated in FIG. 32. The input tokens 2820 may be provided as input to the prompt composer 2822.
[0313] The prompt composer 2822 may receive the input tokens 2820 developed from the sensor data 2802, the text data 2804, the user selection data 2808, and other sources of data available to the Al system 2800 not pictured here. The prompt composer 2822 may also receive input in the form of stored data 2814 from one or more databases 2812. Using these inputs, the prompt composer 2822 may construct a single token, a series of tokens, a body of text, and / or a series of conditional or unconditional commands suitable to use as a prompt 2830 to one or more connected Al models 2824. For example, a series such as “conditional on command A success, send command B, else send command C” may be built and sent all at once given a specific data precondition, rather than being built and sent separately.Docket No. FSP2504PCT
[0314] The prompt composer 2822 may generate tokens or other prompt elements that identify a requested or desired output modality (e.g., text, audio / visual, computer / robotic commands, etc.). The prompt composer 2822 may generate an embedding which may be provided separately from the prompt 2830 for use in an intermediate layer of the Al model 2824. The prompt composer 2822 may generate multiple tokenized sequences at once that constitute a series of conditional commands. In some embodiments, the prompt composer 2822 may utilize a formal prompt composition language such as Microsoft Guidance. In such a case, the composition language may utilize one or more formal structures that facilitate deterministic prompt composition as a function of mixed modality inputs. For example, the prompt composer 2822 may contain subroutines that process raw signal data, such as may be included in sensor data 2802 and user selection data 2808, and utilize this data to modify prompts 2830 in order to ensure specific types of Al outputs 2828.
[0315] In one embodiment, the prompt composer 2822 may be configured to identify among multiple available Al models 2824 the one best suited (specifically trained, historically successful, etc.) for addressing the need represented by the data in the prompt 2830. In this case, the prompt composer 2822 may provide a model identifier 2832 interpretable by a model selector 2826. The model selector 2826 may be configured to route the prompt 2830 to the specific Al model 2824 indicated by the model identifier 2832. In this manner, the Al system 2800 may be able to optimize model usage for the most quickly generated and most accurate Al output 2828.
[0316] Al models 2824 may include large language models (LLMs), a Generative Pre-trained Transformer (GPT) model, a generalist agent such as Gato, and other models such as may be readily anticipated by one of ordinary skill in the art. These Al model 2824 may be configured to supply Al outputs 2828 based on the prompt 2830. The Al output 2828 may be multimodal in some embodiments, and may take the form of text output, audio output, visual output, programmatic output such as computer or robotic commands, etc. Al output 2828 may be sent to users 2806, which may be the person providing the text data 2804 or user selection data 2808 inputs or other users similarly in contact with the Al system 2800 through a computational interface. The Al output 2828 may be sent to interconnected computing devices or systems 2836 such as these users 2806 may utilize. The Al output 2828 may be sent to robotic systems 2834 configured to perform automated tasks based on the Al output 2828.Docket No. FSP2504PCT
[0317] As shown in FIG. 29, computer system / server 2902 in cloud computing node 2900 is shown in the form of a general-purpose computing device. The components of computer system / server 2902 may include, but are not limited to, one or more processors or processing units 2906, a system memory 2904, and a bus 2926 that couples various system components, including system memory 2904, to processing units 2906.
[0318] Bus 2926 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnects (PCI) bus.
[0319] Computer system / server 2902 typically includes a variety of computer system readable media. Such media may be any available media that is accessible by computer system / server 2902, and it includes both volatile and non-volatile media, removable and non-removable media.
[0320] System memory 2904 may include computer system readable media in the form of volatile memory, such as Random access memory (RAM) 2908 and / or cache memory 2912. Computer system / server 2902 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example, a storage system 2920 may be provided for reading from and writing to a non-removable, non-volatile magnetic media (not shown and typically called a “hard drive”). Although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a “floppy disk”) and an optical disk drive for reading from or writing to a removable, non-volatile optical disk such as a CD-ROM, DVD-ROM or other optical media may be provided. In such instances, each may be connected to bus 2926 by one or more data media interfaces. As will be further depicted and described below, system memory 2904 may include at least one program product having a set (e.g., at least one) of program logic 2924 configured to carry out the functions of the present disclosure.
[0321] Program / utility 2922 having a set (at least one) of program logic 2924 may be stored in system memory 2904 by way of example, and not limitation, as well as an operating system, one or more application programs, other program logic 2924, and program data. Each of the operating system, one or more application programs, other program logic 2924, and programDocket No. FSP2504PCTdata or some combination thereof, may include an implementation of a networking environment. Program logic 2924 generally carry out the functions and / or methodologies of the present disclosure as described herein.
[0322] Computer system / server 2902 may also communicate with one or more external devices 2914 such as a keyboard, a pointing device, a display 2916, etc.; one or more devices that allow a user to interact with computer system / server 2902; and / or any devices (e.g., network card, modem, etc.) that allow computer system / server 2902 to communicate with one or more other computing devices. Such communication may occur via I / O interfaces 2910. Computer system / server 2902 may communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet) via network adapter 2918. As depicted, network adapter 2918 communicates with the other components of computer system / server 2902 via bus 2926. It may readily be understood that although not shown, other hardware and / or software components may be used in conjunction with computer system / server 2902. Examples include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0323] Referring now to FIG. 30, illustrative cloud computing environment 3000 is depicted. As shown, cloud computing environment 3000 comprises one or more cloud computing nodes 2900 with which computing devices such as, for example, a laptop 3002, a personal digital assistant (PDA) or cellular telephone 3004, an automobile computer system 3006, and / or a desktop computer 3008 may communicate. This allows for infrastructure, platforms, and / or software to be offered as services from cloud computing environment 3000, so that each client need not separately maintain such resources. It is understood that the types of computing devices shown in FIG. 30 are intended to be illustrative and that a cloud computing environment 3000 may communicate with any type of computerized device over any type of network and / or network / addressable connection (e.g., using a web browser).
[0324] Referring now to FIG. 31, a set of functional abstraction layers provided by a cloud computing environment 3000 such as is illustrated in FIG. 30. It may be understood in advance that the components, layers, and functions shown in FIG. 31 are intended to be illustrative, and the present disclosure is not limited thereto. As depicted, the following layers and corresponding functions are provided:Docket No. FSP2504PCT
[0325] Hardware and software layer 3102 includes hardware and software components.Examples of hardware components include mainframes. In one example, IBM® zSeries® systems and RISC (Reduced Instruction Set Computer) architecture based servers. In one example, IBM pSeries® systems, IBM xSeries® systems, IBM BladeCenter® systems, storage devices, networks, and networking components. Examples of software components include network application server software. In one example, IBM WebSphere® application server software and database software. In one example, IBM DB2® database software. (IBM, zSeries, pSeries, xSeries, BladeCenter, WebSphere, and DB2 are trademarks of International Business Machines Corporation in the United States, other countries, or both.)
[0326] Virtualization layer 3104 provides an abstraction layer from which the following exemplary virtual entities may be provided: virtual servers; virtual storage; virtual networks, including virtual private networks; virtual applications; and virtual clients.
[0327] Management layer 3106 provides the exemplary functions described below. Resource provisioning provides dynamic procurement of computing resources and other resources that are utilized to perform tasks within the Cloud computing environment. Metering and Pricing provide cost tracking as resources are utilized within the Cloud computing environment, and billing or invoicing for consumption of these resources. In one example, these resources may comprise application software licenses. Security provides identity verification for users and tasks, as well as protection for data and other resources. User portal provides access to the Cloud computing environment for both users and system administrators. Service level management provides Cloud computing resource allocation and management such that desired service levels are met. SLA planning and fulfillment provides pre-arrangement for, and procurement of, Cloud computing resources for which a future need is anticipated in accordance with an SLA.
[0328] Workloads layer 3108 provides functionality for which the Cloud computing environment is utilized. Examples of workloads and functions which may be provided from this layer include: mapping and navigation; software development and lifecycle management; virtual classroom education delivery; data analytics processing; transaction processing; and resource credit management. As mentioned above, all of the foregoing examples described with respect to FIG. 31 are illustrative, and the present disclosure is not limited to these examples.
[0329] FIG. 32 illustrates an exemplary tokenizer 3200 in accordance with one embodiment, data to be tokenized 3202 may be provided to the exemplary tokenizer 3200 in the form of aDocket No. FSP2504PCTtext string typed by a user. In one embodiment, the text string may be generated by performing voice-to-text conversion on an audio stream. The exemplary tokenizer 3200 may detect tokenizable elements 3204 within the data to be tokenized 3202. Each tokenizable element 3204 may be converted into a token 3206. The set of tokens 3206 created from the tokenizable elements 3204 of the data to be tokenized 3202 may be sent from the exemplary tokenizer 3200 as tokenized output 3208. The set of tokens 3206 may be such as are used to create the prompt FIG. 28.
[0330] For structured historical data such as plaintext, database, or web-based textual content, tokens may consist of the numerical indexes in an embedded or vectorized (e.g., word2vec or similar) representation of the text content such as are shown here. In some embodiments, a machine learning technique called an autoencoder may be utilized to transform plaintext inputs into high dimensional vectors that are suitable for indexing and tokenization ingestion by the prompt composer 2822 introduced with respect to FIG. 28.
[0331] In some embodiments, data to be tokenized may include audio, visual, or other multimodal data. For images, video, and similar visual data, tokenization may be performed using a convolution-based tokenizer such as a vision transformer. In some alternate embodiments, multimodal data may be quantized and converted into tokens 3206 using a codebook. In yet other alternate embodiments, multimodal data may be directly encoded and for presentation to a language model as a vector space encoding. An exemplary system that utilizes this tokenizer strategy is Gato, a generalist agent capable of ingesting a mixture of discrete and continuous inputs, images, and text as tokens.
[0332] FIG. 33 illustrates the training and deployment of a deep neural network 3300, such as the basic deep neural network 3500 illustrated in FIG. 35, according to at least one embodiment. In at least one embodiment, untrained neural network 3306 is trained using a training dataset 3302. In at least one embodiment, training framework 3304 is a PyTorch framework, whereas in other embodiments, training framework 3304 is a TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j , or another training framework. In at least one embodiment, training framework 3304 trains an untrained neural network 3306 and allows it to be trained using processing resources described herein to generate a trained neural network 3308. In at least one embodiment, weights may be chosen randomly or by pre-training using a deep belief network. In at least one embodiment, training may be performed in either a supervised, partially supervised, or unsupervised manner.Docket No. FSP2504PCT
[0333] In at least one embodiment, untrained neural network 3306 is trained using supervised learning, wherein training dataset 3302 includes an input paired with a desired output for the input, or where training dataset 3302 includes input having a known output and an output of untrained neural network 3306 is manually graded. In at least one embodiment, untrained neural network 3306 is trained in a supervised manner, processes inputs from training dataset 3302, and compares resulting outputs against a set of expected or desired outputs. In at least one embodiment, errors are then propagated back through untrained neural network 3306. In at least one embodiment, training framework 3304 adjusts weights that control untrained neural network 3306. In at least one embodiment, training framework 3304 includes tools to monitor how well untrained neural network 3306 is converging towards a model, such as trained neural network 3308, suitable to generating correct answers, such as in result 3312, based on input data such as a new dataset 3310. In at least one embodiment, training framework 3304 trains untrained neural network 3306 repeatedly while adjusting weights to refine an output of untrained neural network 3306 using a loss function and adjustment algorithm, such as stochastic gradient descent. In at least one embodiment, training framework 3304 trains untrained neural network 3306 until untrained neural network 3306 achieves the desired accuracy. In at least one embodiment, trained neural network 3308 may then be deployed to implement any number of machine learning operations.
[0334] In at least one embodiment, untrained neural network 3306 is trained using unsupervised learning, wherein untrained neural network 3306 attempts to train itself using unlabeled data. In at least one embodiment, an unsupervised learning training dataset 3302 will include input data without any associated output data or “ground truth” data. In at least one embodiment, untrained neural network 3306 may learn groupings within training dataset 3302 and may determine how individual inputs are related to other data in the training dataset 3302. In at least one embodiment, unsupervised training may be used to generate a self-organizing map in a trained neural network 3308 capable of performing operations useful in reducing the dimensionality of the new dataset 3310. In at least one embodiment, unsupervised training may also be used to perform anomaly detection, which allows the identification of data points in new dataset 3310 that deviate from normal observation agent of new dataset 3310.
[0335] In at least one embodiment, semi-supervised learning may be used, which is a technique in which training dataset 3302 includes a mix of labeled and unlabeled data. In at least one embodiment, training framework 3304 may be used to perform incremental learning,Docket No. FSP2504PCTsuch as through transferred learning techniques. In at least one embodiment, incremental learning allows trained neural network 3308 to adapt to new dataset 3310 without forgetting knowledge instilled within trained neural network 3308 during initial training. A trained neural network 3308 such as the one described may be used as the basis for Al and ML models such as may be used in computational systems to analyze complex data and provide results based on that analysis productive toward the improved knowledge or task action performance of people and computational and robotic systems.
[0336] The following figures set forth, without limitation, exemplary artificial intelligencebased systems that may be used to implement at least one embodiment.
[0337] FIG. 34A illustrates inference and / or training logic 3400a used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding training logic / hardware structure 3410 are provided below in conjunction with FIG. 34A and / or FIG.34B.
[0338] In at least one embodiment, training logic / hardware structure 3410 may include, without limitation, code and / or data storage 3402 to store forward and / or output weight and / or input / output data, and / or other parameters to configure neurons or layers of a neural network trained and / or used for inferencing in aspects of one or more embodiments. In at least one embodiment, training logic / hardware structure 3410 may include or be coupled to code and / or data storage 3402 to store graph code or other software to control the timing and / or order in which weight and / or other parameter information is to be loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture of a neural network to which such code corresponds. In at least one embodiment code and / or data storage 3402 stores weight parameters and / or input / output data of each layer of a neural network trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 3402 may be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory.
[0339] In at least one embodiment, any portion of code and / or data storage 3402 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 3402 may be cache memory, dynamicDocket No. FSP2504PCTrandomly addressable memory (DRAM), static randomly addressable memory (SRAM), nonvolatile memory (e.g., flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 3402 is internal or external to a processor, for example, or comprising DRAM, SRAM, flash, or some other storage type may depend on available storage on-chip versus off-chip, latency needs of training and / or inferencing functions being performed, batch size of data used in inferencing and / or training of a neural network, or some combination of these factors.
[0340] In at least one embodiment, training logic / hardware structure 3410 may include, without limitation, a code and / or data storage 3406 to store backward and / or output weight and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inferencing in aspects of one or more embodiments. In at least one embodiment, code and / or data storage 3406 stores weight parameters and / or input / output data of each layer of a neural network trained or used in conjunction with one or more embodiments during backward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, training logic / hardware structure 3410 may include or be coupled to code and / or data storage 3406 to store graph code or other software to control the timing and / or order in which weight and / or other parameter information is to be loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)).
[0341] In at least one embodiment, code, such as graph code, causes loading of weight or other parameter information into processor ALUs based on an architecture of a neural network to which such code corresponds. In at least one embodiment, any portion of code and / or data storage 3406 may be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 3406 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 3406 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 3406 is internal or external to a processor, for example, or comprising DRAM, SRAM, flash memory or some other storage type may depend on available storage on-chip versus off-chip, latency needs of training and / or inferencing functions being performed, batch size of data used in inferencing and / or training of a neural network, or some combination of these factors.Docket No. FSP2504PCT
[0342] In at least one embodiment, code and / or data storage 3402 and code and / or data storage 3406 may be separate storage structures. In at least one embodiment, code and / or data storage 3402 and code and / or data storage 3406 may be a combined storage structure. In at least one embodiment, code and / or data storage 3402 and code and / or data storage 3406 may be partially combined and partially separate. In at least one embodiment, any portion of code and / or data storage 3402 and code and / or data storage 3406 may be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory.
[0343] In at least one embodiment, training logic / hardware structure 3410 may include, without limitation, one or more arithmetic logic units 3412, including integer and / or floating point units, to perform logical and / or mathematical operations based, at least in part on, or indicated by, training and / or inference code (e.g., graph code), a result of which may produce activations (e.g., output values from layers or neurons within a neural network) stored in an activation storage 3414 that are functions of input / output and / or weight parameter data stored in code and / or data storage 3402 and / or code and / or data storage 3406. In at least one embodiment, activations stored in activation storage 3414 are generated according to linear algebraic and or matrix-based mathematics performed by arithmetic logic units 3412 in response to performing instructions or other code, wherein weight values stored in code and / or data storage 3406 and / or code and / or data storage 3402 are used as operands along with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 3406 or code and / or data storage 3402 or another storage on or off-chip.
[0344] In at least one embodiment, arithmetic logic units 3412 are included within one or more processors or other hardware logic devices or circuits, whereas in another embodiment, arithmetic logic units 3412 may be external to a processor or other hardware logic device or circuit that uses them (e.g., a co-processor). In at least one embodiment, arithmetic logic units 3412 may be included within a processor’s execution units or otherwise within a bank of ALUs accessible by a processor’s execution units either within the same processor or distributed between different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, code and / or code and / or data storage 3402, code and / or data storage 3406, and activation storage 3414 may share a processor or other hardware logic device or circuit, whereas, in another embodiment, they may be in different processors or other hardware logic devices or circuits, or some combinationDocket No. FSP2504PCTof same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 3414 may be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory. Furthermore, inferencing and / or training code may be stored with other code accessible to a processor or other hardware logic or circuit and fetched and / or processed using a processor’s fetch, decode, scheduling, execution, retirement, and / or other logic circuits.
[0345] In at least one embodiment, activation storage 3414 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, activation storage 3414 may be completely or partially within or external to one or more processors or other logic circuits. In at least one embodiment, a choice of whether activation storage 3414 is internal or external to a processor, for example, or comprising DRAM, SRAM, flash memory or some other storage type may depend on available storage on-chip versus off-chip, latency needs of training and / or inferencing functions being performed, batch size of data used in inferencing and / or training of a neural network, or some combination of these factors.
[0346] In at least one embodiment, the training logic / hardware structure 3410 illustrated in FIG. 34A may be used in conjunction with an application-specific integrated circuit (ASIC), such as a TensorFlow® Processing Unit from Google, an inference processing unit (IPU) from Graphcore™, or a Nervana® (e.g., “Lake Crest) processor from Intel Corp. In at least one embodiment, the training logic / hardware structure 3410 illustrated in FIG. 34A may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware, such as field programmable gate arrays (FPGAs).
[0347] FIG. 34B illustrates inference and / or training logic 3400b, according to at least one embodiment. In at least one embodiment, training logic / hardware structure 3410 may include, without limitation, hardware logic in which computational resources are dedicated or otherwise exclusively used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, the training logic / hardware structure 3410 illustrated in FIG. 34B may be used in conjunction with an application-specific integrated circuit (ASIC), such as TensorFlow® Processing Unit from Google, an inference processing unit (IPU) from Graphcore™, or a Nervana® (e.g., “Lake Crest”) processor from Intel Corp. In at least one embodiment, the training logic / hardware structure 3410 illustrated in FIG. 34B may be used in conjunction with CPU hardware, GPUDocket No. FSP2504PCThardware, or other hardware, such as field programmable gate arrays (FPGAs). In at least one embodiment, training logic / hardware structure 3410 includes, without limitation, code and / or data storage 3402 and code and / or data storage 3406, which may be used to store code (e.g., graph code), weight values and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment illustrated in FIG. 34B, each of code and / or data storage 3402 and code and / or data storage 3406 is associated with a dedicated computational resource, such as computational hardware 3404 and computational hardware 3408, respectively. In at least one embodiment, each of computational hardware 3404 and computational hardware 3408 comprises one or more ALUs that perform mathematical functions, such as linear algebraic functions, on information stored in code and / or data storage 3402 and code and / or data storage 3406, respectively, the result of which is stored in activation storage 3414.
[0348] In at least one embodiment, each of code and / or data storage 3402 and 3406 and corresponding computational hardware 3404 and 3408, respectively, correspond to different layers of a neural network, such that resulting activation from one storage / computational pair 3402 / 3404 of code and / or data storage 3402 and computational hardware 3404 is provided as an input to a next storage / computational pair 3406 / 3408 of code and / or data storage 3406 and computational hardware 3408, in order to mirror a conceptual organization of a neural network. In at least one embodiment, each of the storage / computational pairs 3402 / 3404 and 3406 / 3408 may correspond to more than one neural network layer. In at least one embodiment, additional storage / computation pairs (not shown) subsequent to or in parallel with storage / computation pairs 3402 / 3404 and 3406 / 3408 may be included in training logic / hardware structure 3410.
[0349] FIG. 35 illustrates a basic deep neural network 3500 in accordance with one embodiment. A basic deep neural network 3500 is based on a collection of connected units or nodes called artificial neurons which loosely model the neurons in a biological brain. Each connection, like the synapses in a biological brain, may transmit a signal from one artificial neuron to another. An artificial neuron that receives a signal may process it and then signal additional artificial neurons connected to it.
[0350] In common implementations, the signal at a connection between artificial neurons is a real number, and the output of each artificial neuron is computed by some non-linear function (the activation function) of the sum of its inputs. The connections between artificial neurons are called 'edges' or axons. Artificial neurons and edges typically have a weight that adjusts asDocket No. FSP2504PCTlearning proceeds. The weight increases or decreases the strength of the signal at a connection. Artificial neurons may have a threshold (trigger threshold) such that the signal is sent if the aggregate signal crosses that threshold. Typically, artificial neurons are aggregated into layers. Different layers may perform different kinds of transformations on their inputs. Signals travel from the first layer (the input layer 3502), to the last layer (the output layer 3506), possibly after traversing one or more intermediate layers, called hidden layers 3504.
[0351] Referring to FIG. 36, an artificial neuron 3600 receiving inputs from predecessor neurons consists of the following components:• inputs Xi;• weights Wi applied to the inputs;• an optional threshold (b), which stays fixed unless changed by a learning function; and • an activation function 3602 that computes the output from the previous neuron inputs and threshold, if any.
[0352] An input neuron has no predecessor but serves as input interface for the whole network of artificial neurons 3600, such as may form a basic deep neural network 3500. Similarly an output neuron has no successor and thus serves as output interface of the whole network.
[0353] The network includes connections, each connection transferring the output of a neuron in one layer to the input of a neuron in a next layer. Each connection carries an input x and is assigned a weight w.
[0354] The activation function 3602 often has the form of a sum of products of the weighted values of the inputs of the predecessor neurons.
[0355] The learning rule is a rule or an algorithm which modifies the parameters of the neural network, in order for a given input to the network to produce a favored output. This learning process typically involves modifying the weights and thresholds of the neurons and connections within the network.CONCLUSION
[0356] The disclosed system may leverage Kubernetes for scalability, Docker for portability, and Git for version control in various embodiments. In this manner, the system may integrate seamlessly with external frameworks, supports both cloud and on-premise LMs, and provides robust interfaces for managing and scaling Al workflows.Docket No. FSP2504PCT
[0357] Within this disclosure, different entities (which may variously be referred to as "units," "circuits," other components, etc.) may be described or claimed as "configured" to perform one or more tasks or operations. This formulation — [entity] configured to [perform one or more tasks] — is used herein to refer to structure (i.e., something physical, such as an electronic circuit). More specifically, this formulation is used to indicate that this structure is arranged to perform the one or more tasks during operation. A structure may be said to be "configured to" perform some task even if the structure is not currently being operated. A "credit distribution circuit configured to distribute credits to a plurality of processor cores" is intended to cover, for example, an integrated circuit that has circuitry that performs this function during operation, even if the integrated circuit in question is not currently being used (e.g., a power supply is not connected to it). Thus, an entity described or recited as "configured to" perform some task refers to something physical, such as a device, circuit, memory storing program instructions executable to implement the task, etc. This phrase is not used herein to refer to something intangible.
[0358] The term "configured to" is not intended to mean "configurable to." An unprogrammed field programmable gate array (FPGA), for example, would not be considered to be "configured to" perform some specific function, although it may be "configurable to" perform that function after programming.
[0359] Reciting in the appended claims that a structure is "configured to" perform one or more tasks is expressly intended not to invoke 35 U.S.C. § 112(f) for that claim element.Accordingly, claims in this application that do not otherwise include the "means for" [performing a function] construct should not be interpreted under 35 U.S.C § 112(f).
[0360] As used herein, the term "based on" is used to describe one or more factors that affect a determination. This term does not foreclose the possibility that additional factors may affect the determination. That is, a determination may be solely based on specified factors or based on the specified factors as well as other, unspecified factors. Consider the phrase "determine A based on B." This phrase specifies that B is a factor that is used to determine A or that affects the determination of A. This phrase does not foreclose that the determination of A may also be based on some other factor, such as C. This phrase is also intended to cover an embodiment in which A is determined based solely on B. As used herein, the phrase "based on" is synonymous with the phrase "based at least in part on."Docket No. FSP2504PCT
[0361] As used herein, the phrase "in response to" describes one or more factors that trigger an effect. This phrase does not foreclose the possibility that additional factors may affect or otherwise trigger the effect. That is, an effect may be solely in response to those factors, or may be in response to the specified factors as well as other, unspecified factors. Consider the phrase "perform A in response to B." This phrase specifies that B is a factor that triggers the performance of A. This phrase does not foreclose that performing A may also be in response to some other factor, such as C. This phrase is also intended to cover an embodiment in which A is performed solely in response to B.
[0362] As used herein, the terms "first," "second," etc. are used as labels for nouns that they precede, and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.), unless stated otherwise. For example, in a register file having eight registers, the terms "first register" and "second register" may be used to refer to any two of the eight registers, and not, for example, just logical registers 0 and 1.
[0363] When used in the claims, the term "or" is used as an inclusive or and not as an exclusive or. For example, the phrase "at least one of x, y, or z" means any one of x, y, and z, as well as any combination thereof.
[0364] As used herein, a recitation of “and / or” with respect to two or more elements should be interpreted to mean only one element, or a combination of elements. For example, “element A, element B, and / or element C” may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. In addition, “at least one of element A or element B” may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, “at least one of element A and element B” may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0365] Reference has been made in detail to implementations and embodiments of various aspects and variations of systems and methods described herein. Although several exemplary variations of the systems and methods are described herein, other variations of the systems and methods may include aspects of the systems and methods described herein combined in any suitable manner having combinations of all or some of the aspects described.
[0366] In this disclosure, it is to be understood that the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It is also to be understood that the term “and / or” as used herein refers to and encompasses any andDocket No. FSP2504PCTall possible combinations of one or more of the associated listed terms. It is further to be understood that the terms “includes,” “including,” “comprises,” and / or “comprising,” when used herein, specify the presence of stated features, integers, steps, operations, elements, components, and / or units but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, units, and / or groups thereof.
[0367] Certain aspects of the present disclosure include process steps and instructions described herein in the form of an algorithm. It should be noted that the process steps and instructions of the present disclosure could be embodied in software, firmware, or hardware and, when embodied in software, could be downloaded to reside on and be operated from different platforms used by a variety of operating systems. Unless specifically stated otherwise as apparent from the following discussion, it is appreciated that, throughout the description, discussions utilizing terms such as “processing,” “computing,” “calculating,” “determining,” “displaying,” “generating” or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within computer system memories or registers or other such information storage, transmission, or display devices.
[0368] The subject matter of the present disclosure is described with specificity herein to meet statutory requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the applicants have contemplated that the claimed subject matter might also be embodied in other ways, to include different steps or combinations of steps similar to the ones described in this document, in conjunction with other present or future technologies.Moreover, although the terms “step” and / or “block” may be used herein to connote different elements of methods employed, the terms should not be interpreted as implying any particular order among or between various steps herein disclosed unless and except when the order of individual steps is explicitly described.
[0369] The present disclosure in some embodiments also relates to a device for performing the operations herein. This device may be specially constructed for the purposes needed, or it may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-transitory, storage medium, such as, but not limited to, any type of disk, including floppy disks, USB flash drives, external hard drives, optical disks, CD-ROMs, magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), EPROMs, EEPROMs, magnetic orDocket No. FSP2504PCToptical cards, application-specific integrated circuits (ASICs), or any type of media suitable for storing electronic instructions, and each connected to a computer system bus. Furthermore, the computing systems referred to in the specification may include a single processor or may be architectures employing multiple processor designs, such as for performing different functions or for increased computing capability. Suitable processors include CPUs, GPUs, field programmable gate arrays (FPGAs), and ASICs.
[0370] The methods, devices, and systems described herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may also be used with programs in accordance with the teachings herein, or it may prove convenient to construct a more specialized apparatus to perform the method steps needed. The structure for a variety of these systems will appear in the description below. In addition, the present disclosure is not described with reference to any particular programming language. It will be appreciated that a variety of programming languages may be used to implement the teachings of the present disclosure as described herein.
[0371] Having thus described illustrative embodiments in detail, it will be apparent that modifications and variations are possible without departing from the scope of the disclosure as claimed. The scope of disclosed subject matter is not limited to the depicted embodiments but is rather set forth in the following Claims.
Claims
Docket No. FSP2504PCTCLAIMSWhat is claimed is:
1. A method for analyzing data comprising, at a computing system:receiving a user input comprising an indication of the data to be analyzed and at least one request related to the data;accessing, by an input evaluator agent, the data in available data silos; determining, by the input evaluator agent, whether the at least one request is at least partially satisfied by a cached result previously generated using at least one language model of a plurality of language models; andon condition the at least one request is not at least partially satisfied by the cached result:receiving, at a response generator agent, the user input and the data; selecting, by the response generator agent, a set of language models from the plurality of language models;provisioning the set of language models as artificial intelligence (Al) custom agents to analyze the data in view of the user input; andanalyzing the data with the Al custom agents to generate a response to the at least one request.
2. The method of claim 1, analyzing the data with the Al custom agents further comprises: performing a high-level analysis to generate the response, wherein the response includes at least one high-level recommendation;presenting the at least one high-level recommendation to a user;performing a detailed analysis to generate at least one detailed response, wherein the at least one detailed response includes detailed recommendations related to the at least one high-level recommendation; andon condition the user requests details related to the at least one high-level recommendation :presenting the at least one detailed response to the user.Docket No. FSP2504PCT3. The method of claim 1, wherein analyzing the data with the Al custom agents includes generating an abstract semantic tree (AST) representation of unstructured data from the available data silos.
4. The method of claim 3, further comprising:generating knowledge graphs based on semantic relationships among the unstructured data; andmapping the AST to the knowledge graphs.
5. The method of claim 4, wherein generating the knowledge graphs includes generating a semantic graph for each primitive type of the unstructured data, wherein the primitive type is at least one of text, visual relationships, image relationships, and tables.
6. The method of claim 1, further comprising:determining, by a methodology evaluator agent, that the response complies with at least one of business rules, business constraints, and the data indicated by the user input; and validating the response against structural logic provided by at least one of:an abstract semantic tree (AST) representation of unstructured data from the available data silos;knowledge graphs based on semantic relationships among the unstructured data; andthe AST mapped to the knowledge graphs.
7. The method of claim 1, further comprising:receiving, by an output agent, the response to the at least one request from the Al custom agents; andgenerating, by the output agent, an output comprising the response to the at least one request.
8. The method of claim 7, wherein the output comprises at least one of a document, a presentation, text data, an image, a video, audio data, embedded content, and web content.
9. The method of claim 1, further comprising storing at least one Al custom agent in a cache.
10. The method of claim 1, further comprising validating at least one Al custom agent.Docket No. FSP2504PCT11. The method of claim 10, wherein validating the at least one Al custom agent comprises comparing the at least one Al custom agent with a cached Al custom agent.
12. The method of claim 1, wherein provisioning the set of language models as Al custom agents includes:populating at least one Al custom agent with context based on the user input; assigning at least one tool to the at least one Al custom agent, the at least one tool allowing the at least one Al custom agent to interact with the available data silos;assigning a language model of the set of language models as a reasoning model of the at least one Al custom agent; andinstructing the reasoning model of the at least one Al custom agent on how to interpret the data within the populated context based on the user input.
13. The method of claim 1, wherein selecting the set of language models comprises evaluating at least one of computational power, memory, storage, and network bandwidth of the computing system.
14. The method of claim 1, wherein selecting the set of language models comprises selecting the set of language models based on at least one of capability, latency, cost, and uptime.
15. The method of claim 1, wherein selecting the set of language models comprises:determining a complexity of the at least one request; andselecting the set of language models based on the complexity of the at least one request.
16. The method of claim 1, wherein the response comprises code.
17. A system comprising one or more processors and memory storing instructions for execution by the one or more processors for causing the system to:receive a user input comprising an indication of data to be analyzed and at least one request related to the data;access, by an input evaluator agent, the data in available data silos;determine, by the input evaluator agent, whether the at least one request is at least partially satisfied by a cached result previously generated using at least one language model of a plurality of language models; andDocket No. FSP2504PCTon condition the at least one request is not at least partially satisfied by the cached result:receive, at a response generator agent, the user input and the data; select, by the response generator agent, a set of language models from the plurality of language models;provision the set of language models as artificial intelligence (Al) custom agents to analyze the data in view of the user input; andanalyze the data with the Al custom agents to generate a response to the at least one request.
18. The system of claim 17, the instructions further including:generate, by the Al custom agents, an abstract semantic tree (AST) representation of unstructured data from the available data silos; andmap the AST to the knowledge graphs.
19. The system of claim 17, the instructions further including:populate at least one Al custom agent with context based on the user input;assign at least one tool to the at least one Al custom agent, the at least one tool allowing the at least one Al custom agent to interact with the available data silos;assign a language model of the set of language models as a reasoning model of the at least one Al custom agent; andinstruct the reasoning model of the at least one Al custom agent on how to interpret the data within the populated context based on the user input.
20. The system of claim 17, the instructions further including:receive, by an output agent, the response to the at least one request from the Al custom agents; andgenerate, by the output agent, an output comprising the response to the at least one request.