System and method for generating action-oriented enterprise outputs based on spatial memory and multi-AI orchestration
The control system addresses AI-based systems' inefficiencies by integrating multimodal data into structured spatial memory for dynamic AI orchestration, enhancing processing speed, accuracy, and ethical decision-making through transformer-based predictions and API execution.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2026-03-10
AI Technical Summary
Existing AI-based systems struggle with processing heterogeneous data from multiple sources due to limitations in flexibility, compatibility, and lack of orchestration mechanisms, leading to inefficiencies in dynamic task execution and integration of multiple AI models.
A control system that interprets intent from multimodal inputs, stores data in short-term and long-term memory, converts it into structured spatial memory, and uses transformer-based models to predict events, dynamically selecting and executing AI APIs while ensuring ethical considerations.
Enhances processing speed, accuracy, and reliability by enabling dynamic AI orchestration, improving decision-making through trend-based predictions and ethical validation, and facilitating proactive action planning.
Smart Images

Figure 2026041668000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates generally to artificial intelligence (AI) and machine learning (ML) architectures, and more particularly to AI-based systems that combine short-term, long-term, and spatial memory to interpret intent from heterogeneous multimodal inputs and coordinate multiple AI models or enterprise application programming interfaces (APIs) to generate outputs that drive subsequent human or machine actions.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 681,468, filed August 9, 2024, the entire contents of which are incorporated herein by reference for all purposes. [Background technology]
[0003] As digital transformation progresses across various sectors, artificial intelligence (AI) is increasingly being utilized. For example, AI is used to perform a variety of operations, including data analysis, automating decision-making processes, and complex computational tasks. Performing such operations requires processing large amounts of heterogeneous data originating from multiple input sources and formats, including structured and unstructured data. In response to the increasing complexity of operations and data diversity, there has been a shift toward utilizing AI-based components to process specific analytical or functional tasks in a distributed environment. However, as AI models become more specialized and diverse, AI-based systems using a single AI model may be insufficient to process all categories of input or meet various system goals.
[0004] Furthermore, because a single AI-based system relies on static configurations and predefined model associations, limiting its flexibility and responsiveness to evolving demands, the AI-based system may not be able to support coordinated execution, dynamic selection, or seamless switching between multiple AI models based on the nature of a given input or task. Furthermore, integrating multiple AI models into an AI-based system can often be hindered by compatibility constraints, a lack of standardized interfaces, and the absence of an orchestration mechanism that can manage interactions between models. The rapid development cycle of AI technology creates additional challenges as new models may offer enhanced performance or capabilities that AI-based systems cannot easily adopt. Summary of the Invention [Means for solving the problem]
[0005] This Summary is provided to present selected concepts in a simplified form that are further described in the Detailed Description. It is not intended to identify key features or delineate the scope of the claims.
[0006] In one aspect, the present disclosure provides a control system that interprets intent from multimodal inputs and generates enterprise outputs. The system ingests text, conversation logs, audio streams, images, videos, and sensor data and stores raw text and binary content metadata in a short-term memory. The system invokes one or more AI models via a connector layer to convert the short-term memory into a summarized and standardized long-term memory. The long-term memory is then converted, again via an AI model, into a structured spatial memory organized along feature axes such as time, location, actor, action, and motivation. Features and historical correlations are extracted from the spatial memory, sequences with similar behaviors are clustered to derive trends, and a transformer-based event prediction model is trained based on the trend data. Using the resulting model, the system predicts upcoming events, generates and tests hypotheses against the archived spatial memory, and records the tested hypotheses in a layered meta-memory.
[0007] Based on predicted events and validated hypotheses, the system autonomously selects one or more approved AI or enterprise APIs through the same connector layer, executes them, and provides at least one output (e.g., a report, spreadsheet, code snippet, or robotic process instruction). Each output is validated by a responsible AI module to mitigate bias, hallucination, or policy violations. In this way, the disclosed architecture integrates intent interpretation, memory management, and dynamic AI orchestration to extend machine capabilities and facilitate subsequent human or system decision-making.
[0008] In another aspect, the present disclosure relates to a non-transitory computer-readable medium comprising machine-executable instructions that may be executable by a processor to perform the methods discussed herein.
[0009] It is understood that methods according to the present disclosure can include any combination of aspects and features described herein, i.e., methods according to the present disclosure are not limited to the combinations of aspects and features specifically described herein, but also include any combination of the aspects and features provided.
[0010] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other features and advantages of the disclosure will be apparent from the description and drawings, and from the claims.
[0011] Various examples according to the present disclosure are described with reference to the following drawings.
[0012] Like reference numbers and designations in the various drawings indicate like elements. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a block diagram illustrating an exemplary control system that integrates short-term memory, long-term memory, and spatial memory and interfaces with multiple artificial intelligence (AI) models or enterprise application programming interfaces (APIs), according to embodiments of the present disclosure. [Figure 2] 1 is a high-level process flow illustrating how a control system interprets and executes intents from heterogeneous multi-modal inputs to generate enterprise outputs, according to an implementation of the present disclosure. [Figure 3] 1 is a detailed flowchart of an exemplary method for extracting 5W features, bundling trends, and predicting subsequent events through a transformer-based model, according to implementations of the present disclosure. [Figure 4] 1 is a flowchart illustrating a method for storing prediction results generated by a transformer-based event prediction model in a prediction history for continuous learning and auditability, according to an implementation of the present disclosure. [Figure 5]1 is a flowchart illustrating a method for training a model to predict an upcoming event using features and trends extracted from a spatial memory, according to an implementation of the present disclosure. [Figure 6] 1 is a flowchart illustrating a method for extracting features and trends from multiple inputs and using hypotheses and tests to turn collective intelligence into meta-memory, according to an implementation of the present disclosure. [Figure 7] FIG. 1 is a process flow diagram illustrating an exemplary method in which a push action agent actively prompts a human or system to take an action using past memories stored in spatial memory and hypotheses stored in meta-memory extracted from the past memories, according to an embodiment of the present disclosure. [Figure 8] FIG. 1 is a flow diagram illustrating an exemplary method for generating enterprise outputs from multimodal inputs, consistent with implementations of the present disclosure. [Figure 9] FIG. 2 illustrates an example computer system for implementing the control system disclosed in the example system of FIG. 1 in accordance with an implementation of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0014] In the following description, various examples are described by way of example, and not limitation, in the accompanying drawing figures. References to various examples in this disclosure do not necessarily refer to the same example, and such references mean at least one. While specific implementations and other details are discussed, it should be understood that this is done for illustrative purposes only. Those skilled in the relevant art will recognize that other components and configurations may be used without departing from the scope and spirit of the claimed subject matter.
[0015] Any reference herein to an "example" (e.g., "for example," "an example of," "by way of example," or the like) should be considered a non-limiting example, whether explicitly stated or not.
[0016] The terms used herein generally have their ordinary meaning in the art, within the context of this disclosure and in the specific context in which each term is used. Alternative terms and synonyms may be used for any one or more of the terms discussed herein, and no special significance is attached to whether a term is elaborated or discussed herein. Synonyms for particular terms are provided. The listing of one or more synonyms does not exclude the use of other synonyms. The use of examples anywhere in this specification, including examples of any term discussed herein, is merely illustrative and is not intended to further limit the scope and meaning of the disclosure or the scope and meaning of the exemplified term. Similarly, the disclosure is not limited to the various examples provided herein.
[0017] Without intending to limit the scope of the present disclosure, examples of instruments, devices, methods, and their related results according to the examples of the present disclosure are given below. For the convenience of the reader, the examples may use titles or subtitles, but please note that this does not limit the scope of the present disclosure in any way. Unless otherwise defined, the technical and scientific terms used herein have the meanings that are commonly understood by those skilled in the art to which this disclosure pertains. In the event of any discrepancy, the present document, including definitions, shall prevail.
[0018] When the term "comprising" is used, it means "including, but not necessarily limited to" and specifically indicates open-ended inclusion or membership in such listed combinations, groups, series, and the like.
[0019] The term "a" means "one or more" unless the context clearly indicates a single element.
[0020] "First," "second," etc. are labels to distinguish between components or blocks that are similarly named by nature, and do not imply any sequential or numerical constraints.
[0021] "And / or" of two possibilities means either one or both of the listed possibilities ("A and / or B" covers A only, B only, or both A and B together), and when present with more than two listed possibilities means any individual possibility only, all possibilities together, or a combination of fewer than all possibilities. Where A through N are possibilities, phrases of the form "at least one of A... and N" mean "and / or" of the listed possibilities (e.g., at least one A, at least one N, at least one A and at least one N, etc.).
[0022] It should also be noted that in some alternative implementations, the functions / acts described may occur out of the order depicted. For example, two steps disclosed or shown as successive may, in fact, be performed substantially concurrently or may sometimes be performed in the reverse order, depending on the functions / acts involved.
[0023] The following description provides specific details to provide a thorough understanding of the examples. However, those skilled in the art will understand that the examples may be practiced without these specific details. For example, systems may be shown in block diagrams to avoid obscuring the examples in unnecessary detail. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail to avoid obscuring the illustrative examples.
[0024] Implementations of the present disclosure provide systems and methods for interpreting and executing intents from multimodal input data to generate enterprise outputs (also referred to as enterprise solutions), including enterprise-level outputs. To generate the enterprise outputs, data from input data sources, including but not limited to text documents, conversation logs, audio logs, image logs, and video logs, may be received and stored in a short-term memory. The short-term memory may be transformed and stored into a summarized and organized long-term memory via connectors that interface with various artificial intelligence (AI) models, and the long-term memory may be further structured into a spatial memory using the AI models via the connectors. Furthermore, features, including temporal attributes, spatial attributes, contextual attributes, causal attributes, and entity-related attributes, may be extracted from both the multimodal input data and historical data. Based on the features, trends may be determined and stored in the spatial memory, and a transformer-based event prediction model may be trained. Based on the learned trends, upcoming events may be predicted, hypotheses may be developed and verified using historical data, and a layered meta-memory may be generated to guide proactive decision-making. Additionally, implementations of the present disclosure may enable the execution of enterprise-approved AI and application programming interfaces (APIs) based on predicted events, which may enable the generation of dynamic outputs such as reports, spreadsheets, code, or robotic process automation configurations. Such collaboration of internal and external systems may improve the accuracy, responsiveness, and viability of enterprise decision-making while ensuring that enterprise-level outputs are validated for ethical considerations and ill-conceived risks.
[0025] In summary, multimodal input data (e.g., text, audio, images, and video) may be received, processed, and interpreted to dynamically generate ethically informed enterprise outputs. Implementations may employ transformer-based event prediction models to extract temporal, spatial, contextual, causal, and entity-related features from both multimodal input data and historical data, which may be used to identify trends and predict subsequent events. Such predictions may be organized into a layered meta-memory to improve decision-making and proactive action planning.
[0026] Implementations can utilize short-term and long-term memory, as well as spatial memory, enabling AI models to collaborate with multiple external APIs and handle complex tasks beyond the scope of a single AI agent. As a result, implementations of the present disclosure can improve processing speed through distributed collaboration, improve accuracy and reliability through trend-based prediction and hallucination detection, and provide efficient memory management through structured long-term memory transformation and spatial organization. Furthermore, implementations integrate ethical considerations into outputs and cross-validate outputs using historical patterns and predefined parameters, ensuring reliable, context-aware automation for enterprise-level decision-making.
[0027] FIG. 1 is a block diagram illustrating an exemplary control system that integrates short-term memory, long-term memory, and spatial memory and interfaces with multiple artificial intelligence (AI) models or enterprise application programming interfaces (APIs) in accordance with an implementation of the present disclosure. The exemplary system 100 illustrated in FIG. 1 includes a control system 102, multimodal input data 104 received from various input data sources, and a connector 106. The control system 102 includes a memory 108 that includes a short-term memory 110 (short-term memory (text) 110-1 and short-term memory (audio) 110-N), a long-term memory 112 (long-term memory (text) 112-1 and long-term memory (audio) 112-N), a spatial memory 114, and a metamemory 116. The control system 102 further includes a main control service 118, an action executor 120, a hippocampal agent 122, a metamemory agent 124, a push action agent 126, and a responsible AI (RAI) check service 128. For simplicity, only one control system 102 is shown in Figure 1. However, in some implementations, the example system 100 may include multiple control systems.
[0028] In some implementations, the input data sources may also be referred to as heterogeneous upstream data repositories. Examples of input data sources may include enterprise applications, databases, servers, cloud platforms, repositories, Customer Relationship Management (CRM) systems, websites, service history databases, Internet of Things (IoT) platforms, customer support platforms, management systems, and / or the like. The input data sources may store multimodal input data 104. The multimodal input data 104 may be structured data, semi-structured data, and / or unstructured data. Non-limiting examples of structured data, semi-structured data, and unstructured data may include structured tables, semi-structured documents (e.g., JavaScript Object Notation (JSON) or Extensible Markup Language (XML) documents), and unstructured text or multimedia content.
[0029] The multimodal input data 104 may include one or more of text data 104-1, audio logs 104-2, image logs 104-3, video logs 104-N, and / or the like. Examples of multimodal input data 104 may include, but are not limited to, corporate documents, chat logs, customer support transcripts, emails, voice call recordings, closed-circuit television (CCTV) or surveillance video feeds, user interaction logs from mobile or web applications, sensor data, social media content, and / or uploaded images.
[0030] The audio log 104-2 may include speech-to-text (STT) data and raw audio data, while the image log 104-3 may include contextual information, image tag metadata, and image data. Similarly, the video log 104-N may include video tag data, video content, and contextual information extracted from recorded or streaming video. Additionally, the audio log 104-2 may be captured via one or more microphones integrated into a user device, communication system, or IoT-enabled environment. The image log 104-3 may be derived from still images taken via a camera, scanner, or other imaging device. The video log 104-N may be obtained from a continuous video stream or recorded video, for example, from a movie, a surveillance camera, or a screen recording.
[0031] For example, control system 102 may extract STT summary data and audio data pointers from audio log 104-2, context summary data and image data pointers from image log 104-3, and context summary data and video data pointers from video log 104-N. The pre-processed and indexed data is utilized for downstream feature extraction, trend analysis, and prediction.
[0032] In one embodiment, the control system 102 may be a server system. Some examples of a server system may be, but are not limited to, a cloud server, a centralized server, a rack server, a network server, a computer-based server, an on-premise server, a dedicated server, a remote server, and the like. In some examples, the control system 102 may be implemented as an on-premise system operated by a company or a third party engaged in cross-platform interaction and data management. In some examples, the control system 102 may be implemented as an off-premise system (e.g., a cloud or on-demand system) operated by a company or a third party on behalf of a company. In some examples, the control system 102 may be implemented in a cloud environment. For simplicity, the control system 102 shown in FIG. 1 may be a cloud environment, which is intended to represent various forms of servers, including web servers, application servers, proxy servers, network servers, server pools, and / or the like. Furthermore, in some other implementations, the control system 102 may be implemented by a single device or by a combination of multiple devices that may be operatively connected or networked. The control system 102 may be implemented in hardware or an appropriate combination of hardware and software.
[0033] The control system 102 includes a processor (not shown in FIG. 1 ). In some implementations, the control system 102 includes two or more processors. A processor may include, for example, a microprocessor, a microcomputer, a microcontroller, a digital signal processor, a central processing unit, a state machine, a logic circuit, and / or any device that manipulates data or signals based on operational instructions. The processor may execute machine-readable program instructions for generating an enterprise solution stored in a memory (not shown in FIG. 1 ). The processor's execution of the machine-readable program instructions may enable the control system 102 to perform one or more operations described herein related to interpreting intent from multiple inputs to generate an enterprise solution. “Hardware” may include discrete components, integrated circuits, application-specific integrated circuits, field-programmable gate arrays, digital signal processors, or other suitable hardware combinations. “Software” may include one or more objects, agents, threads, lines of code, subroutines, individual software applications, two or more lines of code, or one or more software applications or other suitable software structures running on one or more processors.
[0034] Additionally, control system 102 may include memory 108 including short-term memory 110 (including short-term memory data 110-1 through 110-N), long-term memory 112 (including long-term memory data 112-1 through 112-N), spatial memory 114, and meta-memory 116.
[0035] In some embodiments, the short-term memory 110 stores large binary data corresponding to the multimodal input data 104, such as text data 104-1, audio logs 104-2, image logs 104-3, and video logs 104-N, in separate storage systems, and retains only the corresponding storage paths and metadata in a non-relational database for fast lookup and retrieval. The long-term memory 112 retains standardized structured data and core feature keywords in a non-relational database to enable semantic-level search and feature-based aggregation. The spatial memory 114 is initially organized in a chronological format by the "five Ws"—for example, when, where, who, what, and why (optionally including how)—and stored in a relational database (RDB) data mart. Subsequently, during hypothesis development, additional axes can be introduced to enhance multidimensional analysis, and the "five Ws" are reapplied with the new axes. The structured data in the spatial store 114 captures the relationships between entities along a timeline, enabling efficient bundling of trends by extracting and grouping data from the RDB by time and location, thereby avoiding brute force processing of large datasets and enabling scalable analysis of temporal and contextual trends.
[0036] In an exemplary embodiment, the control system 102 receives multimodal input data 104 from various input data sources. The multimodal input data 104 may include, but is not limited to, text data 104-1, audio logs 104-2, image logs 104-3, video logs 104-N, and / or the like. The raw binary input data (the multimodal input data 104 in its raw form) is stored in storage, and metadata is stored in a short-term memory 110, where the metadata and access paths are maintained for efficient lookup. The control system 102 further converts the short-term memory 110 into a summarized and organized long-term memory 112 using an artificial intelligence (AI) model via the connector 106. The long-term memory 112 is then structured into a spatial memory 114 using the AI model. The control system 102 extracts one or more features from the multimodal input data 104 and correlates the multimodal input data 104 with past trends stored in the spatial memory 114. These features and trends are used to train a transformer-based subsequent event prediction model, enabling accurate prediction of future events.
[0037] The connectors 106 may connect the control system 102 to multiple systems or AI models (AI), such as externally generated AI models (generated AI), internal custom AI models (internal custom AI), external data (internet data), and internal data (enterprise / company data). The connectors 106 facilitate integration, data exchange, and collaboration between the control system 102 and multiple systems or AI models to efficiently generate enterprise outputs. Externally generated AI is a cloud-based or third-party AI service capable of generating text, images, or other outputs. For example, the control system 102 may use an externally generated AI, such as OpenAI's GPT-4, to generate enterprise reports or suggest action items based on customer support transcripts. Internal custom AI models are proprietary AI models developed and trained within an enterprise to address specific enterprise requirements or domain-specific tasks. For example, an internal fraud detection AI model analyzes transaction data and flags anomalies specific to the enterprise's operating region or regulations. External data sources (internet data) are publicly available or licensed external datasets from the web or third-party APIs. For example, the control system 102 may retrieve live market trends, weather data, or news articles via public or corporate APIs to enrich the decision-making process. Internal data sources (corporate data) include private, secure databases and knowledge repositories maintained within the enterprise. For example, internal CRM databases, Enterprise Resource Planning (ERP) logs, or employee performance records may be accessed to generate Human Resources (HR) reports, inventory forecasts, or automated workflows.
[0038] The AI model of the control system 102, while interacting with multiple external AIs and data via connectors 106, identifies intent based on various input data, including not only text but also images and voice, and functions based on past memories (processing results) or past data. The control system 102 dynamically generates output that takes human ethics into consideration. The control system 102 has two important features: (1) the AI model in charge of the command chain formulates subsequent action plans and executes tasks based not only on external input but also on actively suggested instructions from past memories, and (2) the control system 102 actively prompts actions using accumulated past memories and hypotheses extracted and verified from those memories.
[0039] The control system 102 further includes various modules, such as a main control service 118, an action executor 120, a hippocampal agent 122, a metamemory agent 124, a push action agent 126, and a responsible AI (RAI) check service 128. The main control service 118 serves as a central coordination engine managing the flow of data and control signals in the control system 102. The main control service 118 is responsible for receiving multimodal input data 104 from various input sources, coordinating with internal modules, and interfacing with various external AI systems and APIs via connectors 106. The main control service 118 interprets user intent, formulates an action plan based on the multimodal input data 104 and historical data, and dynamically modifies the action plan in real time based on feedback from external systems. Historical data can be associated with past inputs and retrieved from the connectors 106 via AI models. The main control service 118 also queries a catalog database of available AI services to identify and trigger the execution of appropriate APIs or AIs corresponding to predicted or detected events. Additionally, the main control service 118 interfaces with a push action agent 126 to trigger proactive suggestions when a verified hypothesis is met or aligned.
[0040] The action executor 120 serves as an operational intermediary between the main control service 118 and the analysis subsystem, which includes the hippocampus agent 122 and the metamemory agent 124. The action executor 120 is responsible for routing data and control instructions between components, executing commands derived from interpreted user intent, and returning processed results to the orchestration layer. The action executor 120 enables the execution of action plans, aggregating response data from AI components, and providing results for ethical review or further transformation.
[0041] The hippocampal agent 122 performs memory-based feature extraction and prediction of subsequent events based on both the multimodal input data 104 and historical data. The hippocampal agent 122 extracts semantic features, including temporal, spatial, causal, contextual, and entity-specific attributes, from the multimodal input data. The hippocampal agent 122 also predicts subsequent events using a transformer-based large language model (LLM) trained using previously extracted trends and patterns stored in the spatial memory 114. The predictions generated by the hippocampal agent 122 are used to guide the control system 102 in formulating enterprise solutions or initiating automated processes.
[0042] The metamemory agent 124 is responsible for managing and querying the metamemory 116, a structured memory layer that stores tested hypotheses abstracted from long-term trends and historical patterns. The metamemory agent 124 generates candidate hypotheses based on trends and historical data extracted from the spatial memory. The metamemory agent 124 tests the candidate hypotheses by cross-referencing them with previously stored or historical data patterns in the spatial memory to determine their consistency and relevance. For each candidate hypothesis, the metamemory agent 124 assigns a confidence score that reflects its consistency with past behavioral, contextual, and causal patterns.
[0043] The push action agent 126 is configured to query the meta-memory 116 to identify whether a tested hypothesis is consistent with recent data stored in the spatial memory 114. Upon detecting a correlation between the most recent input data and a previously tested hypothesis, the push action agent 126 determines that an associated event condition has been met. In response, the push action agent 126 proactively notifies the main control service 118 to initiate execution of an appropriate action plan. Execution may include retrieving and invoking an appropriate API or AI service from a catalog database to address the identified scenario. The push action agent 126 continuously monitors for matching hypotheses and triggers actions based on real-time data alignment, enabling the control system 102 to initiate intelligent, context-aware actions without explicit user input, thereby improving system autonomy and responsiveness.
[0044] Based on the executed action plan, the main control service 118 may generate output data 130 as an enterprise solution including ethical consideration data. The output data 130 may include a wide range of actionable deliverables to the user, such as strategic reports 130-1 (e.g., performance summaries or operational insights), financial or operational spreadsheets 130-2, automatically generated scripts or software code modules 130-3, or robotic process automation configuration data 130-N. The output data 130 may be formatted and transmitted via designated enterprise interfaces, user dashboards, system APIs, or workflow automation tools depending on the target deployment environment.
[0045] The RAI check service 128 evaluates output data 130 (e.g., enterprise solutions) to ensure alignment with ethical, legal, and enterprise-specific compliance standards. The RAI check service 128 performs validation checks on generated output data 130 by evaluating the output data 130 for issues such as bias, harmful language, false information, or confidential data leaks. The RAI check service 128 also cross-references the output data 130 with predefined ethical parameters and past processing results to detect inconsistencies or deviations from acceptable norms. If anomalies or violations are detected, the RAI check service 128 can correct, filter, or flag the output before it is finalized and delivered. The validation process enhances operational safety and governance by ensuring all outputs are trustworthy, accountable, and aligned with RAI principles.
[0046] 1 discloses only a few components and subsystems, there may be additional components and subsystems not shown, such as, but not limited to, ports, network devices, databases, network-attached storage devices, assets, machines, equipment, facility facilities, emergency management devices, photography devices, cooling devices, heating devices, compressors, any other devices, and combinations thereof. Those skilled in the art should not limit the components / subsystems shown in FIG. 1.
[0047] Those skilled in the art will appreciate that the hardware depicted in FIG. 1 may vary depending on the particular implementation. For example, other peripheral devices, such as optical disk drives and the like, local area network (LAN), wide area network (WAN), wireless (e.g., wireless-fidelity (Wi-Fi)) adapters, Bluetooth adapters, graphics adapters, disk controllers, and input / output (I / O) adapters may be used in addition to or in place of the hardware depicted. The depicted example is provided for purposes of explanation only and is not intended to imply architectural limitations with respect to the present disclosure.
[0048] Those skilled in the art will recognize that, for simplicity and clarity, this specification does not depict or describe the complete structure and operation of every data processing system suitable for use in the present disclosure. Instead, only those portions of control system 102 that are unique to or necessary for an understanding of the present disclosure are depicted and described. The remaining configuration and operation of control system 102 may conform to any of a variety of current implementations and practices known in the art.
[0049] Various examples of generating enterprise solutions are described with reference to FIGS.
[0050] Figure 2 is a high-level process flow illustrating how a control system interprets and executes intent from heterogeneous multi-modal inputs to generate enterprise outputs in accordance with implementations of the present disclosure. Figure 2 is described in conjunction with Figure 1. Process flow 200 may be executed within system 100.
[0051] The process flow 200 includes receiving multimodal input data 104 from input data sources including, but not limited to, text data, audio logs, image logs, and video logs via a main control service 118 that serves as a central orchestration module of the control system 102.
[0052] The short-term memory 110 stores the multimodal input data 104, including text data 204-1, speech-to-text (STT) data and corresponding audio data 204-2, context information and image tag metadata and image data 204-3, and context information and video tag metadata and video data 204-4. The long-term memory contains structured and summarized representations of the raw data originally stored in the short-term memory. The long-term memory 112 includes pointers to summary data and text data 206-1, pointers to STT summary data and audio data 206-2, pointers to context summary data and image data 206-3, and pointers to context summary data and video data 206-N.
[0053] The received multimodal input data 104 is then forwarded to the hippocampal agent 122 via the action executor 120, which functions as a dynamic task dispatch unit. The hippocampal agent 122 may perform feature extraction by analyzing the multimodal input data 104 in combination with past data (prediction history 210) retrieved via the connector 106. The hippocampal agent 122 extracts features, including temporal features, spatial features, contextual features, causal relationships, and entity-related attributes. The features are stored in the spatial memory 114 to construct a structured time-series data representation. The hippocampal agent 122 utilizes a transformer-based subsequent event prediction model 208 integrated with a collective intelligence framework to predict one or more subsequent events. The one or more subsequent events may be predicted based on patterns and trends derived from the extracted features and correlations between past data. The prediction results are returned to the main control service 118 via the action executor 120.
[0054] The main control service 118 then evaluates whether the predicted event requires the execution of an enterprise action or action plan. If execution is required, the main control service 118 retrieves the appropriate external or internal AI model, API, or service from the catalog database 202 and initiates the execution of the action plan. The resulting output data 130, such as a report, spreadsheet, executable code, or robotic process automation (RPA) configuration, may be considered an enterprise solution. The output data 130 may be further validated by the RAI check service 128, which ensures that the generated output or enterprise solution complies with predefined ethical parameters, regulatory standards, and internal governance policies. The RAI check service 128 analyzes the output data 130 to detect potential inconsistencies, illusions, confidential content, or harmful instructions and performs any necessary corrections before final delivery. In parallel, the metamemory agent 124 continuously analyzes historical data and trends stored in the spatial memory 114 to develop and validate one or more candidate hypotheses. Hypotheses are generated by correlating extracted trends with historical data patterns, and each hypothesis is assigned a confidence score based on its degree of consistency with previously tested company outcomes. Tested hypotheses are stored in a structured meta-memory 116 that supports higher-level abstract reasoning.
[0055] Additionally, the push action agent 126 periodically queries the meta-memory 116 to determine whether validated hypotheses are consistent with the most recent input features or data fluctuations. Upon detecting a match indicating a previously confirmed pattern is reoccurring in real time, the push action agent 126 proactively sends a signal to the main control service 118. This signal triggers the selection and execution of a corresponding AI / API module appropriate for the matched scenario, which is also subject to ethical validation by the RAI check service 128. The continuous feedback loop enables the control system 102 to autonomously generate and validate enterprise solutions by reasoning based on multi-modal data, historical trends, and confirmed hypotheses, thereby enhancing operational intelligence and proactive decision-making.
[0056] FIG. 3 is a detailed flowchart of an exemplary method for extracting five W features, bundling trends, and predicting subsequent events using a transformer-based model, according to an embodiment of the present disclosure. For example, raw data is text 302-1. The text 302-1 and data in the short-term memory 110 or short-term memory data are summarized into the long-term memory 112. The long-term memory 112 includes features such as five Ws 304. The five Ws 304 are when, where, what, why, and who. Trends are extracted from the five Ws 304. The extracted trends are passed to the transformer-based subsequent event prediction model 208 to obtain a subsequent event prediction 308. After recording the subsequent event prediction 308 in the spatial memory 114, it is converted into hypotheses and validation data for the wisdom of crowds 306 together with past data accumulated in the spatial memory 114 and stored in the hierarchical meta-memory 116.
[0057] For example, a product is out of stock on a particular day of the week. The out-of-stock products are clustered, e.g., products with the same trend as a previously hypothesized product are picked. The data is sliced by vending machine, product, and time of day (e.g., the processor slices the data by week). A pattern of fluctuations in the inventory of the hypothesized product by time of day is created. Groups of products with similar trends but different trends from the previously hypothesized product are picked.
[0058] The data is sliced by area, product, and time zone (for example, the processor slices the data by week). A pattern of inventory fluctuations by time of day for each area and product is created, and products that have similar trends to products covered by previous hypotheses are picked out. Products that show different trends from those identified in previous hypotheses but exhibit similar trends are identified. The data is sliced by vending machine, product, and time zone, for example, by week. A pattern of fluctuations in inventory levels by time of day is created for each vending machine and product. Based on these patterns, products that have similar trends to previous hypotheses are selected.
[0059] A group of products that have similar trends but different overall trends from the target products of the previous hypothesis is identified. The data is segmented based on these trend-similar products selected in the previous step. Buyer attribute data is collected for each target product, vending machine, and time zone. Information on when, where, and by whom the target products were purchased is collected, along with data on out-of-stocks and backlogs. The data is further divided into product groups that share similarities but different trends from the target products of the previous hypothesis. Buyer attributes are analyzed for each target product, vending machine, and time zone. Furthermore, information on purchasing behavior, timing of out-of-stocks, and backlogs is collected.
[0060] FIG. 4 shows a flowchart of a method 400 for storing prediction results in a prediction history according to an implementation of the present disclosure. For example, assume that raw data is text. First, the text is registered in the short-term memory 110. Further, the registered short-term memory 110 can be summarized and organized to create the long-term memory 112. The long-term memory 112 can be organized by features such as the five Ws (5W1H) to create the spatial memory 114. Non-limiting examples of the features are when, where, what, why, who, and how. Furthermore, trends can be extracted from the spatial memory 114 structured by the features. The extracted trends are bundled, and the latest data from the spatial memory 114 can be input to the transformer-based subsequent event prediction model 208 to predict subsequent events.
[0061] FIG. 5 illustrates a process flow of an exemplary method 500 for bundling extracted trends into a historical storage structure to facilitate training of a transformer-based subsequent event prediction model 208, according to an embodiment of the present disclosure.
[0062] In one implementation, trends are extracted from spatial memory 114. Trends are structured according to feature dimensions such as, but not limited to, where, who, and time. Each trend represents a time-evolving pattern derived from multimodal input data, including text logs, audio input, video data, and historical data.
[0063] Once trends are extracted, they undergo time series clustering 508 to identify groups of trends that exhibit similar temporal dynamics or behavioral signatures. Time series clustering 508 serves as an abstraction 510 of low-level signals into higher-level semantic groups, allowing for the identification of recurring scenarios, seasonal patterns, or causal relationships.
[0064] The resulting clustered data 512 forms the basis for generating structured training data 514, which is then used to train the Transformer-based subsequent event prediction model 208. The Transformer-based subsequent event prediction model 208 is implemented using a Transformer-based architecture and utilizes both contextual embeddings and temporal dependencies to accurately predict likely future events or actions within the enterprise environment.
[0065] In one embodiment, if new features 516 are discovered during time series clustering 508, i.e., attributes that significantly contribute to trend grouping but were not previously part of the spatial memory's feature set, such new features 516 are dynamically added to spatial memory 114. For example, if trend clusters are found to be highly dependent on device type or weather conditions, those features are incorporated into spatial memory 114 as new classification axes, dynamically expanding the features (including the "5 Ws"). Dynamic expansion of the "5 Ws" (Who, What, When, Where, Why) allows system 100 to evolve its dimensional understanding of enterprise events over time. As new contextual signals become relevant, they are incorporated into the trend abstraction pipeline, improving the comprehensiveness of the predictive modeling process.
[0066] The updated trend groups are used to continuously refine the spatial memory 114 and retrain the transformer-based subsequent event prediction model 208, resulting in increased self-improving predictive capabilities. This architecture supports the generation of increasingly accurate, context-aware predictions of future events, which are used to drive autonomous decision-making, enterprise task automation, or human-in-the-loop recommendations.
[0067] 6 is a process flow of an exemplary method 600 for converting bundles of historical trend data into structured hypotheses to create and refine the meta-memory 116, according to an embodiment of the present disclosure. In one implementation, the method 600 begins by accessing previously abstracted and clustered bundled trend data from the spatial memory 114, as described in connection with FIG. 5 .
[0068] Bundles of trends contain recurring patterns extracted from the multimodal input data 104 and are categorized by dimensions such as location 502 (where), actor 504 (who), time 506 (when), action (what), and motivation (why). The method 600 performs time series clustering 602 on these bundles of trends to identify similar historical sequences or groups of behaviors. Time series clustering 602 enables the system 100 to recognize recurring phenomena that are temporally or spatially separated but exhibit statistically or semantically similar patterns. From the groups, candidate hypotheses are automatically generated that represent generalized causal or correlation statements. For example, a hypothesis might be, "People tend to buy soda when they show signs of fatigue." Each such hypothesis is stored as an untested hypothesis in the metamemory 116, which functions as a higher-level cognitive model designed to store abstract knowledge 604. The method 600 then proceeds to verify the validity of each untested hypothesis using the sequences of past events held in the spatial memory 114. Testing involves cross-referencing hypotheses with time-ordered, feature-rich data to assess whether they are supported or contradicted by actual past events.
[0069] Hypotheses that are supported by data are classified as verified hypotheses, while hypotheses without empirical support are classified as rejected hypotheses. Upon classification, the metamemory 116 is updated accordingly. Furthermore, verified hypotheses are given increased confidence and are used as knowledge components to support downstream inferences, predictions, or the initiation of actions. Rejected hypotheses are retained for traceability but are flagged as invalid and can be downgraded over time based on system heuristics or thresholds.
[0070] Additionally, structural feedback is applied to the spatial memory 114. For example, new feature dimensions related to validated hypotheses (e.g., "fatigue" as a trigger) are added to the spatial memory 114, allowing future trends to incorporate and benefit from newly discovered causal factors. Conversely, feature dimensions correlated only with rejected hypotheses can be removed from the spatial memory 114 after a rejection persistence threshold is met, thereby optimizing the memory model and reducing noise. By performing this process iteratively, the system builds a hierarchical metamemory 116 that evolves over time through continuous learning and self-correction. The metamemory 116 serves as a repository of collective intelligence, enabling the control system 102 to reason based on abstract concepts, generalize across various data inputs, and adapt to dynamic enterprise environments.
[0071] FIG. 7 is a process flow diagram of an exemplary computer-implemented method in which the push action agent 126 utilizes validated hypotheses stored in the metamemory 116 and corresponding structured events maintained in the spatial memory 114 to proactively prompt human intervention or automated responses, according to an embodiment of the present disclosure. In one implementation, the method 700 begins when recurring event data, such as repeated customer complaints or inventory shortages, accumulates in the spatial memory 114. This data is structured across multiple axes, including, but not limited to, location, product type, timestamp, and user behavior signals. Over time, consistent patterns emerging from this data are aggregated and abstracted into bundles of trends, which are then evaluated (as described in connection with FIGS. 5 and 6 ) to generate hypotheses. Once a hypothesis is validated against historical patterns (e.g., “Product X runs out every Friday night, causing complaints”), it is promoted to a confirmed hypothesis in the metamemory 116.
[0072] The push action agent 126 operates as a proactive decision-making component. Periodically, or in response to a trigger, the push action agent 126 queries the meta-memory agent 124 to determine whether the latest structured data in the spatial memory 114 reflects a match with any confirmed hypotheses. For example, if current inventory and complaint data indicate signals similar to those associated with a known out-of-stock event, the meta-memory agent 124 confirms that the hypothesis has been re-instantiated by the real-time data. Upon receiving this confirmation, the push action agent 126 notifies the main control service 118, which evaluates whether the matched hypothesis warrants human attention or system-driven remediation. This evaluation is based on predefined thresholds, historical resolution success rates, and ethical or regulatory parameters stored in the system policy memory. If action is warranted, the main control service 118 selects and executes an appropriate API or AI model from the catalog database 202, which may include logistics automation scripts, chatbot-based customer support escalations, supply chain coordination logic, or alert generation to a human administrator. After execution, the generated output 130 is subjected to a responsible AI (RAI) check service 128, which audits the results for compliance with ethical standards, fairness metrics, illusion filtering, and security guidelines. Only if the output 130 passes this validation is it published or implemented via outputs 130-1...N, which may include messages, dashboards, control signals, or automated workflows. In one exemplary scenario, vending machines at a particular location regularly run out of a popular item on Friday nights, causing customer dissatisfaction. The control system 102 slices the historical data by product type, vending machine ID, and time interval.The control system 102 then structures this data to identify patterns of variation, which form the basis of validated hypotheses, such as, for example, "Customers frequently complain about shortages of product X every Friday after 6:00 PM." If this pattern reappears in new data, the push action agent 126 triggers an alert via an API call to the logistics management system to prompt preemptive replenishment. In particular, through repeated application of the event prediction and push action mechanisms, if previously recurring events are resolved (e.g., inventory levels are maintained consistently and complaints cease), the associated event patterns gradually fade from the spatial memory reconstruction. As a result, confirmed hypotheses are demoted and marked as potentially rejectable upon reevaluation, so that only active and relevant hypotheses persist in the metamemory 116. Furthermore, the transformer-based subsequent event prediction model 208 is continuously retrained using updated training data, including the evolving state of confirmed and rejected hypotheses.
[0073] Figure 8 is a flow diagram illustrating an example computer-implemented method 800 for generating enterprise outputs (also referred to as enterprise solutions) according to implementations of the present disclosure. In some implementations, the computer-implemented method 800 may be executed by a processor of the control system 102, as described in connection with Figures 1 and 2. Figure 8 is described in connection with Figures 1-7.
[0074] At step 802, the computer-implemented method 800 may include receiving multimodal input data 104 from various input data sources. The multimodal input data 104 may include one of text data, audio logs, image logs, video logs, and / or the like. The multimodal data is preprocessed to generate structured data or summarized long-term memory data. Using the LLM via the connector 106, the short-term memory data (raw data) is converted into one of summarized and organized long-term memory data. For example, in one implementation, a speech recognition module is used to extract speech-to-text summary data and audio data pointers from audio logs associated with the input data sources. An image processing module is used to extract context summary data and image data pointers from image logs associated with the input data sources. A video analysis module is used to extract context summary data and video data pointers from video logs associated with the input data sources.
[0075] At step 804, the computer-implemented method 800 may include retrieving historical data corresponding to enterprise solutions related to the historical input from a connector that interfaces with multiple artificial intelligence (AI) models. The AI models may include one of an externally generated artificial intelligence model, an internal custom artificial intelligence model, an external data source, and an internal enterprise data source. At step 806, the computer-implemented method 800 may include extracting features from the multimodal input data and the historical data. The features may include one of temporal features, spatial features, contextual features, causal features, and entity-related attributes.
[0076] At step 808, the computer-implemented method 800 may include determining a trend from the extracted features associated with the multimodal input data and the historical data. The trend includes a pattern corresponding to data fluctuations associated with the multimodal input data. The trend is determined by structuring the data based on one of temporal attributes, spatial attributes, and contextual attributes. To determine the trend, in some implementations, intent attributes are interpreted from the multimodal input data received from the input source. An action plan is then formulated based on the interpreted intent attributes and the stored processing results of the short-term memory data and the long-term memory data. The formulated action plan is then dynamically modified based on real-time feedback from the external AI system. Collaboration with each of the external AI systems is then established to align the modified action plan. The modified action plan is executed based on collaboration with each of the external AI systems, and response data is obtained from a hippocampal agent (e.g., hippocampal agent 122) associated with each of the external AI systems. Output data, including ethical consideration data, is then generated based on the executed action plan. The output data generated is verified by cross-referencing it with stored processing results and responses obtained.
[0077] In some implementations, inconsistencies in the enterprise solution are detected to validate the output data. Further, the enterprise solution is corrected based on predefined ethical parameters and historical data patterns. In some implementations, the generated output data, including the ethical consideration data, is validated based on an executed action plan. The validated output data is analyzed for at least one of detecting hallucinations, restricting harmful content, and filtering sensitive information.
[0078] In some implementations, to determine trends, the extracted features are partitioned by one of a time interval, a geographic area, a product, and an entity identifier. To partition the features, groups of multimodal input data having similar data variation patterns are identified. The groups of multimodal input data are correlated with historical data. The consistency of the trends is verified based on the correlation. Once the features are partitioned, changes in the state of the partitioned features over the time interval are aggregated into patterns. Furthermore, key factors, including one of a time zone, a location, a person, and a product, associated with the patterns are identified.
[0079] At step 810, the computer-implemented method 800 may include predicting subsequent events based on the determined trends using a transformer-based large-scale language model (LLM). The predicted subsequent events are stored in a spatial store (e.g., spatial store 114).
[0080] At step 812, the computer-implemented method 800 may include generating an enterprise solution based on the predicted subsequent events. The enterprise solution includes one of a report, a spreadsheet, code, and a robotic process automation configuration.
[0081] Furthermore, in some implementations, the long-term memory data is structured into a spatial memory using the LLM via the connector 106. The spatial memory 114 is organized by features including one of time, place, person, and action. Furthermore, trends are extracted from the historical data and features stored in the spatial memory using time series clustering. The extracted trends are used to train a transformer-based subsequent event prediction model 208 associated with the transformer-based LLM. Subsequent events are predicted based on the extracted trends and multimodal input data. Hypotheses are formulated and verified based on the extracted trends using the historical data stored in the spatial memory 114, thereby generating meta-memory data. An alert is sent to the user to take action based on the verified hypotheses and predicted subsequent events.
[0082] To generate and test multiple hypotheses, candidate hypotheses are generated based on the extracted trends and cross-referenced with historical data patterns stored in spatial memory 114. Each candidate hypothesis is assigned a confidence score based on its consistency with the historical data patterns.
[0083] The present disclosure provides significant advantages by integrating multimodal input data, such as text, audio, images, video, and memory streams, with historical enterprise data to generate intelligent, highly contextual enterprise solutions. The present disclosure enables more accurate event prediction and decision-making through advanced feature extraction and trend analysis across temporal, spatial, and contextual dimensions. The inclusion of ethical validation, real-time collaboration with external AI systems, and the use of meta-memory improves trust, compliance, and transparency. Furthermore, the present disclosure provides the ability to detect hallucinations, correct inconsistencies, and filter sensitive information, ensuring that outputs are trustworthy and actionable, creating a robust framework for dynamic, data-driven enterprise automation and insight generation.
[0084] 9 illustrates a computer system 900 that may be used to implement system 100. More specifically, computing machines such as desktops, laptops, smartphones, tablets, wearables, etc. may be used to generate enterprise solutions. Computer system 900 may include additional components not shown, and some of the described process components may be removed and / or modified. In other examples, computer system 900 may be deployed on an external cloud platform, such as the cloud, an in-house cloud computing cluster, an organization's computing resources, and / or the like.
[0085] Computer system 900 includes processor(s) 902, such as a central processing unit, application specific integrated circuit (ASIC), or other type of processing circuit; input / output devices 904, such as a display, a mouse, and a keyboard; a network interface 906, such as a local area network (LAN), a wireless 802.11x LAN, a 3G or 4G or 5G mobile WAN, or a WiMax WAN; and storage medium(s) 908. Each of these components may be operably coupled to a computer bus 910. Storage medium(s) 908 may be any suitable medium that participates in providing instructions to processor(s) 902 for execution. For example, storage medium(s) 908 may be a non-transitory or non-volatile medium such as a magnetic disk, or a solid-state non-volatile memory or a volatile medium such as RAM. The instructions or modules stored on the storage medium(s) 908 may include machine-readable instructions 912 that are executed by the processor(s) 902 to cause the processor(s) 902 to implement the methods and functions of the system 100.
[0086] System 100 may be implemented as software stored on a non-transitory processor-readable medium and executed by processor(s) 902. For example, storage medium(s) 908 may store an operating system 914, e.g., MAC OS, MS WINDOWS, UNIX, or LINUX, and code for system 100. Operating system 914 may be multi-user, multi-processing, multi-tasking, multi-threaded, real-time, and the like. For example, at run time, operating system 914 operates and code for system 100 is executed by processor(s) 902.
[0087] The computer system 900 may include data storage 916, which may include non-volatile data storage. The data storage 916 stores any data used or generated by the system 100.
[0088] The network interface 906 connects the computer system 900 to internal systems, for example, via a LAN. The network interface 906 may also connect the computer system 900 to the Internet. For example, the computer system 900 may connect to a web browser and other external applications and systems via the network interface 906.
[0089] What has been described and illustrated herein, together with some of its variations, is by way of example only. The terms, descriptions, and figures used herein are set forth for purposes of illustration only and are not meant to be limiting. Many variations are possible within the spirit and scope of the present subject matter, which is intended to be defined by the following claims and their equivalents.
[0090] The implementations and all functional operations described herein may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. The implementations may be realized as one or more computer program products (i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or controlling the operation of a data processing apparatus). The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter providing a machine-readable propagated signal, or one or more combinations thereof. The term "computing system" includes all apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, this apparatus may include code that creates an execution environment for a subject computer program (e.g., code comprising processor firmware, a protocol stack, a database management system, an operating system, or any suitable combination of one or more of these). A propagated signal is an artificially generated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that encodes information to be transmitted to an appropriate receiver apparatus.
[0091] A computer program (also known as a program, software, software application, script, or code) may be written in any suitable form of programming language, including compiled or interpreted languages, and may be deployed in any suitable form, for example, as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple associated files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program may be deployed to be executed on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.
[0092] The processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry (e.g., an FPGA (field programmable gate array) or an ASIC).
[0093] Processors suitable for executing a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any suitable kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. Elements of a computer may include a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices (e.g., magnetic, magneto-optical, or optical disks) for storing data, or be operatively coupled to receive data from or transfer data to them, or both. However, a computer need not have such devices. Furthermore, a computer may be incorporated into another device (e.g., a mobile phone, a personal digital assistant (PDA), a portable audio player, a Global Positioning System (GPS) receiver). Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor(s) 902 and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0094] To provide for interaction with a user, an implementation may be realized on a computer having a display device (e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse, trackball, touchpad) by which the user may provide input to the computer. Other types of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any suitable form of sensory feedback (e.g., visual feedback, auditory feedback, haptic feedback), and input from the user may be received in any suitable form, including acoustic, speech, or tactile input.
[0095] An implementation may be realized in a computing system that includes back-end components (e.g., as a data server), middleware components (e.g., an application server), and / or front-end components (e.g., a client computer having a graphical user interface or web browser through which a user may interact with an implementation), or any suitable combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any suitable form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN") and a wide area network ("WAN"), e.g., the Internet.
[0096] A computing system may include clients and servers. Clients and servers are generally remote from each other and interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0097] While this specification contains numerous details, these should not be construed as limitations on the scope of the disclosure or what may be claimed, but rather as descriptions of features specific to particular implementations. Certain features described herein in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented separately in multiple implementations or in any suitable subcombination. Furthermore, while features may be described above as working in particular combinations and even initially claimed as such, one or more features of a claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to subcombinations or variations of the subcombination.
[0098] Similarly, although operations are shown in a particular order in the figures, this should not be understood as requiring that such operations be performed in the particular order shown, or in a sequential order, or that all of the illustrated operations be performed, to achieve desirable results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the division of various system components in the above-described implementations should not be understood as requiring such division in all implementations, and it should be understood that the described program components and systems may generally be integrated into a single software product or packaged into multiple software products.
[0099] Although several implementations have been described, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. For example, various forms of the flows illustrated above may be used in which steps are rearranged, added, or deleted. Accordingly, other implementations are within the scope of the following claims.
Claims
1. a processor; a memory communicatively coupled to the processor; 1. A system comprising: the memory including processor-executable instructions that, when executed by the processor, cause the processor to: receiving multimodal input data from a plurality of input data sources, the multimodal input data including at least one of text data, audio logs, image logs, and video logs; Retrieving historical data corresponding to an enterprise solution related to a plurality of historical inputs from a plurality of connectors that interface with a plurality of artificial intelligence (AI) models; extracting a plurality of features from the multimodal input data and the historical data, the plurality of features comprising at least one of temporal features, spatial features, contextual features, causal features, and entity-related attributes; determining a plurality of trends from the extracted features associated with the multimodal input data and the historical data, the plurality of trends comprising a plurality of patterns corresponding to a plurality of data variations associated with the multimodal input data, the plurality of trends being determined by structuring data based on at least one of temporal attributes, spatial attributes, and contextual attributes; predicting a subsequent event based on the determined trends using a transformer-based large-scale language model (LLM), wherein the predicted subsequent event is stored in a spatial memory; and generating at least one enterprise solution based on the predicted subsequent events, the at least one enterprise solution including at least one of a report, a spreadsheet, code, and a robotic process automation configuration; A system that allows the following to be performed.
2. To determine the trends from the extracted features associated with the multi-modal input data and the historical data, the processor further comprises: interpreting intent attributes from the multimodal input data received from the plurality of input data sources; Formulating an action plan based on the interpreted intention attributes and the stored processing results of the short-term memory data and the long-term memory data; Dynamically modifying the developed action plan based on real-time feedback from a plurality of external AI systems; coordinating with each external AI system of the plurality of external AI systems to align the modified action plan; Executing the modified action plan based on the collaboration with each of the external AI systems and obtaining response data from a hippocampal agent associated with each of the external AI systems; generating output data including ethical consideration data based on the executed action plan; validating the generated output data by cross-referencing it with the stored processing results and the obtained responses; The system of claim 1 ,
3. To verify the generated output, the processor further comprises: Detecting inconsistencies in the enterprise solution; correcting the enterprise solution based on predefined ethical parameters and historical data patterns; The system of claim 2 , further comprising:
4. The processor further comprises: verifying the generated output data including the ethical consideration data based on the executed action plan; analyzing the validated output data for at least one of hallucination detection, harmful content restriction, and sensitive information filtering; The system of claim 2 , further comprising:
5. To determine the plurality of trends, the processor further comprises: Segmenting the extracted features by at least one of a plurality of time intervals, a plurality of geographic areas, a plurality of products, and a plurality of entity identifiers; aggregating the changes in state of the separated plurality of features over the plurality of time intervals into the plurality of patterns; identifying significant factors including at least one of a time zone, a location, a person, and the product associated with the plurality of patterns; The system of claim 1 ,
6. To separate the extracted features, the processor further comprises: identifying multiple groups of multimodal input data having similar data variation patterns; correlating the groups of multimodal input data with the historical data; verifying trend consistency based on said correlation; The system of claim 5 , further comprising:
7. The processor further comprises: extracting speech-to-text summary data and audio data pointers from the speech log associated with the plurality of input data sources using a speech recognition module; extracting, using an image processing module, context summary data and image data pointers from the image log associated with the plurality of input data sources; extracting context summary data and video data pointers from the video logs associated with the plurality of input data sources using a video analytics module; The system of claim 1 ,
8. The processor further comprises: converting short-term memory data into at least one of summarized long-term memory data and organized long-term memory data using an AI model via the plurality of connectors; structuring the long-term memory data into a spatial memory using the LLM via the plurality of connectors, the spatial memory being organized by a plurality of features including at least one of time, place, person, and action; extracting the plurality of trends from the historical data and the plurality of features stored in the spatial memory using time series clustering, wherein the extracted plurality of trends are used to train a Transformer-based subsequent event prediction model associated with the Transformer-based LLM; predicting the subsequent event based on the extracted trends and the multi-modal input data; generating meta-memory data by formulating and testing a plurality of hypotheses based on the extracted trends using the past data stored in the spatial memory; alerting a user to take action based on the tested hypotheses and the predicted subsequent event; The system of claim 1 ,
9. To generate and test the plurality of hypotheses, the processor further comprises: generating a plurality of candidate hypotheses based on the extracted trends; cross-referencing the plurality of candidate hypotheses with historical data patterns stored in the spatial memory; assigning a confidence score to each of the plurality of candidate hypotheses based on their consistency with the historical data patterns; The system of claim 8 , further comprising:
10. 2. The system of claim 1, wherein the plurality of AI models comprises at least one of an externally generated artificial intelligence model, an internal custom artificial intelligence model, an external data source, and an internal enterprise data source.
11. receiving, by a processor, multimodal input data from a plurality of input data sources, the plurality of input data sources including at least one of text data, audio logs, image logs, and video logs; retrieving, by the processor, historical data corresponding to an enterprise solution associated with a plurality of historical inputs from a plurality of connectors interfacing with a plurality of artificial intelligence (AI) models; extracting, by the processor, a plurality of features from the multimodal input data and the historical data, the plurality of features including at least one of temporal features, spatial features, contextual features, causal features, and entity-related attributes; determining, by the processor, trends from the extracted features associated with the multimodal input data and the historical data, the trends comprising patterns corresponding to data variations associated with the multimodal input data, the trends being determined by structuring data based on at least one of temporal attributes, spatial attributes, and contextual attributes; predicting, by the processor, a subsequent event based on the determined trends using a transformer-based large-scale language model (LLM), wherein the predicted subsequent event is stored in a spatial memory; and generating, by the processor, at least one enterprise solution based on the predicted subsequent event, the at least one enterprise solution including at least one of a report, a spreadsheet, code, and a robotic process automation configuration; A method comprising:
12. Determining the trends from the extracted features related to the multi-modal input data and the historical data includes: interpreting, by the processor, intent attributes from the multimodal input data received from the plurality of input sources; formulating, by the processor, an action plan based on the interpreted intent attributes and the stored processing results of the short-term memory data and the long-term memory data; dynamically modifying, by the processor, the developed action plan based on real-time feedback from a plurality of external AI systems; coordinating, by the processor, with each external AI system of the plurality of external AI systems to align the modified action plan; executing, by the processor, the modified action plan based on the collaboration with each of the external AI systems and obtaining response data from a hippocampal agent associated with each of the external AI systems; generating, by the processor, output data including ethical consideration data based on the executed action plan; validating, by the processor, the generated output data by cross-referencing it with the stored processing results and the obtained responses; The method of claim 11 further comprising:
13. Validating the generated output includes: detecting, by the processor, an inconsistency in the enterprise solution; correcting, by the processor, the enterprise solution based on predefined ethical parameters and historical data patterns; 13. The method of claim 12, further comprising:
14. validating, by the processor, the generated output data including the ethical consideration data based on the executed action plan; analyzing, by the processor, the verified output data for at least one of detecting hallucinations, limiting harmful content, and filtering sensitive information; 13. The method of claim 12, further comprising:
15. Determining the plurality of trends includes: segmenting, by the processor, the extracted features by at least one of a plurality of time intervals, a plurality of geographic areas, a plurality of products, and a plurality of entity identifiers; aggregating, by the processor, changes in state of the separated plurality of features over the plurality of time intervals into the plurality of patterns; identifying, by the processor, significant factors including at least one of a time zone, a location, a person, and the product associated with the plurality of patterns; The method of claim 11 further comprising:
16. Separating the extracted features includes: identifying, by the processor, a plurality of groups of multi-modal input data having similar data variability patterns; correlating, by the processor, the groups of multi-modal input data with the historical data; verifying, by the processor, trend consistency based on the correlation; 16. The method of claim 15, further comprising:
17. extracting, by the processor, speech-to-text summary data and audio data pointers from the speech log associated with the plurality of input data sources using a speech recognition module; extracting, by the processor, context summary data and image data pointers from the image log associated with the plurality of input data sources using an image processing module; extracting, by the processor, context summary data and video data pointers from the video logs associated with the plurality of input data sources using a video analytics module; The method of claim 11 further comprising:
18. converting, by the processor, short-term memory data into at least one of summarized long-term memory data and organized long-term memory data using the plurality of AI models via a plurality of connectors; structuring, by the processor, the long-term memory data into a spatial memory using the LLM via the plurality of connectors, the spatial memory being organized by a plurality of features including at least one of time, place, person, and action; extracting, by the processor, trends from the historical data stored in the spatial store and the features using time series clustering; training, by the processor, a transformer-based subsequent event prediction model associated with the transformer-based LLM using the extracted trends; predicting, by the processor, a subsequent event based on the extracted trends and multimodal input data, wherein the multimodal input data includes at least one of text data, audio data, image data, and video data; and generating, by the processor, meta-memory data by formulating and testing a plurality of hypotheses based on the extracted trends using the past data stored in the spatial memory; alerting a user to take action based on the tested hypotheses and the predicted subsequent event, by the processor; The method of claim 11 further comprising:
19. Formulating and verifying the multiple hypotheses generating, by the processor, a plurality of candidate hypotheses based on the extracted trends; cross-referencing, by the processor, the plurality of candidate hypotheses with historical data patterns stored in the spatial memory; assigning, by the processor, a confidence score to each of the plurality of candidate hypotheses based on their consistency with the historical data patterns; 20. The method of claim 18, further comprising:
20. A non-transitory computer-readable medium containing processor-executable instructions, the processor-executable instructions causing a processor to: receiving multimodal input data from a plurality of input data sources, the plurality of input data sources including at least one of text data, audio logs, image logs, and video logs; Retrieving historical data corresponding to an enterprise solution related to a plurality of historical inputs from a plurality of connectors that interface with a plurality of artificial intelligence (AI) models; extracting a plurality of features from the multimodal input data and the historical data, the plurality of features comprising at least one of temporal features, spatial features, contextual features, causal features, and entity-related attributes; determining trends from the extracted features associated with the multi-modal input data and the historical data, the trends comprising patterns corresponding to data variations associated with the multi-modal input data; predicting a subsequent event based on the determined trends using a transformer-based large-scale language model (LLM), wherein the predicted subsequent event is stored in a spatial memory; and generating at least one enterprise solution based on the predicted subsequent events, the at least one enterprise solution including at least one of a report, a spreadsheet, code, and a robotic process automation configuration; A non-transitory computer-readable medium for causing