System and method for multi-layered graph integration with domain-specific ai agents for disparate data sources

WO2026207116A1PCT designated stage Publication Date: 2026-10-01GILEAD SCIENCES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/020773
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-25
Publication Date
2026-10-01

Smart Images

  • Figure US2026020773_01102026_PF_FP_ABST
    Figure US2026020773_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented system integrates data from disparate sources into a multi¬ layered graph comprising domain-specific layers and an entity layer. The system includes an entity resolution engine configured to reconcile entity references by performing hash-based lookups for definitive identifiers and vector-based similarity searches for partial or ambiguous data. An ingestion pipeline parses incoming data, extracts attributes using neural network models, and updates the graph with new or matched nodes. Domain- specific Al agents, each associated with a respective layer, retrieve relevant data via retrieval-augmented generation (RAG) subsystems and generate insights using fine-tuned large language models (LLMs). A coordinating Al agent receives cross-layer queries, distributes sub-queries to domain-specific agents, and synthesizes a unified response by referencing the entity layer.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 1589-WO-PCT SYSTEM AND METHOD FOR MULTI-LAYERED GRAPH INTEGRATION WITH DOMAIN-SPECIFIC Al AGENTS FOR DISPARATE DATA SOURCESCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 778,639, filed March 27, 2025, the entire contents of which is hereby incorporated by reference in its entirety.BACKGROUND

[0002] Organizations across industries face the challenge of integrating disparate data sources that exist in isolated silos, such as finance, manufacturing, quality management, human resources, and supply chain systems. These silos store information in various formats and structures, ranging from structured databases to free-text documents. Conventional methods of data integration often rely on primary keys or relational joins, which are only effective when clear identifiers are shared between datasets. When such identifiers are absent or when data sources include unstructured descriptions, traditional approaches become inefficient and error-prone.

[0003] Moreover, interpreting data across different domains requires contextual understanding that varies significantly depending on the subject matter. For example, manufacturing data may include machine logs and production records that use specialized terms, while financial data involves invoices and payment records that follow accounting conventions. The lack of a unified framework capable of handling cross-domain queries and drawing connections between related entities across silos impedes organizational insights and decision-making.

[0004] Existing graph-based data integration systems have limitations in handling semantic complexity and domain- specific context. These systems typically treat data as homogeneous nodes and edges, often losing important distinctions between domain- specific relationships. Additionally, traditional data integration frameworks do not support dynamic orchestration of artificial intelligence (Al) agents specialized in understanding domain-specific jargon and logic.

[0005] The present disclosure addresses these limitations by introducing a heterogeneous multilayered network (HMN) framework, where a multi-layered graph structure integrates domainspecific data sources. This structure is coupled with specialized Al agents fine-tuned for each domain, allowing efficient cross-domain reasoning, anomaly detection, and query resolution. Embodiments of the disclosure enable users to seamlessly traverse relationships betweenAttorney Docket No. 1589-WO-PCT different layers, such as tracing a part number from a purchase order in the finance layer to a machine failure in the manufacturing layer and associated quality -control records in the complaint layer. By leveraging retrieval-augmented generation (RAG) systems, embeddingbased query matching, and graph traversal retrievals (can be local or global traversals), the disclosure facilitates both real-time and batch ingestion of structured and unstructured data, ensuring accurate and context-aware query responses.SUMMARY

[0006] This patent application describes a system and method for integrating disparate data sources into a unified, multi-layered graph structure (also referred to as a heterogeneous multilayered network (HMNs)) augmented by specialized artificial intelligence (Al) agents. The purpose of this system is to bridge data silos — such as finance, manufacturing, and quality management — by dynamically connecting entities through descriptive relationships rather than traditional primary keys. The architecture relies on large language models (LLMs) that are tailored to each domain, allowing for nuanced interpretation and querying.

[0007] The core of certain embodiments of the disclosure is the multi-layered graph, where each layer represents a distinct domain (e.g., finance, manufacturing). Within each layer, nodes and edges define domain- specific entities and their relationships, such as invoices linked to employees in the finance layer or machines linked to production batches in the manufacturing layer. In some embodiments, a central “entity layer” connects nodes across domains, unifying references to the same real-world objects (e.g., an employee linked to both a managerial role in the finance layer and an operational role in the manufacturing layer).

[0008] Data ingestion is achieved via pipelines that extract, tokenize, and parse data, identifying key entities and linking them to their corresponding nodes in the graph. This can occur in realtime through streaming solutions such as Apache Kafka or in batch processes. For each data layer, an LLM-bascd Al agent is fine-tuned with domain-specific corpora, enabling it to interpret jargon, acronyms, and contextual nuances. These agents work in tandem with retrieval-augmented generation (RAG) subsystems, which use embedding-based indices to perform semantic similarity searches and retrieve relevant documents or data points. This enables efficient querying, reducing inaccuracies or hallucinations by grounding the model's outputs in specific retrieved data.

[0009] Another notable feature of some embodiments of the disclosure may include the use of a coordinating agent to manage cross-domain queries. When a user query spans multiple domains, the coordinating agent distributes subtasks to the appropriate layer-specific agents, aggregatesAttorney Docket No. 1589-WO-PCT their responses, and synthesizes a coherent answer. This orchestration ensures that domainspecific insights are preserved while enabling a comprehensive cross-domain analysis. For example, a query about defects linked to late raw material deliveries might involve data from the supply chain, manufacturing, and quality management layers, all harmonized through the coordinating agent.

[0010] The system also supports modularity and scalability. Each layer can be updated or expanded independently without disrupting the overall architecture. New data sources or improved LLM models can be integrated seamlessly. Additionally, caching mechanisms improve performance by storing frequently accessed data for faster retrieval.

[0011] Practical applications of this framework include enhanced traceability, anomaly detection, and compliance management in complex environments such as pharmaceutical manufacturing. For instance, the system can detect anomalies or potential fraud through domainspecific neural networks (Al agents) trained on historical data. These networks identify anomalies in real time by scoring incoming data against learned behavior patterns. Upon detecting an anomaly, the domain-specific Al agent generates an alert, which is sent to the coordinating Al agent for cross-layer analysis. The coordinating Al agent verifies the anomaly by querying related layers and consolidating partial results. If confirmed, the system updates the multi-layered graph to freeze suspicious nodes, block related deliveries, and label sources as high-risk. The system also records query parameters, partial responses, and final results, maintaining versioned updates and tagging incomplete information to trigger follow-up queries for further validation.

[0012] The disclosure further extends to multi-modal data by supporting non-textual inputs like images or voice data, which can be indexed and queried similarly to textual information. This adaptability makes the system applicable across various industries requiring robust data integration, reasoning, and operational insights.BRIEF DESCRIPTION OF THE DRAWINGS

[0013] FIG. 1 illustrates an example multi-layered graph or network for integrating disparate data sources that exist in isolated silos, in accordance with some embodiments.

[0014] FIG. 2 illustrates an example multi-layered graph or network configured with domainspecific Al agents, in accordance with some embodiments.

[0015] FIG. 3 illustrates example system architectures for implementing the multi-layered graph or network configured with domain-specific Al agents, in accordance with some embodiments.Attorney Docket No. 1589-WO-PCT

[0016] FIGs. 4A and 4B illustrate an example process performed by the multi-layered graph or network configured with domain- specific Al agents, in accordance with some embodiments.

[0017] FIG. 5 illustrates an example computing system that may be used in implementing various features of embodiments of the disclosed technology.DET ILED DESCRIPTION

[0018] As mentioned in the background section, many organizations have separate “silos” of data, such as finance data, human resources data, manufacturing data, and 5G data, each stored in different formats and domains. Conventionally, linking these data uses relational databases and relies on a primary key or similar identifier to join database tables. However, when data must be connected without obvious shared keys or when the data is represented in free -text descriptions, conventional approaches become inefficient.

[0019] For example, a financial report may reference a purchase order associated with a specific employee ID, while a separate manufacturing database may log equipment usage under a machine operator’s name. If the financial report describes the operator only by their role or job title (e.g., “line supervisor”), rather than by a unique identifier, a traditional database join cannot easily connect these records. This issue becomes more pronounced when linking unstructured text entries, such as complaint logs or audit reports, where the same entity (e.g., a machine, part, or employee) may be described using different terms across domains.

[0020] The present disclosure addresses this issue by building a multi-layered graph that dynamically connects objects cross these data silos based on their descriptive relationships, supported by fine-tuned large language models (LLMs) or “agents” that specialize in each distinct layer.

[0021] FIG. 1 illustrates an example multi-layered graph or network for integrating disparate data sources that exist in isolated silos, in accordance with some embodiments. In this example, the multi-layered graph includes three distinct layers corresponding to three individual research groups: machine and deep learning research, bioscience research, and self-driving car research. These groups may represent different departments within an organization, a community, or even international, cross-country joint research groups. Each layer contains domain-specific objects and their corresponding relationships, which are represented as nodes and edges within that layer.

[0022] In the machine and deep learning research layer, the nodes represent entities such as papers (publications), authors, and labs (experiments or test results). Relationships between theAttorney Docket No. 1589-WO-PCT nodes may include authorship (an author writing one or more papers), lab results being cited across multiple publications, or collaborative research stemming from a shared lab experiment.

[0023] In the bioscience research layer, the nodes include papers, authors, labs, and organizations. The inclusion of organizations in this layer reflects the fact that bioscience research may involve patient data or proprietary experimental data that is often protected and not shared between organizations. The edges within this layer may indicate connections such as joint research between authors from different institutions, citations of lab results across publications, and contributions from multiple labs to a single study.

[0024] In the self-driving car research layer, the nodes represent papers, authors, and organizations. Unlike the bioscience layer, this layer lacks lab-related nodes but includes organizations to account for the collaborative and proprietary nature of self-driving car research, where different entities may independently collect and manage sensor data from their fleet of vehicles. The edges within this layer may represent relationships such as co-authorship between researchers from different organizations or the use of shared datasets referenced in multiple publications.

[0025] Cross-layer connections further enhance the multi-layered graph structure by linking related nodes across different research domains. For instance, an author node in the machine learning layer may also appear in the bioscience and self-driving car layers, indicating that the author has contributed to research in all three domains. Additionally, two authors from different layers may be connected through an indirect relationship, such as having studied under the same professor or collaborating through a shared institutional affiliation.

[0026] The technical improvement of using the multi-layer graph structure can be illustrated using a specific example in FIG. 1. FIG. 1 illustrates, in bold, a case where three authors from three different layers collaborate on a paper represented in the self-driving car research layer. These cross-layer collaborations are captured by connections that preserve the distinct properties of the relationships between the authors across layers.

[0027] If the three layers were merged into a single-layer graph representation (as shown in the right bottom comer of FIG. 1) , these authors and their paper would be represented homogeneously, losing the contextual subjectivity expressed by the original multi-layered structure. Specifically, the nuanced relationships, such as the differing strengths of collaboration between authors in the machine learning and self-driving car research layers versus authors in the bioscicncc and machine learning research layers, would be flattened into indistinguishable links.

[0028] By retaining the heterogeneity and multi-layered nature of the graph, the system can assign distinct attributes (e.g., weights, collaboration frequency, or context-specific relevance) toAttorney Docket No. 1589-WO-PCT cross-layer connections, preserving the unique relationships between nodes across layers. For example, the collaboration between an author in the machine learning research layer and an author in the self-driving car research layer may have a different weight or significance than a collaboration between an author in the machine learning research layer and an author in the bioscience research layer.

[0029] FIG. 2 illustrates an example multi-layered graph or network configured with domainspecific Al agents, in accordance with some embodiments.

[0030] From a structural standpoint, the multi-layered graph (in FIG. 1 or FIG. 2) defines nodes at different levels of abstraction, which may include a manufacturing or process maps layer, a finance or supply chain layer, a complaints or quality-management layer, or any other domainspecific layer.

[0031] As shown in FIG. 2, each layer has intra-layer connectivity and is further connected to an entity layer 200 that captures key concepts or “nouns” that appear in the data, whether those concepts represent a person, a product, an equipment item, a chemical, or another definable term. Once these nodes are linked in a unified graph, users can traverse the relationships to see how a single entity appears in multiple, otherwise unconnected sources. For instance, if a person is mentioned in a financial report regarding purchase orders, which in turn references manufacturing data describing equipment, the graph reveals the cross-domain path from the person to the financial report to the manufacturing data.

[0032] In some embodiments, the multi-layered graph architecture can be implemented using a variety of graph database technologies, such as Neo4j, JanusGraph, or other property graph platforms. In example scenarios, the system ingests data from different sources — finance, manufacturing, quality management, human resources, and so forth — through an extraction pipeline. This pipeline may include natural language processing modules that tokenize and parse incoming text, identifying named entities (people, products, equipment, and other terms describing objects). Once identified, these entities are linked to unique nodes in the entity layer 200 (either linked to the existing nodes, or creating new nodes). In addition, the nodes representing domain-specific data remain in their respective layers — finance layer nodes might correspond to invoices, purchase orders, or line items, while manufacturing layer nodes can represent batches, lots, machines, or production steps. In other words, the nodes from the domain-specific layers (e.g., from the same layer or cross different layers) may be linked to the same node in the entity layer 200, and the node in the entity layer 200 may be labeled with attributes (as vectorized or hashed fingerprints) aggregated from the domain-specific nodes.

[0033] For each layer, edges within each layer indicate relevant relationships within that domain. In a manufacturing layer, edges might capture a “producedBy” or “requires”Attorney Docket No. 1589-WO-PCT relationship between a production process node and a particular machine node. In a finance layer, edges could represent “authorizedBy" relationships between an invoice node and an employee node. All these domain-specific nodes connect to entity-layer nodes for more universal concepts — an “employee” node in the entity layer could be linked to both a “machine operator” node in the manufacturing layer and a “manager authorization record” node in the finance layer if both references describe the same person.

[0034] Data ingestion to the multi-layer graph can be performed in real time or batched. In a real-time scenario, streaming solutions like Apache Kafka or similar technologies may feed the pipeline. As new records arrive, named entity recognition models (e.g., LLMs) identify the key terms, such as references to a raw material or a specific employee ID. The system then searches an entity resolution engine to either map those terms to existing entity-layer nodes or create new ones in the entity layer 200, if necessary. In some embodiments, When new references surface in distinct data sources — for example, a supply chain system logging a defective material code — those nodes can be linked back to existing complaint or quality-management data.

[0035] A technical challenge in implementing the searching and mapping mechanism for the entity layer (e.g., using the entity resolution engine) stems from the fact that traditional hashbased solutions are generally designed for one-to-one lookups using a single, unambiguous key. In many real- world scenarios, however, a single object often has multiple identifying attributes, and searches may not always require exact matches. For example, in a network environment, the same machine might be referenced by its hostname, IP address, or MAC address, each of which could appear in different data sources or logs. To reconcile these disparate references so that they map to the same machine, the registry must store all known attributes for the machine in a single, canonical record. By consolidating these variations within a unified framework, the system can effectively handle multi -attribute lookups and partial matches, ensuring that separate references to the same underlying entity ultimately point to the correct, unified record.

[0036] In some embodiments, the entity resolution engine may use a hybrid lookup mechanism that combines both direct and fuzzy matching strategies. It means the entity resolution engine may include a hash table for exact searches and an embedding database for fuzzy searches. When a new record arrives, the system may first perform a direct search using hash-based indexing if a definitive attribute is present, such as a known MAC address for searching a machine or server. This allows for constant-time lookups in cases where the data source provides an unambiguous identifier. If no exact match is found — or if the new record provides only descriptive attributes like a partially specified hostname — the system moves to a fuzzy matching process. Here, it may embed the machine’s name and other attributes (e.g., network footprints) in a vector space and carry out a similarity search, comparing the incoming data withAttorney Docket No. 1589-WO-PCT existing records in the registry. If the similarity score between the new record and a stored entry exceeds a defined threshold, the system concludes that both refer to the same machine and updates the canonical record accordingly. Conversely, if the best match remains below the threshold, a new record is created.

[0037] In the context of a pharmaceutical company, an example for the entity resolution engine may involve batch records for a drug or raw material used in production. Each batch can be definitively identified by a unique batch ID or described using a combination of attributes, such as the manufacturing date, lot number, supplier name, and production line. When a new batch-related record arrives, the system extracts key fields from the incoming data. If a batch ID is present, the system performs a hash-based lookup using the batch ID to check if the batch already exists in the registry. This hash-based search is highly efficient, enabling constant-time lookups due to the unique and deterministic nature of the batch ID.

[0038] However, if the batch ID is missing or ambiguous, the system resorts to a more flexible, vector-based search by embedding a combination of attributes in the incoming record — such as the manufacturing date, supplier name, and production line — into a high-dimensional vector space. The system computes a vector representation of the incoming record and compares it to the vector representations of existing batch records stored in the registry. The similarity between vectors is scored, and if the highest similarity score exceeds a predefined threshold, the system determines that the new record corresponds to an existing batch. For example, a quality-control report may describe an issue using phrases such as “batch manufactured on June 12 from Supplier A, processed on Line 3” instead of providing the exact batch ID. The embedding-based search matches this description against existing batch records, enabling the system to identify the corresponding batch despite the absence of a precise identifier.

[0039] This dual approach ensures robustness when integrating data from disparate sources, each of which may reference the same batch differently. For instance, the supplier’s system may use a lot number to refer to a batch, while the manufacturing system may reference the same batch using its batch ID and production line information, and the quality management system may describe the batch based on test dates and observations of deviations. By combining hashbased and vector-based searches, the system can reconcile these records into a unified batch entry in the canonical registry, avoiding duplication and providing a comprehensive view of the batch across systems. This hybrid design ensures that records with definitive identifiers are matched efficiently, while partial or descriptive data can still be linked to the correct batch through context-aware matching, preserving the integrity and accuracy of the multi-layered graph.Attorney Docket No. 1589-WO-PCT Domain-specific Al Agents

[0040] As shown in FIG. 2, each domain- specific layer and the entity layer 200 in the multilayer graph is configured with a layer-specific Al agent. In an example, once the multi-layer graph is populated, Al agents can explore it through queries that traverse relationships across domains. A user request might be automatically triaged to the manufacturing or processing layer searching for the history of a specific part number, and to the quality-management layer searching for complaints or deviations involving that part have been recorded. The user request may also trigger the Al agent in the supply chain layer searching for purchase orders, suppliers, or payment details linked to that same part. The triaging may happen sequentially or in parallel.

[0041] In an implementation with agent-based or LLM orchestration, each domain or layer may have its own fine-tuned domain-specific Al agent that understands the jargon and logic of that domain. When a user request spans multiple domains — such as discovering whether certain supplier issues correlate with increased scrap rates in manufacturing — the system’s coordinating agent dispatches subtasks to each domain-specific agent, which queries its portion of the graph and returns specialized summaries.

[0042] In some embodiments, the coordinating agent may be implemented as another Al agent at the entity layer 200. Since the nodes (objects) in the entity layer 200 may be encoded with the cross-layer knowledge (knowing one object is related to multiple objects from multiple layers), the coordinating agent may be able to transform a user request for an object in the entity layer into a plurality of subtasks for the domain-specific layers. The coordinating agent then compiles these summaries into a final, unified response, giving the user a more comprehensive and context-aware answer than would be possible with siloed or disconnected datasets.Domain-specific RAGs

[0043] While these domain- specific Al agents may be fine-tuned using the domain-relevant text corpora, they may still suffer the inherent limitations of LLMs such as hallucinations. To mitigate this issue, each layer’s agent is paired with a layer-specific retrieval-augmented generation (RAG) subsystem that leverages an embedding-based index of domain-relevant documents, records, and metadata. Upon receiving an incoming query or subtask directed to a particular layer — such as a manufacturing-layer query concerning specific production batches — the associated agent may first extract the key terms (e.g., using the fine-tuned LLM), encode these terms using a transformer-based or similar embedding model, and then compute semantic similarity scores between the terms and the indexed corpus. By retrieving the most contextually relevant segments of text or data structures, the agent then transforms the user query based on these retrieved segments, thereby grounding the inference and reducing the likelihood of hallucinations or inaccuracies.Attorney Docket No. 1589-WO-PCT

[0044] In some examples, this architecture with domain-specific Al agent and RAG can be realized with vector databases or hybrid indexing systems that store both traditional inverted indexes and embedding vectors. Such databases enable high-dimensional nearest-neighbor searches for semantic retrieval. To further improve query efficiency, each layer-specific Al agent also maintains a cache of frequently accessed documents or metadata, optimizing system throughput and response time. When a query spans multiple layers, a coordinating process — by the Al agent at the entity layer 200 — distributes sub-queries to each relevant layer’s RAG subsystem. Each layer agent then independently performs retrieval, executes domain-specific reasoning, and generates partial or complete answers. The Al agent at the entity layer 200 aggregates or reconciles these responses to produce a unified, cross-layer analysis that can be returned to the user or ingested by downstream systems.

[0045] By segregating the retrieval logic within each domain-level RAG subsystem, the system preserves modularity and domain specialization. Over time, updated domain data can be appended to or removed from a particular layer’s index without requiring reconfiguration of the entire architecture. This modular approach also allows for incremental improvements in each layer’s fine-tuned language model or RAG engine, ensuring that any enhancements in indexing, embedding generation, or model architecture remain localized to that domain.Layer-specific queries

[0046] After training or fine-tuning is complete, each layer-specific Al agent (LLM) is integrated with its vector database through a retrieval interface. When a user query is received by a layer, the LLM first determines whether the query can be fully satisfied by querying the RAG. This determination is based on analyzing the structure, intent, and required data specificity of the query. If the query pertains directly to information available within the RAG (e.g., structured or semi-structured data indexed in the vector database), the system vectorizes the query and performs a semantic similarity search within the RAG. The top-ranked documents or data segments are retrieved and used to directly construct a response, which is promptly delivered to the user.

[0047] For layer-specific queries requiring additional reasoning or cross-referencing beyond the scope of the RAG’s indexed corpus, the system transforms the query into a more domainspecific form. This transformation involves identifying general-domain terms in the user's input — such as “medication availability” or “staffing levels” — and replacing them with equivalent domain-specific terms that align with the indexed data and fine-tuned LLM’s vocabulary. For instance, a user query might refer to determining trends in medication shortages across multiple regions over a specified time frame, and the transformed prompt would directAttorney Docket No. 1589-WO-PCT the LLM to incorporate both retrieved data and its fine-tuned reasoning capabilities to generate a comprehensive response.

[0048] By maintaining a modular architecture, new documents or data sets can be ingested into the vector index without retraining the LLM. Similarly, the fine-tuned LLM can be updated or replaced independently of the RAG infrastructure, ensuring system scalability and adaptability over time. This dual approach — direct RAG-based retrieval for simple queries and transformed prompt-driven LLM reasoning for complex queries — optimizes both efficiency and accuracy, providing robust domain-specific responses to user queries.Cross-layer queries

[0049] When a cross-layer query arrives, the coordinating agent (e.g., the Al agent at the entity layer 200) first performs a semantic analysis on the query to identify which layers are likely implicated. This may be done by examining specific keywords or entities that belong to each domain, such as references to raw material delivery in a manufacturing or supply chain layer, alongside mentions of product complaints in a quality management layer. Based on the detected semantics, the coordinating agent distributes sub-queries to each of the relevant layers. Each layer-specific subsystem then conducts an internal retrieval step, leveraging its vector-based index to find the most pertinent documents or data points. Next, the subsystem either returns raw or partially processed information to the coordinating agent or, if necessary, applies its own finetuned LLM to inteipret and produce structured insights from the retrieved data.

[0050] Once the coordinating agent receives partial results from each implicated layer, it applies a cross-layer query transformation if the user’s query requires integrated reasoning that spans different domains. During this transformation, domain-specific terminologies are aligned to ensure that information from separate layers can be accurately matched and correlated. For example, if the supply chain layer reports that a particular raw material was delivered late, and the quality management layer logs that the same material was associated with elevated defect rates, the coordinating agent will reconcile the internal identifiers and correlate the overlapping entities. This process often entails referencing the entity layer 200 that stores canonical references for critical terms like batch numbers, machine IDs, or part codes. By anchoring each domain’s output to these canonical references, the coordinating agent creates a unified perspective of the data.

[0051] Having aggregated and harmonized the sub-query results, the coordinating agent may invoke a specialized cross-layer language model, or it can compile the data into a cohesive response. The choice depends on whether further reasoning is required. For instance, if the user asks a question that demands historical trend analysis, root-cause inference, or a predictive forecast, then a higher-level reasoning component might integrate the individual outputs,Attorney Docket No. 1589-WO-PCT drawing on historical precedents or domain-specific logic. In contrast, if the query simply requests a consolidated set of facts — such as the number of delayed raw material shipments correlated with corresponding defect reports — the aggregated data might suffice without additional inference.

[0052] From a performance standpoint, the system is designed to operate efficiently by processing sub-queries in parallel. Each layer-specific RAG subsystem independently performs its vector-based retrieval, allowing the coordinating agent to receive multiple partial results concurrently. The coordinating agent then merges and reconciles these partial outputs as soon as they become available. In cases where additional contextual details are needed — for example, if the manufacturing layer returns incomplete data regarding a batch number — the coordinating agent can issue follow-up sub-queries. This incremental refinement ensures that the final response is both accurate and complete, without placing unnecessary load on every layer from the outset.

[0053] By employing this hierarchical coordination and query transformation approach, the platform efficiently resolves queries that span multiple data layers. Users benefit from a comprehensive view of how disparate data sources interrelate, and the system remains both modular and scalable, as newly added layers or updated data ingestion pipelines can be integrated without disrupt i ng the existing architecture.GraphRAG

[0054] In addition to the vector-based layer-specific RAGs discussed above, some embodiments of the present disclosure may incorporate graph-based RAG or GraphRAG. The GraphRAG may be deployed independent from the layer-specific RAGs or in combination with the layerspecific RAGs. The GraphRAG may support intra-layer graph-based retrieval as well as crosslayer graph-based retrieval.

[0055] In some embodiments, the system implements GraphRAG at both the intra-layer and cross-layer levels to facilitate structured, relationship-driven information retrieval. Each domainspecific layer, such as manufacturing, supply chain, finance, or quality management, maintains a knowledge graph that represents objects and their relationships. Within a given layer, nodes may represent domain-specific entities such as production batches, machines, employees, raw materials, or supplier records, while edges capture relationships like dependencies, usage history, process flows, or ownership. This structured representation allows the system to retrieve relevant information not only based on direct matches but also by following logical connections between related entities.

[0056] In an example, when a domain-specific agent processes an incoming query or subtask, it may determine whether a graph-based traversal is more appropriate than a text-similarityAttorney Docket No. 1589-WO-PCT lookup. For instance, if a query seeks to uncover relationships among suppliers, manufacturing processes, and product defect reports, the agent can traverse the corresponding nodes and edges in its domain layer’s knowledge graph. This traversal allows the agent to resolve multi-hop relationships, such as connecting a delayed raw material shipment to a specific production batch that used the material and then tracing that batch to elevated defect rates recorded in the qualitymanagement layer. By contrast, a direct vector-based RAG approach might only retrieve semantically similar documents but fail to explicitly reveal these relational paths. Graph traversal enables a deeper, contextual understanding of dependencies within the same domain, ensuring that queries uncover root causes, cascading effects, and hidden correlations rather than isolated pieces of information.

[0057] Since each layer in the multi-layer graph system represents a distinct domain, the system also needs to support cross-layer graph-based retrieval, where a coordinating agent at the entity layer (200) unifies references across multiple subgraphs. Each domain maintains its own structured knowledge graph (i.e., the subgraphs), but inter-layer edges or canonical identifiers (i.e., the global graph), such as batch numbers, supplier codes, machine IDs, or part references, enable relationships between layers to be navigated seamlessly. For example, if the knowledge graph in the manufacturing layer links equipment failures to certain product batches, and the supply chain layer links raw materials to suppliers, then the entity layer can unify these references. This enables GraphRAG queries that traverse multiple subgraphs to answer crossdomain questions. A query investigating supplier-related defects could retrieve a history of supplier shipments from the supply chain layer, production batch outcomes from the manufacturing layer, and quality-control deviations from the quality-management layer.

[0058] In some embodiments, GraphRAG and the vector-based RAG can coexist within each domain layer. In some cases, a request may involve retrieving unstructured text (e.g., equipment maintenance reports in plain text) via vector embeddings, then linking those documents to specific nodes in the domain knowledge graph. Conversely, a user might initiate a graph-based query but require additional text context from external documents. In these scenarios, the domain- specific agent can combine graph traversal with semantic similarity searches, ensuring a comprehensive and context-rich response.

[0059] In an example, a user in a pharmaceutical setting may issue a request: “Identify’ common factors in recent production deviations and determine whether any of these factors overlap with supply chain anomalies. ” In response, the system can leverage GraphRAG to retrieve and correlate relevant information across multiple layers.

[0060] The manufacturing layer identifies and connects production deviations, mapping them to associated machines and relevant staff involved in the processes. Simultaneously, the supplyAttorney Docket No. 1589-WO-PCT chain layer analyzes raw materials linked to each production deviation, flagging any suppliers with quality-control issues or delayed shipments. Once these insights are gathered, the coordinating agent at the entity layer aligns references between the two domains, such as matching batch numbers and supplier IDs, to uncover cross-domain dependencies.Practical applications

[0061] In a practical implementation, one or more neural network models can be trained at the supply chain layer (or another suitable domain layer) to detect anomalies across shipments, finance, or manufacturing processes. For instance, the supply chain layer’s Al agent may rely on a deep learning model — trained on historical data sets of raw material deliveries, shipment sizes, and supplier reliability metrics — to learn typical behavior patterns under normal operating conditions. During training, the neural network captures correlations among features such as shipment frequency, shipment volume, supplier track records, and seasonality. The training data is processed via forward propagation, and the model parameters are iteratively updated during back-propagation until convergence. Once trained, the model is deployed in the supply chain layer' s Al agent to score incoming shipments in real time, flagging those that deviate significantly from learned norms.

[0062] When this trained model flags an unusually large shipment from a newly onboarded supplier, the supply chain layer’ s Al agent designates it as a potential anomaly and notifies the coordinating agent at the entity layer. The coordinating agent then queries related nodes and edges across multiple layers — for instance, batch usage logs and quality-control data in the manufacturing layer, and purchase orders or payment records in the finance layer. If these crosslayer sub-queries confirm the anomaly (e.g., incomplete lot documentation in manufacturing or mismatched payment amounts in finance), the system classifies the shipment as high-risk or fraudulent, triggering a broadcast of the anomaly across the various domains. The coordinating agent also updates the multi-layer graph to reflect the anomaly status, effectively freezing the use of the suspect raw material and blocking future deliveries from that supplier.

[0063] In some cases, the multi-layered graph infrastructure empowers the pharmaceutical company (or another enterprise) to take decisive action by marking or “freezing” the affected inventories in the manufacturing layer. The system automatically updates relevant nodes and edges in real time, preventing further use of that raw material. Simultaneously, the supply chain layer issues a directive to suspend additional shipments from the flagged supplier, pending further investigation. Because these cross-domain updates arc captured within the unified graph, authorized personnel — whether in quality assurance, procurement, or finance — can instantly see the anomaly and the resulting freeze in the shared dashboard.Attorney Docket No. 1589-WO-PCT

[0064] Beyond identifying fraudulent or suspect shipments, the described system may also be used to support robust monitoring of production deviations and change controls. For instance, the manufacturing layer can record specific events — equipment malfunctions, process deviations, or instances of product nonconformance — as nodes or edges tied to relevant manufacturing processes. Each of these events can be tagged with metadata (e.g., timestamps, equipment identifiers, affected product batches, severity levels), creating a network of interlinked records that capture a detailed history of production incidents. Meanwhile, a connected quality-management layer can log the corresponding change controls — such as revised standard operating procedures, updated safety protocols, or newly instituted corrective and preventive actions — and link these records to the originating events in the manufacturing layer.

[0065] By organizing these production and quality-management data sources in separate but interconnected layers, the multi-layer graph allows the system’s coordinating agent to share updates in real time. Whenever a new equipment malfunction is detected in the manufacturing layer, the system triggers a cross-layer notification that prompts the quality -management layer to create or update change control documentation. This real-time connectivity helps decisionmakers and operators immediately see if a malfunction has recurred multiple times, identify whether certain repairs or CAPAs were already put in place, and track the overall effectiveness of remedial actions over time.

[0066] Because each production event and each change control record resides in its own domain-specific layer, the hierarchical agent structure can analyze cause-and-effect relationships within and across layers. For example, an agent in the manufacturing layer can automatically correlate repeated deviations on a particular machine with an elevated product defect rate, while the quality -management layer flags incomplete or overdue repairs related to that machine. The coordinating agent at the entity layer consolidates these findings, giving personnel a unified view of where persistent breakdowns are occurring, which interventions have been attempted, and whether those interventions resolved the underlying issue.

[0067] FIG. 3 depicts several example system architectures for implementing the multi-layered graph or network with domain-specific Al agents, in accordance with some embodiments. As illustrated, these multi-layer Al agents can be organized in cooperative, competitive, mixed, or hierarchical frameworks. In the embodiments described herein, the Al agents follow a hierarchical structure, with the leading Al agent located at the entity layer (200 in FIG. 2).

[0068] In addition, there are multiple ways to facilitate collaboration between large language models (LLMs) and Al agents, such as Collaborative Problem- solving Data Exchange (CPDE) or Distributed Problem- solving Data Exchange (DPDE). Under CPDE, a single LLM mayAttorney Docket No. 1589-WO-PCT communicate with and instruct multiple Al agents. Under DPDE, each Al agent is coupled with its own specialized LLM. In some embodiments, the Al agents either do not communicate with each other or share a common memory or RAG subsystem. In the present disclosure, the domain-specific layer Al agents communicate solely with the Al agent residing in the entity layer, thereby preserving a clear hierarchical flow of information and control.Technical Improvement

[0069] The embodiments described in this application leverage the concept of heterogeneous multi-layered graphs or networks (HMN) to create a powerful framework for enterprise data integration, augmented by domain-specific LLMs. This combination addresses the inherent complexity of modern enterprise systems, where disparate data domains operate with unique acronyms, terminologies, and processes. The HMN model enables the representation of these diverse data sources as distinct yet interconnected layers, preserving their contextual and semantic nuances.

[0070] A critical improvement lies in the ability of certain embodiments to integrate heterogeneity and multi-layered structures. Traditional networks struggle with context retention, often merging diverse entities into a single, homogeneous framework, thereby losing essential layer-specific associations. For example, in scientific research, the HMN model can represent distinct layers for machine learning, bioscience, and self-driving cars, each with its own types of nodes — papers, authors, labs, or organizations — and directed interconnections that preserve cross-disciplinary citations. The HMN framework ensures that these relationships remain intact, allowing for a nuanced understanding of inter-layer dependencies, such as collaboration trends across disciplines or shared resource allocations.

[0071] In the context of enterprise applications, an embodiment introduces a hierarchical, multi-agent architecture where each layer is paired with one or more fine-tuned LLMs. These agents interpret and process data within their specific domain, such as financial data in accounting principles or operational data in manufacturing workflows. For instance, an agent specialized for a finance layer can accurately process numeric data and monetary terms, while a manufacturing agent understands domain-specific acronyms and procedural terminologies. This layer-specific intelligence ensures that data integration respects the semantic integrity of each domain, mitigating the risks of misinterpretation or generic results that lack contextual relevance.

[0072] An embodiment of the disclosur also incorporates a coordination agent or master orchestrator, enabling seamless collaboration among the domain-specific agents. This orchestrator routes cross-layer queries, aggregates results from individual layers, and synthesizes a unified response. For example, a query investigating the root causes of product defects mightAttorney Docket No. 1589-WO-PCT involve retrieving data from a manufacturing layer (deviation logs), a supply chain layer (raw material delays), and a quality management layer (complaints and CAPAs). The coordination agent harmonizes these inputs, resolving semantic conflicts across layers and presenting a coherent, multi-faceted analysis to the user or feeding the results into automated decisionmaking systems.

[0073] By embedding domain-specific LLMs within the HMN framework, some embodiments significantly enhance anomaly detection, compliance management, and operational optimization. For instance, in a pharmaceutical environment, the system can link complaints about a raw material to manufacturing deviations, vision-system-detected defects, and historical corrective actions, creating a comprehensive traceability map. Furthermore, by enabling agents to collaborate and reason across domains, the system facilitates predictive insights, such as identifying patterns of recurring quality issues tied to specific suppliers or production processes.

[0074] FIGs. 4A and 4B illustrate an example process performed by the multi-layered graph or network configured with domain- specific Al agents, in accordance with some embodiments. In some implementations, one or more process blocks of FIG. 4 may be performed by a device.

[0075] As shown in FIGs. 4A and 4B, process 400 may include storing, in a graph storage, a multi-layer network that may include a plurality of domain-specific layers and an entity layer, where each domain-specific layer may include domain nodes and intra-layer edges representing relationships within a respective domain, and the entity layer may include entity nodes that are linkable to the domain nodes across the plurality of domain-specific layers via cross-layer edges (block 402). For example, the device may store, in a graph storage, a multi-layer network that may include a plurality of domain-specific layers and an entity layer, where each domainspecific layer may include domain nodes and intra-layer edges representing relationships within a respective domain, and the entity layer may include entity nodes that are linkable to the domain nodes across the plurality of domain-specific layers via cross-layer edges, as described above.

[0076] As also shown in FIGs. 4A and 4B, process 400 may include executing domain-specific queries using a plurality of domain- specific Al agents, each domain- specific Al agent being associated with one of the domain- specific layers and having: a fine-tuned Large Language Model (LLM) trained on domain-relevant data; and a domain-specific retrieval-augmented generation (RAG) subsystem configured to retrieve relevant data from at least one of the disparate data sources (block 404). For example, the device may execute domain-specific queries using a plurality of domain- specific Al agents, each domain- specific Al agent being associated with one of the domain- specific layers and having: a fine-tuned large language model (LLM) trained on domain-relevant data; and a domain- specific retrieval-augmented generationAttorney Docket No. 1589-WO-PCT (RAG) subsystem configured to retrieve relevant data from at least one of the disparate data sources, as described above.

[0077] As further shown in FIG. 4, process 400 may include receiving a cross-domain query by a coordinating Al agent associated with the entity layer, where the cross-domain query references one or more entity nodes in the entity layer (block 406). For example, the device may receive a cross-domain query by a coordinating Al agent associated with the entity layer, where the cross-domain query references one or more entity nodes in the entity layer, as described above.

[0078] As also shown in FIGs. 4A and 4B, process 400 may include converting, by the coordinating Al agent, the cross-domain query into a plurality of domain- specific subtasks by: identifying the relevant domain-specific layers based on the entity nodes and corresponding cross-layer edges; and constructing the domain-specific subtasks, each subtask referencing domain-relevant keywords or entities extracted from the cross-domain query (block 408). For example, the device may convert, by the coordinating Al agent, the cross-domain query into a plurality of domain-specific subtasks by: identifying the relevant domain- specific layers based on the entity nodes and corresponding cross -layer edges; and constructing the domain- specific subtasks, each subtask referencing domain-relevant keywords or entities extracted from the cross-domain query, as described above.

[0079] As further shown in FIG. 4, process 400 may include distributing the domain-specific subtasks to the respective domain-specific Al agents for execution (block 410). For example, the device may distribute the domain-specific subtasks to the respective domain-specific Al agents for execution, as described above.

[0080] As also shown in FIGs. 4A and 4B, process 400 may include retrieving partial responses by the coordinating Al agent from the domain-specific Al agents, where each domain-specific Al agent processes its respective subtask by: retrieving contextually relevant information from its domain-specific RAG subsystem; and generating a partial response using the fine-tuned LLM (block 412). For example, the device may retrieve partial responses by the coordinating Al agent from the domain-specific Al agents, where each domain- specific Al agent processes its respective subtask by retrieving contextually relevant information from its domain-specific RAG subsystem; and generating a partial response using the fine-tuned LLM, as described above.

[0081] As further shown in FIG. 4, process 400 may include aggregating, by the coordinating Al agent, the partial responses from the domain-specific Al agents by referencing the entity nodes in the entity layer and resolving cross -layer references (block 414). For example, the device may aggregate, by the coordinating Al agent, the partial responses from the domain- specific AlAttorney Docket No. 1589-WO-PCT agents by referencing the entity nodes in the entity layer and resolving cross-layer references, as described above.

[0082] As also shown in FIGs. 4A and 4B, process 400 may include generating a unified response to the cross-domain query based on the aggregated partial responses (block 416). For example, the device may generate a unified response to the cross-domain query based on the aggregated partial responses, as described above.

[0083] Process 400 may include additional implementations, such as any single implementation or any combination of implementations described below and / or in connection with one or more other processes described elsewhere herein. A first implementation, process 400 further includes generating multiple identifying attributes for each entity node in the entity layer, each identifying attribute being associated with at least one hash-based identifier and one or more vector-based embeddings of an object represented by the entity node; and upon receiving newly extracted attributes from an ingestion pipeline, searching the newly extracted attributes against stored the identifying attributes to determine whether to update an existing entity node or to create a new entity node based on the newly extracted attributes.

[0084] In a second implementation, alone or in combination with the first implementation, the domain-specific Al agents further may include: a domain-specific cache for frequently accessed records, thereby reducing retrieval latency during repeated queries.

[0085] A third implementation, alone or in combination with the first and second implementation, process 400 further includes recording, by the coordinating Al agent, query parameters, sub-query distributions, partial responses, and final synthesized responses during the execution of cross-domain queries; storing, by the coordinating Al agent, versioned updates to the entity layer to maintain a record of entity node changes and cross-layer references; and tagging, by the coordinating Al agent, conflicting or incomplete information in the aggregated partial responses for generating follow-up domain-specific queries.

[0086] A fourth implementation, alone or in combination with one or more of the first through third implementations, process 400 further includes training, at each domain-specific layer, a neural network as the domain-specific Al agent on historical domain-specific data, where the training includes forward propagation of input features extracted from historical system logs or transactions through multiple layers of the neural network, back-propagation to adjust parameters of the neural network, and the training continues until convergence; deploying the trained neural network at the domain-specific layer to score incoming domain- specific data in real time, thereby identifying over-threshold deviations from learned behavior patterns as anomalies; and generating an anomaly alert using the deployed neural network upon detectingAttorney Docket No. 1589-WO-PCT an anomaly, where the anomaly alert is routed to the coordinating Al agent in the entity layer for cross-layer broadcasting.

[0087] Although FIGs. 4 A and 4B show example blocks of process 400, in some implementations, process 400 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in FIG. 4. Additionally, or alternatively, two or more of the blocks of process 400 may be performed in parallel.

[0088] Where components or modules of the application are implemented in whole or in part using software, in one embodiment, these software elements can be implemented to operate with a computing or processing module capable of carrying out the functionality described with respect thereto. One such example computing module is shown in FIG. 5. Various embodiments are described in terms of this example computing module 500. After reading this description, it will become apparent to a person skilled in the relevant art how to implement the application using other computing modules or architectures.

[0089] Referring now to FIG. 5, computing module 500 may represent, for example, computing or processing capabilities found within desktop, laptop, notebook, tablet, cloud and edge, computers: hand-held computing devices (tablets, PDA’s, smart phones, cell phones, palmtops, etc.); mainframes, supercomputers, workstations or servers; or any other type of special-purpose or general-purpose computing devices as may be desirable or appropriate for a given application or environment. Computing module 500 might also represent computing capabilities embedded within or otherwise available to a given device. For example, a computing module might be found in other electronic devices such as, for example, digital cameras, navigation systems, cellular telephones, portable computing devices, modems, routers, WAPs, terminals and other electronic devices that might include some form of processing capability.

[0090] Computing module 500 might include, for example, one or more processors, controllers, control modules, or other processing devices, such as a processor 504. Processor 504 might be implemented using a general-purpose or special-purpose processing engine such as, for example, a microprocessor, controller, or other control logic. In the illustrated example, processor 504 is connected to a bus 502, although any communication medium can be used to facilitate interaction with other components of computing module 500 or to communicate externally. The bus 502 may also be connected to other components such as a display, input devices, or cursor control to help facilitate interaction and communications between the processor and / or other components of the computing module 500.

[0091] Computing module 500 might also include one or more memory modules, simply referred to herein as main memory 508. For example, preferably random-access memory (RAM) or other dynamic memory might be used for storing information and instructions to beAttorney Docket No. 1589-WO-PCT executed by processor 504. Main memory 508 might also be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 504. Computing module 500 might likewise include a read only memory (“ROM”) or other static storage device 510 coupled to bus 502 for storing static information and instructions for processor 504.

[0092] Computing module 500 might also include one or more various forms of information storage devices 510, which might include, for example, a media drive 512 and a storage unit interface 520. The media drive 512 might include a drive or other mechanism to support fixed or removable storage media 514. For example, a hard disk drive, a floppy disk drive, a magnetic tape drive, an optical disk drive, a CD, DVD or Bluray drive (R or RW), or other removable or fixed media drive 512 might be provided. Accordingly, storage media 514 might include, for example, a hard disk, a floppy disk, magnetic tape, cartridge, optical disk, a CD or DVD, or other fixed or removable medium that is read by, written to or accessed by media drive 512. As these examples illustrate, the storage media 514 can include a computer usable storage medium having stored therein computer software or data.

[0093] In alternative embodiments, information storage devices 510 might include other similar instrumentalities for allowing computer programs or other instmetions or data to be loaded into computing module 500. Such instrumentalities might include, for example, a fixed or removable storage unit 522 and a storage unit interface 520. Examples of such storage units and storage unit interfaces can include a program cartridge and cartridge interface, a removable memory (for example, a flash memory or other removable memory module) and memory slot, a PCMCIA slot and card, and other fixed or removable storage units and interfaces that allow software and data to be transferred from the storage unit to computing module 500.

[0094] Computing module 500 might also include a communications interface or network interface(s). Communications or network interface(s) interface might be used to allow software and data to be transferred between computing module 500 and external devices. Examples of communications interface or network interface(s) might include a modem or soft modem, a network interface (such as an Ethernet, network interface card, WiMedia, WiFi, IEEE 502.XX or other interface), a communications port (such as for example, a USB port, IR port, RS232 port Bluetooth® interface, or other port), or other communications interface. Software and data transferred via communications or network interface(s) might typically be carried on signals, which can be electronic, electromagnetic (which includes optical) or other signals capable of being exchanged by a given communications interface. These signals might be provided to communications interface via a channel. This channel might carry signals and might be implemented using a wired or wireless communication medium. Some examples of a channelAttorney Docket No. 1589-WO-PCT might include a phone line, a cellular link, an RF link, an optical link, a network interface, a local or wide area network, and other wired or wireless communications channels.

[0095] In this document, the terms “computer program medium” and “computer usable medium” are used to generally refer to transitory or non-transitory media such as, for example, memory 508, ROM, and storage unit interface 520. These and other various forms of computer program media or computer usable media may be involved in carrying one or more sequences of one or more instructions to a processing device for execution. Such instructions embodied on the medium, are generally referred to as “computer program code” or a “computer program product” (which may be grouped in the form of computer programs or other groupings). When executed, such instructions might enable the computing module 500 to perform features or functions of the present application as discussed herein.

[0096] Various embodiments have been described with reference to specific exemplary features thereof. It will, however, be evident that various modifications and changes may be made thereto without departing from the broader spirit and scope of the various embodiments as set forth in the appended claims. The specification and FIGs are, accordingly, to be regarded in an illustrative rather than a restrictive sense.

[0097] Although described above in terms of various exemplary embodiments and implementations, it should be understood that the various features, aspects and functionality described in one or more of the individual embodiments are not limited in their applicability to the particular embodiment with which they are described, but instead can be applied, alone or in various combinations, to one or more of the other embodiments of the present application, whether or not such embodiments are described and whether or not such features are presented as being a part of a described embodiment. Thus, the breadth and scope of the present application should not be limited by any of the above-described exemplary embodiments.

[0098] Terms and phrases used in the present application, and variations thereof, unless otherwise expressly stated, should be construed as open ended as opposed to limiting. As examples of the foregoing: the term “including” should be read as meaning “including, without limitation” or the like; the term “example” is used to provide exemplary instances of the item in discussion, not an exhaustive or limiting list thereof; the terms “a” or “an” should be read as meaning “at least one,” “one or more” or the like; and adjectives such as “conventional,” “traditional,” “normal,” “standard,” “known” and terms of similar meaning should not be construed as limiting the item described to a given time period or to an item available as of a given time, but instead should be read to encompass conventional, traditional, normal, or standard technologies that may be available or known now or at any time in the future.Likewise, where this document refers to technologies that would be apparent or known to one ofAttorney Docket No. 1589-WO-PCT ordinary skill in the art, such technologies encompass those apparent or known to the skilled artisan now or at any time in the future.

[0099] The presence of broadening words and phrases such as “one or more,” “at least,” “but not limited to” or other like phrases in some instances shall not be read to mean that the narrower case is intended or required in instances where such broadening phrases may be absent. The use of the term “module” does not imply that the components or functionality described or claimed as part of the module are all configured in a common package. Indeed, any or all of the various components of a module, whether control logic or other components, can be combined in a single package or separately maintained and can further be distributed in multiple groupings or packages or across multiple locations.

[0100] Additionally, the various embodiments set forth herein are described in terms of exemplary block diagrams, flow charts and other illustrations. As will become apparent to one of ordinary skill in the art after reading this document, the illustrated embodiments and their various alternatives can be implemented without confinement to the illustrated examples. For example, block diagrams and their accompanying description should not be construed as mandating a particular architecture or configuration.

Claims

Attorney Docket No. 1589-WO-PCT CLAIMS:What is claimed is:

1. A system for integrating data from a plurality of disparate data sources into a multilayered graph, facilitating domain-specific and cross-domain data mining, the system comprising:at least one processor;a non-transitory computer-readable medium storing instructions executable by the at least one processor; anda graph storage storing a multi-layer network that comprises a plurality of domainspecific layers and an entity layer, wherein each domain-specific layer comprises domain nodes and intra-layer edges, and the entity layer comprises entity nodes linkable to the domain nodes across the plurality of domain-specific layers;wherein:each of the plurality of domain-specific layers comprises a domain-specific Al agent for executing domain-specific queries and tasks, wherein the domain-specific Al agent comprises a fine-tuned Large Language Model (LLM) and a domain-specific retrieval-augmented generation (RAG) subsystem,the entity layer comprises a coordinating Al agent for executing cross-domain queries, wherein the coordinating Al agent is configured to:convert the cross-domain queries into domain-specific subtasks,distribute the domain-specific subtasks to the domain-specific Al agents, and aggregate partial responses from the domain-specific Al agents to generate responses to the cross-domain queries.

2. The system of claim 1 , further comprising an entity resolution engine coupled to the entity layer, wherein the entity resolution engine is configured to:generate multiple identifying attributes for each entity node in the entity layer, each identifying attribute being associated with at least one hash-based identifier and one or more vector-based embeddings of an object represented by the entity node; andupon receiving newly extracted attributes from an ingestion pipeline, search the newly extracted attributes against stored the identifying attributes to determine whether to update an existing entity node or to create a new entity node based on the newly extracted attributes.Attorney Docket No. 1589-WO-PCT 3. The system of claim 1, wherein the domain-specific RAG subsystems arc configured to:maintain a vector-based index of domain-relevant documents or data records: embedding a received domain-specific query into a searching vector;perform semantic similarity searches using the searching vector against the vector-based index coupled with entities based on graph connectivity; andretrieve contextually relevant text segments for use by the respective fine-tuned LLM.

4. The system of claim 1, wherein the domain-specific Al agents further comprise: a domain- specific cache for frequently accessed records, thereby reducing retrieval latency during repeated queries.

5. The system of claim 1, wherein the coordinating Al agent at the entity layer is configured to:extract entity attributes from a cross-domain query;search, using an entity resolution engine, relevant entity nodes in the entity layer based on the extracted entity attributes, to determine domain-specific entities;construct a plurality of sub-queries based on the domain-specific entities; and distribute the plurality of sub-queries to the domain-specific Al agents corresponding to the domain-specific entities for concurrent processing.

6. The system of claim 1 , wherein the coordinating Al agent, upon receiving aggregated partial responses from the domain-specific Al agents, is further configured to:record query parameters, sub-query distributions, partial responses, and final synthesized responses;store versioned updates to the entity layer; andtags conflicting or incomplete information for generating follow-up domain-specific queries.

7. The system of claim 1, wherein at least one domain-specific Al agent is configured to detect anomalies or potential fraud in a respective domain as well as perform anomaly detection algorithms capable of being run on graphs, and the processor is configured to:train a neural network on historical domain- specific data, the training including forward propagation of input features extracted from historical system logs or transactions throughAttorney Docket No. 1589-WO-PCT multiple layers of the neural network, followed by back-propagation to adjust parameters, until convergence;deploy the trained neural network to score incoming domain-specific data in real time, identifying over-threshold deviations from learned behavior patterns as anomalies; and generate, using the deployed neural network, an anomaly alert upon detecting an anomaly, wherein the anomaly alert is routed to the coordinating Al agent in the entity layer for cross-layer broadcasting.

8. The system of claim 7, wherein, upon receiving an anomaly alert from the domainspecific Al agent, the coordinating Al agent is further configured to:confirm the anomaly through cross-layer queries;identify one or more nodes related to the confirmed anomaly;label and freeze nodes related to the detected anomaly by updating the multi-layered graph to stop ongoing usage or consumption, wherein the labeled and frozen nodes represent suspicious materials, batches, or other domain-relevant objects that are related to the detected anomaly; andblock future deliveries or services associated with the labeled and frozen nodes, by marking a corresponding supplier or entity node in the multi-layered graph as high-risk or fraudulent and updating corresponding edges or attributes to restrict subsequent transactions.

9. The system of claim 8, wherein to verify the anomaly through cross-layer queries, the coordinating Al agent is configured to:identify an entity node in the entity layer that is related to the anomaly; determine, based on cross-layer edges of the entity node, one or more domain nodes in one or more domain-specific layers that are related to the entity node;construct one or more queries directed to the domain-specific Al agents of the one or more domain-specific layers, each query including context or keywords derived from the anomaly alert;transmit the constructed queries to the domain-specific Al agents, wherein each domainspecific Al agent is configured to:retrieve relevant domain data from the corresponding RAG subsystem using a vector-based index, andprocess the retrieved domain data using the fine-tuned LLM to generate partial anomaly-assessment results; andAttorney Docket No. 1589-WO-PCT receive the partial anomaly-assessment results from each domain-specific Al agent and correlate the partial anomaly-assessment results with the entity node, thereby determining whether the detected anomaly is confirmed based on cross-layer consistency or discrepancies within the partial anomaly-assessment results.

10. The system of claim 1, wherein the domain-specific Al agent at a domain-specific layer is configured to:extract object attributes from a new domain-specific record or query using the fine-tuned LLM;in response to determining the object attributes do not correspond to any domain node in the domain-specific layer, create at least one new domain nodes based on the object attributes; andconstruct a prompt to the coordinating Al agent in the entity layer to register the at least one new domain nodes.

11. A computer-implemented method for integrating data from a plurality of disparate data sources into a multi-layered graph, facilitating domain-specific and cross-domain data mining, the method comprising:storing, in a graph storage, a multi-layer network that comprises a plurality of domainspecific layers and an entity layer, wherein each domain-specific layer comprises domain nodes and intra-layer edges representing relationships within a respective domain, and the entity layer comprises entity nodes that are linkable to the domain nodes across the plurality of domainspecific layers via cross-layer edges;executing domain-specific queries using a plurality of domain- specific Al agents, each domain-specific Al agent being associated with one of the domain-specific layers and comprising:a fine-tuned Large Language Model (LLM) trained on domain-relevant data; and a domain- specific retrieval-augmented generation (RAG) subsystem configured to retrieve relevant data from at least one of the disparate data sources;receiving a cross-domain query by a coordinating Al agent associated with the entity layer, wherein the cross-domain query references one or more entity nodes in the entity layer;converting, by the coordinating Al agent, the cross-domain query into a plurality of domain-specific subtasks by:identifying the relevant domain-specific layers based on the entity nodes and corresponding cross-layer edges; andAttorney Docket No. 1589-WO-PCT constructing the domain- specific subtasks, each subtask referencing domainrelevant keywords or entities extracted from the cross-domain query;distributing the domain-specific subtasks to the respective domain- specific Al agents for execution;retrieving partial responses by the coordinating Al agent from the domain-specific Al agents, wherein each domain- specific Al agent processes its respective subtask by:retrieving contextually relevant information from its domain-specific RAG subsystem; andgenerating a partial response using the fine-tuned LLM;aggregating, by the coordinating Al agent, the partial responses from the domain-specific Al agents by referencing the entity nodes in the entity layer and resolving cross-layer references; andgenerating a unified response to the cross-domain query based on the aggregated partial responses.

12. The method of claim 11, further comprising:generating multiple identifying attributes for each entity node in the entity layer, each identifying attribute being associated with at least one hash-based identifier and one or more vector-based embeddings of an object represented by the entity node; andupon receiving newly extracted attributes from an ingestion pipeline, searching the newly extracted attributes against stored the identifying attributes to determine whether to update an existing entity node or to create a new entity node based on the newly extracted attributes.

13. The method of claim 11, wherein the domain-specific Al agents further comprise: a domain-specific cache for frequently accessed records, thereby reducing retrieval latency during repeated queries.

14. The method of claim 11, further comprising:recording, by the coordinating Al agent, query parameters, sub-query distributions, partial responses, and final synthesized responses during the execution of cross-domain queries;storing, by the coordinating Al agent, versioned updates to the entity layer to maintain a record of entity node changes and cross-layer references; andtagging, by the coordinating Al agent, conflicting or incomplete information in the aggregated partial responses for generating follow-up domain-specific queries.Attorney Docket No. 1589-WO-PCT15. The method of claim 11, further comprising:training, at each domain-specific layer, a neural network as the domain- specific Al agent on historical domain-specific data, wherein the training includes forward propagation of input features extracted from historical system logs or transactions through multiple layers of the neural network, back-propagation to adjust parameters of the neural network, and the training continues until convergence:deploying the trained neural network at the domain-specific layer to score incoming domain-specific data in real time, thereby identifying over-threshold deviations from learned behavior patterns as anomalies; andgenerating an anomaly alert using the deployed neural network upon detecting an anomaly, wherein the anomaly alert is routed to the coordinating Al agent in the entity layer for cross-layer broadcasting.

16. A non-transitory computer-readable storage medium configured with instructions executable by one or more processors to cause the one or more processors to perform operations comprising:storing, in a graph storage, a multi-layer network that comprises a plurality of domainspecific layers and an entity layer, wherein each domain- specific layer comprises domain nodes and intra-layer edges representing relationships within a respective domain, and the entity layer comprises entity nodes that are linkable to the domain nodes across the plurality of domainspecific layers via cross-layer edges;executing domain-specific queries using a plurality of domain-specific Al agents, each domain- specific Al agent being associated with one of the domain-specific layers and comprising:a fine-tuned Large Language Model (LLM) trained on domain-relevant data; and a domain-specific retrieval-augmented generation (RAG) subsystem configured to retrieve relevant data from at least one of the disparate data sources;receiving a cross-domain query by a coordinating Al agent associated with the entity layer, wherein the cross-domain query references one or more entity nodes in the entity layer;converting, by the coordinating Al agent, the cross-domain query into a plurality of domain-specific subtasks by:identifying the relevant domain-specific layers based on the entity nodes and corresponding cross-layer edges; andconstructing the domain- specific subtasks, each subtask referencing domainrelevant keywords or entities extracted from the cross-domain query;Attorney Docket No. 1589-WO-PCT distributing the domain-specific subtasks to the respective domain- specific Al agents for execution;retrieving partial responses by the coordinating Al agent from the domain-specific Al agents, wherein each domain-specific Al agent processes its respective subtask by:retrieving contextually relevant information from its domain-specific RAG subsystem; andgenerating a partial response using the fine-tuned LLM;aggregating, by the coordinating Al agent, the partial responses from the domain-specific Al agents by referencing the entity nodes in the entity layer and resolving cross-layer references; andgenerating a unified response to the cross-domain query based on the aggregated partial responses.

17. The non-transitory computer- readable storage medium of claim 16, wherein the operations further compromise:generating multiple identifying attributes for each entity node in the entity layer, each identifying attribute being associated with at least one hash-based identifier and one or more vector-based embeddings of an object represented by the entity node; andupon receiving newly extracted attributes from an ingestion pipeline, searching the newly extracted attributes against stored the identifying attributes to determine whether to update an existing entity node or to create a new entity node based on the newly extracted attributes.

18. The non-transitory computer- readable storage medium of claim 16, wherein the domain-specific Al agents further comprise:a domain-specific cache for frequently accessed records, thereby reducing retrieval latency during repeated queries.

19. The non-transitory computer- readable storage medium of claim 16, wherein the operations further compromise:recording, by the coordinating Al agent, query parameters, sub-query distributions, partial responses, and final synthesized responses during the execution of cross-domain queries;storing, by the coordinating Al agent, versioned updates to the entity layer to maintain a record of entity node changes and cross-layer references; andAttorney Docket No. 1589-WO-PCT tagging, by the coordinating Al agent, conflicting or incomplete information in the aggregated partial responses for generating follow-up domain-specific queries.

20. The non-transitory computer-readable storage medium of claim 16, wherein the operations further compromise:training, at each domain-specific layer, a neural network as the domain- specific Al agent on historical domain-specific data, wherein the training includes forward propagation of input features extracted from historical system logs or transactions through multiple layers of the neural network, back-propagation to adjust parameters of the neural network, and the training continues until convergence;deploying the trained neural network at the domain-specific layer to score incoming domain-specific data in real time, thereby identifying over-threshold deviations from learned behavior patterns as anomalies; andgenerating an anomaly alert using the deployed neural network upon detecting an anomaly, wherein the anomaly alert is routed to the coordinating Al agent in the entity layer for cross-layer broadcasting.