Computer-implemented method, computer program and computer system (dynamic adaptation of resources to edge nodes)
By dynamically adapting RAG data distribution to edge nodes using AI-based models, the method addresses inefficiencies and latency issues in edge computing, optimizing performance and reducing overhead.
Patent Information
- Application Number
- JP2025111115
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-22
- Filing Date
- 2025-06-30
- Publication Date
- 2026-02-03
AI Technical Summary
Conventional systems face inefficiencies and increased network traffic due to overprovisioning of search augmentation generation (RAG) data to edge locations to avoid training inaccurate large language models (LLMs), leading to high latency and computational overhead.
A computer-implemented method that dynamically adapts the amount and type of RAG data sent to edge nodes based on real-time performance metrics and conditions using AI-based models, selectively transmitting only relevant subsets.
Improves performance at each edge node and enhances system efficiency by ensuring only relevant RAG data is sent, reducing latency and computational overhead.
Smart Images

Figure 2026016315000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to edge nodes, and more particularly, the present invention relates to transmitting resources to edge nodes. [Background technology]
[0002] Data generation is increasing the overhead associated with data management and processing. AI has been developed in an attempt to mitigate this increase in processing overhead, but advances in AI are also causing machine learning models to become more complex. Increasingly complex machine learning models lead to more intensive workloads and increased load associated with applying the models to received data. The operation of traditional implementations is thereby negatively impacted.
[0003] Cloud computing has been implemented in an attempt to improve the ability to perform computationally intensive operations and process increasing amounts of data. For example, cloud locations can be adjusted to provide dynamic levels of computational throughput that adjust to meet client needs. While this is effective in preventing processing bottlenecks from occurring, it involves transmitting all data to be analyzed to a centralized location, such as a data center or public cloud location. Sending data to a centralized location exposes it to unwanted attacks and unintentional mishandling, thereby significantly increasing the risk of data loss.
[0004] In an attempt to curb this reliance on networks to perform all processing at a central location, edge computing has been implemented to extend computing to endpoints within a system. For example, applications and other types of computational operations are moved to edge locations where data is generated, for the benefit of data privacy and security. However, this also introduces unresolved inefficiencies into traditional products. Summary of the Invention [Problem to be solved by the invention]
[0005] Conventional products are forced to over-provision search augmentation generation (RAG) data to edge locations in an attempt to avoid training and deploying inaccurate large language models (LLMs) at edge locations. [Means for solving the problem]
[0006] According to one approach, a computer-implemented method (CIM) includes receiving information in real time from an edge node, where the information outlines specific search expansion generation (RAG) data to be applied at the edge node as well as conditions at the edge node. A knowledge database mapping RAG data embeddings to various edge node conditions is further updated using the received information. The received information and the knowledge database are then dynamically evaluated using one or more trained artificial intelligence (AI)-based models. A subset of relevant RAG data is output using the AI-based models. The subset of relevant RAG data is then transmitted to the edge node.
[0007] According to another approach, a computer program product (CPP) comprises a set of one or more computer-readable storage media, the CPP also comprising program instructions collectively stored on the set of one or more storage media for causing a set of processors to perform the following computer operations:
[0008] According to yet another approach, a computer system (CS) comprises a set of processors and a set of one or more computer-readable storage media. The CS also comprises program instructions collectively stored on a set of one or more storage media for causing the set of processors to execute the CIM described above.
[0009] Other aspects and implementations of the present invention will become apparent from the following detailed description, which, when taken in conjunction with the drawings, illustrate by way of example the principles of the invention. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram of a computing environment according to one approach.
[0011] [Figure 2A] A representative diagram of a distributed system according to one approach.
[0012] [Figure 2B] FIG. 1 is a representative diagram of dynamically adapting RAG data sent to edge nodes according to one approach.
[0013] [Figure 3A-1] 1 is a flow chart of a method according to one approach. [Figure 3A-2] 1 is a flow chart of a method according to one approach.
[0014] [Figure 3B]3B is a flowchart of sub-operations of one of the operations in the method of FIG. 3A according to one approach.
[0015] [Figure 3C] 1 is a flowchart for predicting future edge node conditions and / or performance metrics according to one approach.
[0016] [Figure 4] FIG. 1 is a representative diagram of a distributed system with use cases. DETAILED DESCRIPTION OF THE INVENTION
[0017] The following description is made for the purpose of illustrating the general principles of this invention and is not intended to limit the inventive concepts claimed herein. Moreover, particular features described herein can be used in combination with other described features in each of the various possible combinations and permutations.
[0018] Unless otherwise specifically defined in this specification, all terms are to be given their broadest possible interpretation, including the meanings implied by this specification and the meanings understood by a person skilled in the art and / or the meanings defined in dictionaries, treatises, etc.
[0019] It should also be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless otherwise specified. It will be further understood that the terms "comprises" and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0020] The following description discloses several preferred approaches to systems, methods, and computer program products for dynamically adapting the amount and / or type of supplies provided to edge nodes. For example, approaches herein involve dynamically determining the amount and / or type of RAG data to send to edge nodes over time based at least in part on real-time performance at the edge nodes themselves. This results in improved performance at each edge node and more efficient operation of the system as a whole, e.g., as described in further detail below.
[0021] In the following description, techniques may be disclosed that improve the efficiency with which multitask model tuning (or "multitask fine tuning") may be performed. It should be understood that the various techniques herein may be implemented with a wide variety of multitask model tuning types, including, for example, multitask prompt tuning, multitask prefix tuning, etc., or any other type of multitask model tuning that will be apparent to one of ordinary skill in the art after reading this specification. To provide context and simply assist the reader, various techniques may be described with reference to a type of multitask model tuning. For example, many techniques are described in the context of multitask prompt tuning (MPT). This is done by way of example only and should not be considered limiting.
[0022] In one general approach, a CIM comprises receiving information in real time from an edge node, the information outlining specific RAG data applied at the edge node as well as conditions at the edge node. A knowledge database mapping RAG data embeddings to various edge node conditions is further updated using the received information. One or more trained AI-based models are then used to dynamically evaluate the received information and the knowledge database. The AI-based models are then used to output a subset of relevant RAG data. The subset of relevant RAG data is then transmitted to the edge node.
[0023] Thus, the techniques herein desirably can dynamically adapt the amount and / or type of supplies provided to edge nodes. For example, the techniques herein involve dynamically determining the amount and / or type of RAG data to send to edge nodes over time based at least in part on real-time performance at the edge nodes themselves. As a result, only the relevant portions of RAG data are sent to each edge node, improving performance at each edge node and causing the system as a whole to operate more efficiently.
[0024] In some implementations, the CIM further includes receiving performance metrics from an edge node. Edge node condition information corresponding to the edge node may also be received therefrom. In response to receiving the performance metrics and / or the edge node condition information from the edge node, the one or more trained AI-based models dynamically evaluate the performance metrics, the edge node condition information, the received information, and the knowledge database. The trained AI-based models also output a subset of the associated RAG data.
[0025] By evaluating performance metrics and / or condition information corresponding to a given edge node in addition to information outlining the specific RAG data applied at the edge node, the techniques herein can dynamically adjust the resources sent to the edge node at a more granular level. In other words, this additional information received from the edge node is used to develop a more detailed understanding of which resources (e.g., RAG data) are relevant (e.g., useful) to the edge node, which is further used to adjust the resources sent to the edge node in real time.
[0026] In some implementations, the knowledge database is formed by dividing the base data into segments and converting the segments into embeddings. The embeddings are then combined with respective information to form a vector storage pool for each embedding. The techniques herein can thus easily search the vector storage pool and use the information therein to train the AI-based model. Furthermore, as the data in the vector storage pool changes over time, the AI-based model can be retrained to incorporate new information and provide groupings of relevant resources (e.g., RAG data) to edge nodes over time.
[0027] In some implementations, the operations are performed by a centralized edge orchestrator. In response, a subset of the relevant RAG data is sent to the edge node along a RAG pipeline extending between the edge node and the centralized edge orchestrator. The centralized edge orchestrator may also be connected to other edge nodes along the same or different RAG pipelines and may selectively send resources to other edge nodes (e.g., in parallel and simultaneously) in a similar manner.
[0028] In some implementations, the CIM further includes having one or more generative AI models predict future edge node conditions, and the one or more trained AI-based models are used to dynamically evaluate the future edge node conditions and the knowledge database, and the trained AI-based models further output a subset of RAG data with expected relevance.
[0029] By determining which resources are expected to be relevant at some point in the future, the techniques herein enable more accurate distribution of resources to remote locations (e.g., edge nodes) than previously possible. An AI model can be trained in other ways by applying a predetermined training dataset to learn how to predict future edge node conditions and / or performance metrics. For example, an AI model can be trained to evaluate future edge node conditions and a knowledge database and output a subset of RAG data with expected relevance. In response, the subset of available RAG data predicted to be relevant to future conditions and / or performance metrics is actually output by the trained AI-based model.
[0030] In some implementations, the subset of RAG data, along with the expected relevance output by the trained AI-based model, influences the RAG data currently being sent to the edge node. The CIM further includes replacing at least a portion of the subset of relevant RAG data with at least a portion of the subset of RAG data. The remainder of the subset of relevant RAG data and the subset of RAG data are then sent to the edge node. The trained AI-based model can thus perform complex evaluations of information on the fly and generate detailed results, which are further used to dynamically adjust resources sent to the edge node on the fly to adapt to conditions occurring at the edge node.
[0031] In another general approach, a CPP comprises a set of one or more computer-readable storage media, the CPP also comprising program instructions collectively stored on the set of one or more storage media for causing a set of processors to perform any combination of the methodologies described above.
[0032] In yet another general approach, a CS comprises a set of processors and a set of one or more computer-readable storage media, the CS also comprising program instructions collectively stored on the set of one or more storage media for causing the set of processors to perform any combination of the methodologies described above.
[0033] In some implementations, a centralized edge orchestrator is connected to one or more remote edge nodes along one or more resource pipelines. In response to receiving a request from one of the edge nodes (e.g., from an application running thereon), the centralized edge orchestrator transmits resources (e.g., RAG data) to the edge node. Information describing how the resources (e.g., RAG data) transmitted to the edge node are actually used at the edge node is returned to the centralized edge orchestrator. In response, the centralized edge orchestrator can evaluate the received information to determine how the transmitted resources are being utilized. The centralized edge orchestrator can adjust the resources transmitted to the edge node on the fly to dynamically adapt to conditions occurring at the edge node. Furthermore, the centralized edge orchestrator can manage any desired number of edge nodes in this manner, thereby significantly improving performance and utilization of available system resources.
[0034] Various aspects of the present disclosure are described by text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in CPP techniques. For any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in a given flowchart. For example, depending again on the technology involved, two operations shown in successive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in an at least partially overlapping manner.
[0035] Computer program product methodology (“CPP methodology” or “CPP”) is a term used in this disclosure to describe any set of one or more storage media (also referred to as “media”), collectively contained in a set of one or more storage devices, that collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A “storage device” is any tangible device that can hold and store instructions for use by a computer processor. The computer-readable storage medium may be, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on a major surface of a disk), or any suitable combination of the foregoing. Computer-readable storage media, as the term is used in this disclosure, is not to be construed as storage in the form of transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses passing through fiber optic cables, electrical signals communicated through wires, and / or other transmission media.As will be appreciated by those skilled in the art, data is typically moved at some infrequent time during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, but this does not make the storage device temporary because the data is not temporary while it is stored.
[0036] Computing environment 100 includes an example environment for the execution of at least a portion of the computer code involved in performing the methods of the present invention, such as the improved relevance determination code at block 150 for dynamically adapting the amount and / or type of supplies provided to edge nodes. For example, the techniques herein involve dynamically determining the amount and / or type of RAG data to send to edge nodes over time based at least in part on real-time performance at the edge nodes themselves. As a result, only the relevant portions of RAG data are sent to each edge node, improving performance at each edge node and causing the system as a whole to operate more efficiently, e.g., as described in further detail below.
[0037] In addition to block 150, computing environment 100 includes, for example, computer 101, wide area network (WAN) 102, end user device (EUD) 103, remote server 104, public cloud 105, and private cloud 106. In this approach, computer 101 includes a set of processors 110 (including processing circuitry 120 and cache 121), a communications fabric 111, volatile memory 112, persistent storage 113 (including operating system 122 and the above-identified block 150), a set of peripheral devices 114 (including a set of user interface (UI) devices 123, storage 124, and a set of Internet of Things (IoT) sensors 125), and a network module 115. Remote server 104 includes a remote database 130. Public cloud 105 includes a gateway 140, a cloud orchestration module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.
[0038] Computer 101 may take the form of a desktop computer, a laptop computer, a tablet computer, a smartphone, a smartwatch or other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device now known or later developed that is capable of executing programs, accessing a network, or querying a database, such as remote database 130. As is well understood in the field of computer technology, and depending on the technology, execution of a computer-implemented method may be distributed among multiple computers and / or multiple locations. However, in this presentation of computing environment 100, to keep the presentation as concise as possible, the detailed discussion focuses on a single computer, specifically computer 101. Although computer 101 is not shown in FIG. 1 within a cloud, it may be located within a cloud. However, computer 101 is not required to reside within a cloud except to any extent that may be expressly indicated.
[0039] Processor set 110 includes one or more computer processors of any type now known or to be developed in the future. Processing circuitry 120 may be distributed across multiple packages, e.g., multiple tailored integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within the processor chip package and is typically used for data or code that should be available for fast access by threads or cores executing on processor set 110. Cache memory is typically organized into multiple levels depending on relative proximity to the processing circuitry. Alternatively, some or all of the cache for a processor set may be located “off-chip.” In some computing environments, processor set 110 may be designed to operate with qubits and perform quantum computing.
[0040] Computer-readable program instructions are typically loaded onto computer 101 and cause processor set 110 of computer 101 to execute a series of operational steps, thereby enabling a computer-implemented method, such that the instructions so executed instantiate the methods specified in the computer-implemented method flowcharts and / or descriptions contained herein (collectively referred to as the "methods of the present invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by processor set 110 to control and direct the execution of the methods of the present invention. In computing environment 100, at least some of the instructions for executing the methods of the present invention may be stored in block 150 within persistent storage 113.
[0041] Communications fabric 111 is the signal-conducting pathway that allows various components of computer 101 to communicate with one another. Typically, this fabric is made up of switches and conductive pathways, such as switches and conductive pathways that make up buses, bridges, physical input / output ports, and the like. Other types of signal communication pathways may be used, such as fiber optic communication pathways and / or wireless communication pathways.
[0042] Volatile memory 112 may be any type of volatile memory now known or later developed. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, although this is not required unless expressly indicated. In computer 101, volatile memory 112 is located in a single package and is internal to computer 101; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located external to computer 101.
[0043] Persistent storage 113 is any form of non-volatile storage for a computer, now known or later developed. The non-volatility of this storage means that stored data is maintained regardless of whether power is supplied to computer 101 and / or to persistent storage 113 directly. Persistent storage 113 can be read-only memory (ROM), but typically at least a portion of persistent storage allows data to be written, data to be erased, and data to be rewritten. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. Operating system 122 can take several forms, including various known proprietary operating systems or open-source Portable Operating System Interface-type operating systems that employ a kernel. The code contained in block 150 typically includes at least some of the computer code involved in performing the methods of the present invention.
[0044] The peripheral device set 114 includes a set of peripheral devices of the computer 101. Data communication connections between the peripheral devices and other components of the computer 101 can be implemented in various ways, such as Bluetooth® connections, Near-Field Communication (NFC) connections, connections made by cable (such as a universal serial bus (USB)-type cable), insertion-type connections (e.g., a secure digital (SD) card), connections made through a local area communication network, and even connections made through a wide area network such as the Internet. In various approaches, the UI device set 123 can include components such as a display screen, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. The storage 124 can be external storage such as an external hard drive or insertable storage such as an SD card. The storage 124 can be persistent and / or volatile. In some approaches, storage 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In approaches where computer 101 is required to have a large amount of storage (e.g., computer 101 stores and manages large databases locally), then this storage may be provided by a peripheral storage device designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple, geographically distributed computers. IoT sensor set 125 is made up of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer and another sensor may be a motion detector.
[0045] Network module 115 is a collection of computer software, hardware, and firmware that enables computer 101 to communicate with other computers over WAN 102. Network module 115 may include hardware such as a modem or Wi-Fi® signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for communicating data over the Internet. In some approaches, the network control and network forwarding functions of network module 115 are performed on the same physical hardware device. In other approaches (e.g., approaches utilizing Software-Defined Networking (SDN)), the control and forwarding functions of network module 115 are performed on physically separate devices, such that the control function manages multiple different network hardware devices. Computer-readable program instructions for implementing the methods of the present invention may be downloaded to computer 101 from an external computer or external storage device, typically through a network adapter card or network interface included in network module 115.
[0046] WAN 102 is any wide area network (e.g., the Internet) capable of communicating computer data over non-local distances using any technology for communicating computer data now known or later developed. In some approaches, WAN 102 may be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi network. WANs and / or LANs typically include copper transmission cables, optical fiber transmissions, wireless transmissions, and computer hardware such as routers, firewalls, switches, gateway computers, and edge servers.
[0047] End-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 101) and may take any of the forms described above with respect to computer 101. EUD 103 typically receives useful and useful data from the operation of computer 101. For example, in the hypothetical case where computer 101 is designed to provide recommendations to the end user, the recommendations would typically be communicated from network module 115 of computer 101 over WAN 102 to EUD 103. In this manner, EUD 103 can display or otherwise present the recommendations to the end user. In some approaches, EUD 103 may be a client device, such as a thin client, a heavy client, a mainframe computer, a desktop computer, and the like.
[0048] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 may be controlled and used by the same entity that operates computer 101. Remote server 104 represents a machine that collects and stores useful and useful data for use by other computers, such as computer 101. For example, in the hypothetical case where computer 101 is designed and programmed to provide recommendations based on historical data, then this historical data may be provided to computer 101 from remote database 130 of remote server 104.
[0049] A public cloud 105 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computer capabilities, particularly data storage (cloud storage) and computing power, without direct active management by users. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. Direct active management of public cloud 105 computing resources is performed by computer hardware and / or software in cloud orchestration module 141. The computing resources provided by public cloud 105 are typically implemented by virtual computing environments running on various computers comprising host physical machine set 142, which is the universe of physical computers within and / or available to public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 143 and / or containers from container set 144. It is understood that these VCEs can be stored as images and can be transferred among and between various physical machine hosts, either as images or after instantiation of the VCEs. Cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of VCE, and manages active instantiations of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that enables public cloud 105 to communicate over WAN 102.
[0050] Some further discussion of virtualized computing environments (VCEs) is now provided. A VCE can be stored as an "image." A new, active instance of a VCE can be instantiated from an image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of an operating system in which the kernel allows the existence of multiple isolated user space instances, called containers. These isolated user space instances typically behave as actual computers from the perspective of programs running within them. A computer program running on a typical operating system can utilize all of the computer's resources, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and of the devices assigned to the container; this feature is known as containerization.
[0051] A private cloud 106 is similar to a public cloud 105, except that its computing resources are available only for use by a single enterprise. While the private cloud 106 is shown in communication with the WAN 102, in other approaches, the private cloud may be completely disconnected from the Internet and accessible only through a local / private network. A hybrid cloud is a composite of multiple clouds of different types (e.g., private, community, or public cloud types), often each implemented by a different vendor. While each of the multiple clouds remains a separate, discrete entity, the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the constituent clouds. In this approach, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.
[0052] Cloud computing services and / or microservices (not separately shown in FIG. 1 ): Private and public clouds 106 are programmed and configured to deliver cloud computing services and / or microservices (unless otherwise indicated, the word “microservices” should be interpreted to include larger “services” regardless of size). Cloud services are typically infrastructure, platforms, or software hosted by a third-party provider and made available to users over the Internet. Cloud services facilitate the flow of user data from front-end clients (e.g., user-side servers, tablets, desktops, laptops) over the Internet to the provider's systems and vice versa. In some approaches, cloud services can be composed and orchestrated according to the “as a service” technology paradigm, where something is presented to internal or external customers in the form of cloud computing services. As-a-service offerings typically provide endpoints with which various customers interface. These endpoints are typically based on a set of APIs. One category of as-a-service offerings is Platform as a Service (PaaS), where a service provider provisions, instantiates, runs, and manages modular bundles of code that customers can use to instantiate a computing platform and one or more applications without the complexities of building and maintaining the infrastructure typically associated with these. Another category is SaaS, where software is centrally hosted and allocated on a subscription basis. SaaS is also known as on-demand software, web-based software, or web-hosted software. The four technology subfields involved in cloud services are deployment, integration, on-demand, and virtual private networks.
[0053] In some embodiments, systems according to various approaches may include a processor and logic integrated with and / or executable by the processor, the logic configured to perform one or more of the process steps enumerated herein. The processor may be of any configuration as described herein, such as a discrete processor or processing circuitry including many components, such as processing hardware, memory, I / O interfaces, etc. "Integrated with" means that the logic is embedded in the processor as hardware logic, such as an application-specific integrated circuit (ASIC), FPGA, etc. "Executable by the processor" means that the logic is hardware logic; software logic, such as firmware, part of an operating system, part of an application program, etc.; or any combination of hardware and software logic that is accessible by the processor and configured to cause the processor to perform some function when executed by the processor. The software logic may be stored in any memory type, local and / or remote, as known in the art. Any processor known in the art may be used, such as a software processor module and / or a hardware processor such as an ASIC, FPGA, central processing unit (CPU), integrated circuit (IC), graphics processing unit (GPU), etc.
[0054] Of course, this logic may be implemented according to various approaches as a method on any device and / or system or as a computer program product.
[0055] As noted above, the increase in data generation has led to an increase in the overhead associated with data management and processing. While AI has been developed in an attempt to mitigate this increase in processing overhead, advances in AI have also led to an increase in the complexity of machine learning models. Increasingly complex machine learning models lead to more intensive workloads and an increase in the load associated with applying the models to received data. The operation of traditional implementations has been negatively impacted thereby.
[0056] Cloud computing has been implemented in an attempt to improve the ability to perform computationally intensive operations and process increasing amounts of data. For example, cloud locations can be adjusted to provide dynamic levels of computational throughput that adjust to meet client needs. While this is effective in preventing processing bottlenecks from occurring, it involves transmitting all data to be analyzed to a centralized location, such as a data center or public cloud location. Sending data to a centralized location exposes it to unwanted attacks and unintentional mishandling, thereby significantly increasing the risk of data loss.
[0057] In an attempt to curb this reliance on networks to perform all processing at a central location, edge computing is implemented to extend computing to endpoints within a system. For example, applications and other types of computational operations are moved to edge locations where data is generated, for data privacy and security benefits. For example, to enhance data security and privacy, data may not be allowed to leave the borders of a particular country. In another example, a business may prefer to store generated data at an edge location (e.g., "on prem") so that it is not shared over the network.
[0058] While these types of data management schemes may improve data integrity, they significantly increase the computational overhead associated with doing so. For example, edge locations often experience different configurations, local constraints, demands, etc., which change rapidly over time. Conventional products are therefore forced to provide edge locations with sufficient supplies to operate across a wide range of conditions. By way of example, which is by no means intended to be limiting, conventional products are forced to overprovision RAG data to edge locations in an attempt to avoid training and deploying inaccurate large language models (LLMs) at edge locations. However, this causes undesirable and significant increases in network traffic and edge node latency. This problem results in conventional implementations experiencing heavy data storage and slow data retrieval times.
[0059] In sharp contrast to the above-mentioned shortcomings experienced by conventional systems, the techniques herein desirably can dynamically adapt the amount and / or type of supplies provided to an edge node, based at least in part on real-time performance experienced at the edge node itself and / or other locations in the distributed system, resulting in improved performance at each edge node and more efficient operation of the system as a whole, e.g., as described in further detail below.
[0060] Referring now to FIG. 2A, a system 200 having a distributed architecture according to one approach is shown. As an option, the system 200 may be implemented in conjunction with features from any other approach listed herein, such as those described with reference to other figures, such as FIG. 1. However, such system 200 and other systems presented herein may be used in a variety of applications and / or permutations that may or may not be specifically described in the example approaches or implementations listed herein. Furthermore, the system 200 presented herein may be used in any desired environment. Thus, FIG. 2A (and other figures) may be considered to include any possible permutations.
[0061] As shown, system 200 includes a central server 202 connected to user devices 204 and edge nodes 206, accessible to users 205 and administrators 207, respectively. Each of the central server 202, user devices 204, and edge nodes 206 is connected to a network 210, which may be located in different geographic locations. Network 210 may be of any type, depending, for example, on the desired approach. For example, in some approaches, network 210 is a WAN, such as the Internet. However, an exemplary list of other network types that network 210 may implement may include, but is not limited to, a LAN, a PSTN, a SAN, an internal telephone network, etc. As a result, any desired information, data, commands, instructions, responses, requests, etc. may be transmitted between user devices 204, edge nodes 206, and / or central server 202, regardless of the amount of separation that exists therebetween, e.g., despite being located in different geographic locations. According to some approaches, the central server 202 is a remote cloud server that is connected to (eg, can be accessed by) the user devices 204 and / or edge nodes 206.
[0062] However, it should be noted that two or more of the user devices 204, edge nodes 206, and central server 202 may be connected in different ways depending on the approach. According to one example, which is not intended to limit the invention in any way, two servers (e.g., nodes) may be located relatively close to each other and connected by, for example, a hardwired connection such as a cable, fiber optic link, wire, or any other type of connection that will be apparent to one of ordinary skill in the art after reading this specification.
[0063] The terms "user" and "administrator" are not intended to be limiting in any way. For example, while users and administrators may be described as individuals in various implementations herein, users and / or administrators may be applications, organizations, pre-configured processes, etc. The use of "data," "dataset," and "information" herein are not intended to be limiting in any way and may include any desired type of detail depending, for example, on the type of operating systems implemented on user device 204, edge node 206, and / or central server 202.
[0064] In some approaches, a portion of the dataset of text inputs (e.g., alphanumeric strings) generated, received, stored, identified, etc. at the central server 202 may be transmitted to the edge node 206. According to one example, the central server 202 may manage a knowledge database that maps multiple RAG data embeddings to corresponding edge node conditions. In other words, the knowledge database may correlate specific RAG data embeddings with the edge node conditions under which the specific RAG data was applied (e.g., received and actually utilized) at the respective edge node. Thus, entries in the knowledge database that match the conditions experienced at the edge node may be used to train one or more AI-based models, thereby identifying specific RAG data relevant to a given situation, e.g., as described in further detail below.
[0065] Continuing with reference to FIG. 2A , central server 202 includes a large (e.g., robust) processor 212 coupled to cache 211, AI module 213, and data storage array 214 having a relatively high storage capacity. AI module 213 may include any desired number and / or type of AI-based models, such as machine learning models, deep learning models, neural networks, etc. In a preferred approach, AI module 213 and / or processor 212 may train one or more AI-based models. For example, an AI model may be trained in some approaches by applying a predetermined training dataset to learn how to evaluate information associated with an edge node. For example, an AI model may be trained to evaluate performance metrics, edge node condition information, information outlining specific resources (e.g., RAG data) applied at the edge node, knowledge databases, etc., and output resources related to the evaluated information. An AI model may be trained in other approaches by applying a predetermined training dataset to learn how to identify resources related to the evaluated information. For example, an AI model may be trained to identify and output RAG data relevant to evaluated performance metrics, edge node condition information, etc. An AI model may be trained in other ways by applying a predetermined training dataset to learn how to predict future edge node conditions and / or performance metrics. For example, an AI model may be trained to evaluate future edge node conditions and a knowledge database and output a subset of RAG data along with expected relevance.
[0066] In some approaches, this can be achieved by implementing prompt tuning, more specifically, MPT. For purposes of this specification, "prompt tuning" refers to the process of adapting a base pre-trained model to each desired task through conditioning on learned prompt vectors. For example, prompt tuning can be used to efficiently adapt an LLM to multiple downstream tasks. Note also that "MPT" refers to a process that initially involves learning a single transferable prompt by extracting knowledge from multiple task-specific source prompts. Furthermore, multiplicative low-rank updates to this shared prompt are learned to efficiently adapt it to each downstream target task, as will be understood by those skilled in the art, for example, after reading this specification. As a result, the techniques herein can utilize prompt vectors in a multi-task learning setting and leverage rich cross-task knowledge.
[0067] According to some approaches, AI module 213 and / or data storage array 214 include a vector storage pool containing multiple data sets, each of which has been applied to multiple encoding models. Each encoding model may correspond to a different LLM supported by the system. In other words, each encoding model may apply a different language space that interprets a given data set in a way that is unique to the respective LLM. The LLMs supported by the system may include, but are in no way limited to, the T5 Transformer model, the Bidirectional Encoder Representations from Transformers (BERT) language model, the ELECTRA language model, etc., or any other LLM (e.g., language space) that will be apparent to one of ordinary skill in the art after reading this specification.
[0068] Each entry in the vector database may be compared against vector information received from other locations. For example, a mean vector received from an edge node 206 may be compared against entries in the vector database to identify the "N" entries that are closest matches to the received mean vector. In some approaches, the entries in the vector database may be organized so that the distance between entries is inversely proportional to how similar the entries are. The received mean vectors may then be plotted in the vector database, and the "N" closest entries may be selected as the datasets that are closest matches to the dataset that generated the mean vector, for example, as will be understood by those skilled in the art after reading this specification.
[0069] 2A , the user device 204 includes a processor 216 coupled to a memory 218. The processor 216 receives input from and interfaces with the user 205. For example, the user 205 may enter information using one or more of the display screen 224, keys on a computer keyboard 226, a computer mouse 228, a microphone 230, and a camera 232. The processor 216 may thereby be configured to receive input (e.g., text, sound, image, motion data, etc.) from any of these components as entered by the user 205. These inputs typically correspond to information presented on the display screen 224 while the entry is being received. Furthermore, the inputs received from the keyboard 226 and the computer mouse 228 may affect the information shown on the display screen 224, data stored in the memory 218, information gathered from the microphone 230 and / or the camera 232, the status of an operating system implemented by the processor 216, etc. The electronic device 204 also includes a speaker 234 that can be used to play (eg, project) audio signals for the user 205 to hear.
[0070] Some data (e.g., non-sensitive data) may be received from user 205 for storage and / or evaluation using AI module 213 at central server 202. The data may be received as a result of user 205 using one or more applications, software programs, temporary communication connections, etc. running on user device 204. For example, user 205 may upload data for storage in data storage array 214 and evaluation using processor 212 and / or AI module 213 of central server 202. As a result, the data is evaluated and processed.
[0071] Referring now to edge node 206, some of the components contained therein may be the same as or similar to those contained in user device 204, some of which are correspondingly numbered. For example, controller 217 is coupled to memory 218, display screen 224, keys on computer keyboard 226, and computer mouse 228. Additionally, controller 217 is coupled to AI module 238.
[0072] As described above with respect to AI module 213, AI module 238 may include any desired number and / or type of AI-based models. Thus, AI module 238 may implement similar, the same, or different characteristics as AI module 213 in central server 202. In some approaches, AI module 238, controller 217, and / or edge node 206 as a whole may be configured to operate in ultra-low latency conditions (e.g., less than about 1 millisecond). Thus, moving more computationally intensive applications to edge locations involves consolidation.
[0073] For example, the process of transmitting RAG data to edge locations for implementation in an "on-the-edge" LLM preferably involves tailoring the content being sent along the RAG pipeline. In other words, only relevant content for the use case is hosted at the edge location, thereby ensuring that the search mechanism does not impair ultra-low latency service. The LLM is also tuned to perform as desired in given conditions. This is also referred to herein as an edge node supporting "slim RAG" to reduce overhead and latency. This also desirably results in faster searches and faster overall response of the LLM, which is particularly important for operations performed on the edge. As noted above, the dynamic nature of the network edge causes the resources captured in the on-the-edge RAG pipeline to be of shifting importance. For example, different edge nodes have different characteristics across the network. Edge nodes may have equipment from different vendors, different types of users, different locations, etc. Conditions at edge nodes may also change throughout the day. For example, different traffic patterns may be observed in the morning compared to the evening, different types of users may be connected on different days, etc. According to one example, VIP users may connect to applications in the afternoon compared to a manufacturing facility that operates during the night. In another example, more coverage issues may be experienced during morning operations, while more quality issues may be experienced for VIP subscribers during afternoon operations. This difference may be used to determine that more documentation of operations for VIP subscribers in the afternoon is desirable. Again, the techniques herein achieve efficient LLM operation at edge locations by implementing a slim RAG pipeline that facilitates fast data retrieval, which has not previously been possible.
[0074] Referring now momentarily to FIG. 2B, a representative diagram 250 is illustrated for dynamically adapting the amount and / or type of RAG data sent to an edge node based at least in part on real-time performance at the edge node itself, according to one approach that is in no way intended to be limiting. As an option, this diagram 250 may be implemented in conjunction with features from any other approach listed herein, such as those described with reference to other figures, such as FIG. 2A. However, such diagram 250 and other diagrams presented herein may be used in various applications and / or permutations that may or may not be specifically described in the exemplary approaches listed herein. Furthermore, diagram 250 presented herein may be used in any desired environment. Thus, FIG. 2B (and other diagrams) may be considered to include any possible permutations.
[0075] The process of learning the correct amount of RAG data to send to the edge node evolves over time as demand changes. In response, the edge node 252 collects performance metrics (see operation 251) and sends them to the centralized edge orchestrator 254. The edge node 252 also collects information outlining the RAG data actually used at the edge node 252. In response, operation 253 includes recording (e.g., storing) details of the RAG data retrieval, along with details outlining how (or whether) the retrieved RAG data was actually used. As shown, the information collected in operation 253 is also sent to the centralized edge orchestrator 254.
[0076] The centralized edge orchestrator 254 then evaluates the information received from the edge node 252. For example, operation 255 includes evaluating the RAG data search quality and the knowledge database. The centralized edge orchestrator 254 also updates the knowledge database using the information received from the edge node 252. Note that the knowledge database and any associated logical components in the callout 260 may actually be located within one or more processors present in the centralized edge orchestrator 254, for example, as will be understood by those skilled in the art after reading this specification.
[0077] The knowledge database is formed by dividing base data 262 into multiple text segments (e.g., chunks). The text segments are further converted into segmentation embeddings. Each embedding thus corresponds to a respective smaller chunk of the original base data. Each of the embeddings is then combined with a respective piece of information to form a vector storage pool 264. As shown, vector storage pool 264 includes multiple entries. In a preferred approach, the number of entries "N" is the same as or similar to the number of embeddings formed by converting the base data into multiple smaller chunks.
[0078] Returning to the centralized edge orchestrator 254, the evaluation performed in operation 255 is passed to one or more AI models trained to dynamically adapt the amount and / or type of RAG data provided to the edge nodes. Accordingly, this on-the-fly adaptation is based at least in part on the real-time performance experienced by the edge nodes themselves and / or other locations in the distributed system. This results in improved performance at each edge node and a more efficient operation of the system as a whole.
[0079] Accordingly, operation 257 includes dynamically determining the amount and / or type of RAG data to send to edge node 252 based at least in part on real-time performance at edge node 252 itself and information accumulated and organized in a knowledge database. Note that operation 257 preferably references (e.g., incorporates) embedded deployment database 258. In some approaches, embedded deployment database 258 may incorporate (e.g., account for) all relevant artifacts for each connected edge node. Embedded deployment database 258 may also incorporate governance schemes that orchestrate the operation of the overall system. These governance schemes may be predetermined by a user, configured based on industry standards, related to the operating language of the running application, output by one or more AI-based models in response to evaluating input information, etc.
[0080] A stream of adjusted RAG data is thereby generated using the results of operation 257. The associated RAG data is then transmitted to edge node 252 for implementation. Thus, the operations illustrated in FIG. 2B may be repeated over time in an iterative manner to maintain a stream of associated resources (e.g., RAG data) transmitted to edge node 252. Again, this allows for significant reductions in network traffic and processing overhead. Consequently, this desirably improves the applicability of edge nodes and data for distributed applications (e.g., processing).
[0081] 3A, a method 300 for dynamically adapting the amount and / or type of supplies provided to an edge node. Specifically, method 300 involves dynamically determining the amount and / or type of RAG data sent to an edge node over time based at least in part on real-time performance at the edge node itself. This results in improved performance at each edge node and more efficient operation of the system as a whole, e.g., as described in further detail below.
[0082] Method 300 may be implemented in a variety of ways in accordance with the present invention, particularly in any of the environments shown in FIGS. 1-2B. Of course, method 300 may include more or fewer operations than those specifically illustrated in FIG. 3A, as will be understood by those skilled in the art upon reading this specification. Each of the steps of method 300 may be performed by any suitable component of an operating environment. For example, nodes 301 and 302 shown in the flowchart of method 300 may correspond to one or more processors located at different locations in a distributed system. Furthermore, each of the one or more processors is preferably configured to communicate with one another.
[0083] In various approaches, method 300 may be performed partially or wholly by a controller, processor, etc., or some other device having one or more processors therein. A processor, e.g., a processing circuit, chip, and / or module, implemented in hardware and / or software and preferably having at least one hardware component, may be utilized in any device to perform one or more steps of method 300. Exemplary processors include, but are not limited to, a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc., combinations thereof, or any other suitable computing device known in the art.
[0084] As mentioned above, FIG. 3A includes different nodes 301 and 302, both of which represent one or more processors, controllers, computers, etc. located at different locations within a distributed system. For example, in some approaches, one or more of the operations in method 300 may involve one or more physical components within an edge node of the distributed system. The edge node may also be one of several coupled to a central server as part of a larger distributed system. Accordingly, node 301 may include one or more processors located at the central server of the distributed system (e.g., see processor 212 in FIG. 2A ). Furthermore, node 302 may include one or more processors located at a first edge node (e.g., see controller 217 in edge node 206 in FIG. 2A ).
[0085] Accordingly, commands, code, data, metadata outlining code updates, etc. may be transmitted between node 301 and node 302 depending on the approach. Note also that the various processes included in method 300 are not intended to be limiting in any way, as would be understood by one of ordinary skill in the art after reading this specification. For example, data transmitted from node 302 to node 301 may, in some approaches, originate from a request transmitted from node 301 to node 302.
[0086] As shown in the flowchart, method 300 includes operation 304 performed at node 301. There, operation 304 includes installing a knowledge database. In other words, operation 304 includes initializing (e.g., establishing) a logical space that can be used to create and maintain the knowledge database. As used herein, a "knowledge database" can be used to store representations of various public data sets. For example, a knowledge database can be created by mapping embeddings of RAG data to various edge node conditions when the RAG data is used.
[0087] Referring momentarily to Figure 3B, exemplary sub-operations for installing a knowledge database in one manner are shown. Accordingly, one or more of these sub-operations may be used to perform operation 304 of Figure 3A. Note, however, that the sub-operations of Figure 3B are illustrated in one manner and are not intended to be limiting in any way.
[0088] As shown, sub-operation 332 includes dividing the set of base data into multiple text segments. In other words, sub-operation 332 includes dividing the RAG source data (e.g., files, documents, etc.) into multiple chunks or segments. In different approaches, the base data may be divided into a predetermined number of text segments, text segments of a predetermined size, text segments of a predetermined type, etc. In other approaches, the size, number, type, etc. of the text segments formed from the set of base data may depend on the size and / or type of the base data, instructions received from a user, predetermined conditions met, etc.
[0089] From suboperation 332, the flowchart proceeds to suboperation 334. There, suboperation 334 involves converting each of the text segments into a segmentation embedding, where in a preferred approach, each embedding corresponds to a respective smaller chunk of the original base data.
[0090] In other words, each embedding is generated by transforming objects (e.g., their representations) such as text, images, and audio. In some approaches, the embeddings are created using deep learning and are vectors of floating-point numbers that represent similarities between objects in a low-dimensional space. Accordingly, each of the embeddings is further combined with its respective information to form a vector storage pool. See suboperation 336. The vector storage pool may contain multiple entries that are the same as or similar to multiple embeddings formed by transforming the base data into multiple smaller chunks, or "embeddings" as used herein. In some approaches, a vector storage pool is formed for each embedding.
[0091] The vector storage pool is preferably combined with a knowledge database that can be referenced to determine the specific RAG data to send to the edge node. For example, one or more AI-based models can be trained to evaluate entries in the knowledge database and determine the specific portions of RAG data (e.g., embeddings) that should be sent to the edge node. Again, sending specific portions of RAG data reduces network load and processing overhead, thereby improving overall performance.
[0092] 3A , the method proceeds from operation 304 to operation 306. There, operation 306 includes receiving an initial request from node 302. The initial request may specify particular RAG data embedding to be sent along the RAG pipeline extending between node 301 and the edge node at node 302. In other approaches, the initial request may not specify RAG data. In some approaches, the initial request is fulfilled based on previous (e.g., stored) performance.
[0093] In response to receiving the initial request, node 301 collects relevant RAG data (see operation 308) and transmits it to node 302 (see operation 310). Thus, the amount of RAG data transmitted to node 302 may vary. In response to receiving the RAG data, node 302 evaluates the RAG data and applies at least a portion of it. See operation 312. In other words, node 302 examines the RAG data received from node 301 and determines whether any of it is relevant (e.g., can be used in the current conditions and taking into account the current performance of the edge node). In some approaches, operation 312 includes applying selected portions of the RAG data determined to be actually relevant to the application running on the edge node.
[0094] The method proceeds from operation 312 to operation 314, where the particular RAG data actually applied at the edge node, as well as information outlining the current conditions at the edge node, is sent back to node 301 in real time. In other words, a central location at node 301 receives information describing which RAG data was used at the edge node, along with the conditions that existed at the edge node while the RAG data was being applied (e.g., received and actually used at the edge node). In addition to the information outlining the RAG data actually used at the edge node, node 302 may send additional information to node 301 that is useful in determining how the edge node is performing. For example, in some approaches, the edge node at node 302 collects and sends performance metrics corresponding to how the edge node performed during a given period of time. In other approaches, the edge node at node 302 collects information describing (e.g., outlining) the conditions at the edge node while the RAG data was being used.
[0095] Proceeding to act 316, the knowledge database is dynamically updated using any information received from node 302. In other words, the knowledge database installed in act 304 is updated over time to include (e.g., incorporate) information outlining how RAG data is used at edge nodes under various conditions and performance characteristics. In some approaches, received performance metrics, received edge node condition information, received information outlining RAG data usage at node 302 (e.g., and other nodes), knowledge database entries, etc. are used to retrain one or more AI-based models implemented in the centralized edge orchestrator. This enables the AI-based models to maintain an accurate understanding of edge nodes and how they utilize RAG data (and other resources) under different conditions.
[0096] In response, operation 318 includes causing one or more trained AI-based models to dynamically evaluate performance metrics, edge node condition information, received information, a knowledge database, etc., or any other desired information that may have been received from node 302. Further, operation 320 includes causing the one or more AI-based models to output a subset of relevant RAG data based at least in part on the dynamic evaluation. In other words, operations 318 and 320 include determining, for at least one edge node, which resources (e.g., RAG data) are relevant based on historical information (e.g., a knowledge database) and current conditions and / or settings at the edge node itself, and transmitting the relevant resources (e.g., RAG data) to the edge node for implementation (e.g., training an LLM for a particular application). For example, as will be understood by one of ordinary skill in the art after reading this specification, this may be accomplished in several manners, such as by transmitting one or more instructions, commands, requests, files, etc. In response, operation 322 includes transmitting the relevant portion of the RAG data from node 301 back to the edge node at node 302. In response to receiving the RAG data, node 302 may utilize the RAG data as desired, for example, as will be understood by those skilled in the art after reading this specification.
[0097] 3A illustrates a central node 301 dynamically controlling the flow of RAG data to one edge node 302 based at least in part on real-time performance at the node 302, it should be noted that this is not intended to be limiting in any way. Rather, the central node 301 may monitor and control the flow of RAG data to any desired number of edge nodes, individually and in parallel. Furthermore, RAG data may flow from the central node 301 to each of the edge nodes along respective pipelines. Thus, the techniques herein desirably enable the implementation of LLMs and other data- and / or training-intensive models while minimizing impact on the overall system.
[0098] The techniques herein can thereby desirably enable the continuous evolution of generative AI and foster new opportunities, particularly in the telecommunications industry. Additionally, networking clients will benefit from the techniques herein by being able to extend their operations with generative AI use cases, as will be understood by those skilled in the art, for example, after reading this specification. The techniques may therefore achieve faster operation to further support services running at edge locations.
[0099] The techniques herein focus on filtering RAG searches to obtain data that is most relevant to a given situation (e.g., a prompt). As noted above, this is achieved herein by implementing dynamic adaptation of RAG content deployed at the edge, particularly as edge conditions change over time. The techniques herein can thus desirably achieve minimal RAG deployment for limited edge resources at each observed time point.
[0100] While current (e.g., real-time) factors may be considered while determining relevant resources (e.g., RAG data) to send to an edge node in real time, other considerations may be made. For example, predictions about future edge node conditions may be made. Referring now momentarily to FIG. 3C , an example flowchart 350 for predicting future edge node conditions and / or performance metrics according to one approach is illustrated. Accordingly, one or more of these sub-operations may be used to supplement the operations of FIG. 3A , such as operation 320. However, it should be noted that the sub-operations of FIG. 3C are illustrated according to one approach, which is in no way intended to be limiting.
[0101] As shown, operation 352 includes having one or more generative AI models evaluate historical performance data and predict future edge node conditions. The generative AI models are preferably trained to make predictions at the edge node based at least in part on historical data (e.g., stored in a knowledge database). For example, as will be understood by those skilled in the art after reading this specification, depending on the approach, operation 352 may be accomplished in several ways by sending one or more instructions, commands, requests, files, etc. In some ways, the generative AI-based models may be trained to predict future edge node performance metrics.
[0102] Further, operation 354 includes having one or more trained AI-based models dynamically evaluate future edge node conditions and the knowledge database. In other words, the AI-based models are trained to evaluate conditions (e.g., current or future) and / or performance metrics (e.g., current or future) to determine relevant portions of available RAG data to send to the edge node. Thus, the trained AI-based models output a subset of the available RAG data. See operation 356. In some approaches, the relevant subset of RAG data may be output along with expected relevance. In other words, the trained AI-based models determine what information is expected to be relevant to the edge node based on predictions output by the generative AI-based model and / or the trained AI-based models.
[0103] In response, a subset of the available RAG data predicted to be relevant to the future conditions and / or performance metrics is actually output by the trained AI-based model. Furthermore, operation 358 includes replacing at least a portion of the RAG data currently being transmitted to the edge node (e.g., along the RAG pipeline) with at least a portion of the RAG data predicted to be relevant. However, in some approaches, operation 358 may include adding to (e.g., supplementing) the RAG data currently being transmitted to the edge node. In other approaches, operation 358 may include removing some of the RAG data currently being transmitted to the edge node. Thus, the RAG data corresponding to the predicted edge node conditions and / or performance metrics may be combined with the RAG data currently being transmitted to the edge node in a number of different ways, depending, for example, on the desired approach. In a further approach, operation 358 is performed in response to waiting a predetermined time. For example, operation 358 may be implemented in response to reaching a date, time, predetermined conditions, etc., corresponding to the predicted edge node conditions and / or performance metrics.
[0104] 3C , flowchart 350 proceeds from operation 358 to operation 360. There, operation 360 includes transmitting (e.g., forwarding) the combined RAG data (e.g., resources) from the centralized edge orchestrator to each edge node. Again, this allows only resources predicted to be relevant to be sent to the edge node. This allows for a reduction in network traffic while also reducing the computational overhead and latency experienced at each edge node.
[0105] Referring now to Figure 4, a distributed system 400 is shown, according to an example use case that is not intended to be limiting in any way. The system 400 may be implemented in conjunction with features from other approaches enumerated herein, such as those described with reference to other figures. Furthermore, the system 400 presented herein may be used in any desired environment. Thus, Figure 4 (and other figures) may be considered to include any possible permutations.
[0106] As shown, in some approaches, the base data 402 is watsonx.data. The base data 402 is received and divided into “N” text segmentations. The N text segmentations are further converted into N segmentation embeddings. N vector storage pools are formed (e.g., created) using the N segmentation embeddings. Each of the vector storage pools includes a unique vector storage identifier (e.g., VS1), one or more segments of the original base data 402 (e.g., S1 segments), at least one embedding (e.g., embedding) corresponding to the one or more segments, and the type of edge environment in which the base data was utilized (e.g., type E1). These vector storage pools thus provide an efficient and segmented process for locating relevant sections of a given base data 402 for a particular situation faced by an edge node.
[0107] The N vector storage pools are further used to form a knowledge database. In a preferred approach, the knowledge database is formed by identifying a mapping between edge conditions and / or edge performance and vector storage and / or segment relevance, e.g., as described herein. Furthermore, one or more AI-based models may be trained with the information contained in the knowledge database and configured to identify relevant portions of the information contained in the knowledge database. The relevant portions of this information may then be transmitted to the edge node for implementation. According to one example, the knowledge database stores RAG data and is configured to generate relevant portions of available RAG data to transmit to the edge node based at least in part on the performance experienced at the edge node, the current state of the edge node, and the expected workload at the edge node (e.g., as output by one or more generative AI models trained to predict future edge node conditions).
[0108] Accordingly, in response to the edge node being initialized to "start," the knowledge database is consulted and embeddings are deployed to the edge node. For example, operation 404 includes deploying the embeddings on a RAG pipeline extending between the centralized edge orchestrator and the edge node. In some approaches, the initial embeddings deployed on the edge node may be selected randomly based at least in part on the last operation executed on the edge node, based at least in part on current conditions at the edge node (e.g., determined during initialization of the edge node), etc.
[0109] Referring to operation 406, the edge node tracks operational metrics, the quality of service (QoS) experienced by users at the edge node, and any other performance-related details. This may be accomplished by collecting sensor measurements, storing outputs generated by trained AI-based models, recording the performance of the edge node itself, etc. This tracked information is then preferably returned to the centralized edge orchestrator. In some approaches, the tracked performance information may be sent to the centralized edge orchestrator periodically (e.g., at fixed or random intervals), in response to a predetermined amount of performance information being collected, in response to receiving a request from the centralized edge orchestrator, in response to predicting upcoming workload and / or edge conditions, etc.
[0110] Proceeding to operation 408, the edge node also tracks edge conditions as well as operations being performed at the edge node itself. In other words, operation 408 includes storing information outlining conditions (e.g., operational states, error condition indicators, workload level warnings, thresholds at which data overflows memory, etc.). Thus, while operation 406 may track particular performance metrics achieved at the edge node, operation 408 may track what the edge node itself is experiencing while achieving the performance metrics. The tracked edge conditions are also preferably returned to the centralized edge orchestrator. In some approaches, the tracked edge conditions may be sent to the centralized edge orchestrator periodically (e.g., at fixed or random intervals), in response to a predetermined amount of performance information being collected, in response to receiving a request from the centralized edge orchestrator, in response to predicting upcoming workload and / or edge conditions, etc.
[0111] Returning to operation 404, the flowchart also proceeds to operations 410, 412, and 414, in parallel with operations 406 and 408. Accordingly, various operations may be performed simultaneously and / or in parallel, depending, for example, on the desired application. As shown, operation 410 includes receiving relevant information from a centralized edge orchestrator along a RAG pipeline extending therebetween. In other words, RAG data is transmitted from the centralized edge orchestrator to edge nodes along the RAG pipeline, for example, as described herein.
[0112] From operation 410, the flowchart proceeds to a RAG search record procedure, which includes operations 412 and 414, as shown. Operation 412 involves an application running on the edge node registering the received RAG search information (e.g., details). In several ways, the RAG search information may be registered by processing information in a RAG search record, for example, as will be understood by those skilled in the art after reading this specification.
[0113] Further, operation 414 includes the edge node sharing the RAG search record with the centralized edge orchestrator. In other words, the edge node notifies the centralized edge orchestrator of which received RAG data was actually applied (e.g., utilized) at the edge node. In other words, while certain RAG data may be sent to the edge node, the edge node may not use some of the received RAG data. By tracking the RAG data used at the edge node, the centralized edge orchestrator may use this tracked information to train AI-based models to recognize patterns and relationships among various information. At least some of these AI-based models may therefore be trained and configured to evaluate information associated with the edge node, learn how to identify resources related to the evaluated information, and / or predict future edge node conditions and / or performance metrics, for example, as described in the techniques herein. For example, the AI models may be trained to evaluate performance metrics, edge node condition information, information outlining specific resources (e.g., RAG data) applied at the edge node, knowledge databases, etc., and output resources related to the evaluated information. In another example, an AI model may be trained to identify and output RAG data relevant to evaluated performance metrics, edge node condition information, etc. In yet another example, an AI model may be trained to evaluate future edge node conditions and knowledge databases and output a subset of RAG data along with expected relevance.
[0114] The flowchart proceeds from operation 414 to operation 416. Operation 416 is executed at the centralized edge orchestrator and includes correlating information received from the edge nodes. In other words, operation 416 includes correlating performance reports, edge condition reports, local RAG search record information, etc. received from the edge nodes in an attempt to identify actual RAG data utilized at the edge nodes. This enables the centralized edge orchestrator to identify RAG data that was sent to the edge nodes but not actually used. This identified unused RAG data can therefore be used for retraining AI-based models, updating the knowledge database (e.g., see the arrow line extending from operation 416 to the knowledge database), etc. Operation 416 can thereby evaluate RAG search quality at the edge nodes as well as management of the knowledge database.
[0115] Referring now to operation 420, the centralized edge orchestrator uses a generative AI model to evaluate historical data (e.g., information stored in a knowledge database) associated with edge node conditions and make predictions regarding future edge node conditions. In other words, operation 420 involves the generative AI model generating predictions about how the edge node (or other edge nodes) will behave and / or what the edge node will experience. As shown, the output of the generative AI model is passed to operation 422. There, operation 422 includes the centralized edge orchestrator reading the knowledge database to identify a desired set of embeddings (e.g., a portion of the RAG data) to be deployed on the edge node for the predicted set of future conditions. In other words, evaluating the mappings in the knowledge database identifies sections of data predicted to be related to conditions the edge node is expected to experience. Furthermore, operation 424 includes the centralized edge orchestrator updating a RAG pipeline extending between the centralized edge orchestrator and the edge node to incorporate the RAG data identified in operation 422. In some approaches, the identified RAG data is used to replace at least a portion of the RAG data currently received at the edge node. In other approaches, the identified RAG data supplements the RAG data already transmitted to the edge node. In still other approaches, some of the RAG data may simply be stopped from being transmitted to the edge node, for example, in response to the identified RAG data not including some of the currently transmitted RAG data. In response, operation 424 is shown returning to operation 404 and updating the knowledge database (e.g., see the arrowed line).
[0116] In some approaches, one or more of the operations in method 300, flowchart 350, and / or system 400 of FIG. 4 may be performed using an AI model trained using a predetermined set of training data. For example, in some approaches, the various operations described above may be performed in a trained state of a trained AI model. In some approaches, training of the AI model may be performed by applying a predetermined training dataset to learn how to evaluate information associated with the edge node. For example, the AI model may be trained to evaluate performance metrics, edge node condition information, information outlining particular resources applied at the edge node (e.g., RAG data), knowledge databases, etc., and output resources related to the evaluated information. In some approaches, training of the AI model may be performed by applying a predetermined training dataset to learn how to identify resources related to the evaluated information. For example, the AI model may be trained to identify and output RAG data related to the evaluated performance metrics, edge node condition information, etc. Training of the AI model may be performed in yet other approaches by applying a predetermined training dataset to learn how to predict future edge node conditions and / or performance metrics. For example, an AI model can be trained to evaluate future edge node conditions and knowledge databases and output a subset of RAG data with expected relevance. Initial training can include reward feedback, which in some approaches can be implemented using subject matter experts (SMEs) who generally understand the relevance of RAG data in training LLMs. However, to limit costs associated with relying on manual action from SMEs, in other approaches, reward feedback can be implemented using techniques for training BERT models, as will be apparent to those skilled in the art after reading this disclosure.During this training, if a determination is made that the AI model has achieved a redeemed threshold for accuracy in performing the operations described herein, a determination can be made that the model is trained and ready to be deployed to perform the techniques and / or operations of method 300, flowchart 350, and / or the operations performed in system 400 of FIG. 4. Because a neuromyotonic AI model may not require an SME and / or iteratively applied training with reward feedback to accurately perform the operations described herein, in some further approaches, the AI model can be a neuromyotonic AI model that can improve the performance of computing devices in an infrastructure associated with using RAG data to train LLMs to function as desired. Instead, the neuromyotonic AI model is itself configured to make the determinations described in the operations herein. In some approaches, the weight values can be used by an AI inference model to collect and analyze information and / or feedback potentially received from edge nodes and / or LLMs that may be implemented therein. Such AI models ensure that relevant resources (e.g., RAG data) are sent to edge nodes regardless of the conditions and / or performance experienced, and the scale of such analysis and determination is not otherwise feasible for a human to perform. This is because a human cannot efficiently evaluate the myriad factors in play, and the process of attempting to do so would otherwise incorporate processing delays and errors in identifying resources associated with given conditions and / or performance metrics. Therefore, the management of operations described herein cannot be achieved by manual human action.
[0117] It will be apparent from the description provided above that the various features of the systems and / or methodologies described above may be combined in any manner, thereby creating multiple combinations.
[0118] It will be further appreciated that implementations of the present invention may be provided in the form of a service that is deployed on behalf of a customer to provide the service on demand.
[0119] The descriptions of various implementations of the present invention are presented for illustrative purposes and are not intended to be comprehensive or limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein has been selected to best explain the principles of the implementations, practical applications, or technical improvements over technologies found in the market, or to enable others skilled in the art to understand the implementations disclosed herein.
Claims
1. 1. A computer implemented method (CIM) comprising: receiving, in real time, from an edge node, specific search expansion generation (RAG) data to be applied at the edge node and information outlining the condition of the edge node; updating a knowledge database with the received information, the knowledge database mapping RAG data embeddings to various edge node conditions; One or more trained artificial intelligence (AI)-based models: dynamically evaluating the received information and the knowledge database; and outputting a subset of the relevant RAG data; and transmitting a subset of the relevant RAG data to the edge node. CIM equipped with.
2. receiving, from the edge node, performance metrics corresponding to the edge node; receiving edge node condition information corresponding to the edge node from the edge node; and to the one or more trained AI-based models: dynamically evaluating the performance metrics, the edge node condition information, the received information, and the knowledge database; and outputting the subset of related RAG data. The CIM of claim 1 further comprising:
3. The knowledge database includes: dividing the base data into a plurality of segments; converting the segments to the embedding; and combining the embedding with the respective information to form a vector storage pool.
3. The CIM of claim 2 formed by:
4. The CIM of claim 3 , wherein one of the vector storage pools is formed for each of the embeddings.
5. The CIM of claim 1 , wherein operations in the CIM are performed by a centralized edge orchestrator.
6. The CIM of claim 5 , wherein the subset of relevant RAG data is transmitted to the edge node along a RAG pipeline extending between the edge node and the centralized edge orchestrator.
7. having one or more generative AI models predict future edge node conditions; and to the one or more trained AI-based models: dynamically evaluating the future edge node conditions and the knowledge database; and outputting a subset of the RAG data along with expected associations The CIM of claim 1 further comprising:
8. replacing at least a portion of the subset of associated RAG data with at least a portion of the subset of RAG data; and transmitting the remainder of the associated subset of RAG data and the subset of RAG data to the edge node. The CIM of claim 7 further comprising:
9. To the processor set: receiving, in real time, from an edge node, specific search expansion generation (RAG) data to be applied at said edge node and information outlining the condition of said edge node; updating a knowledge database with the received information, the knowledge database mapping RAG data embeddings to various edge node conditions; One or more trained artificial intelligence (AI)-based models: dynamically evaluating the received information and the knowledge database; and a procedure for outputting a subset of the relevant RAG data; and transmitting the subset of relevant RAG data to the edge node; A computer program for executing
10. In the processor set: receiving, from the edge node, performance metrics corresponding to the edge node; receiving edge node condition information corresponding to the edge node from the edge node; and to the one or more trained AI-based models: dynamically evaluating the performance metrics, the edge node condition information, the received information, and the knowledge database; and outputting a subset of the associated RAG data.
10. The computer program of claim 9, further comprising:
11. The knowledge database includes: Dividing the base data into segments; converting the segments into the embeddings; and combining said embeddings with respective information to form a vector storage pool; 11. The computer program of claim 10, formed by:
12. The computer program product of claim 11 , wherein one of the vector storage pools is formed for each of the embeddings.
13. The computer program product of claim 9 , wherein the operations performed by the processor set are performed by a centralized edge orchestrator.
14. The computer program product of claim 13 , wherein the subset of relevant RAG data is transmitted to the edge node along a RAG pipeline extending between the edge node and the centralized edge orchestrator.
15. In the processor set: training one or more generative AI models to predict future edge node conditions; and to the one or more trained AI-based models: dynamically evaluating the future edge node conditions and the knowledge database; and Procedure for outputting a subset of RAG data with expected associations 10. The computer program of claim 9, further comprising:
16. In the processor set: replacing at least a portion of the subset of associated RAG data with at least a portion of the subset of RAG data; and transmitting the remainder of the associated subset of RAG data and the subset of RAG data to the edge node.
16. The computer program of claim 15, further comprising:
17. A computer system (CS), Processor set; a set of one or more computer-readable storage media; and Program instructions collectively stored on the set of one or more storage media and configured to cause the set of processors to perform the following computer operations: receiving, in real time, from an edge node, specific search expansion generation (RAG) data to be applied at said edge node and information outlining the condition of said edge node; updating a knowledge database with the received information, the knowledge database mapping RAG data embeddings to various edge node conditions; One or more trained artificial intelligence (AI)-based models: dynamically evaluating the received information and the knowledge database; and outputting a subset of the relevant RAG data; and transmitting the subset of relevant RAG data to the edge node; and a CS comprising program instructions for causing the CS to execute the above steps.
18. The program instructions cause the processor set to perform the following computer operations: receiving, from the edge node, performance metrics corresponding to the edge node; receiving edge node condition information corresponding to the edge node from the edge node; and to the one or more trained AI-based models: dynamically evaluating the performance metrics, the edge node condition information, the received information, and the knowledge database; and outputting a subset of the relevant RAG data; The CS of claim 17, further comprising:
19. The knowledge database includes: Dividing the base data into segments; converting the segments into the embeddings; and combining said embeddings with respective information to form a vector storage pool; 19. The CS of claim 18 formed by
20. The program instructions cause the processor set to perform the following computer operations: Having one or more generative AI models predict future edge node conditions; and to the one or more trained AI-based models: dynamically evaluating the future edge node conditions and the knowledge database; and Output a subset of RAG data with expected associations The CS of claim 17, further comprising: