Virtual memory management in a neural routing apparatus
The neural routing apparatus addresses inefficiencies in content-based publish/subscribe systems by using a large language model to manage memory and resources dynamically, enhancing efficiency and scalability.
Patent Information
- Application Number
- PCT/FI2025/050115
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-10
- Filing Date
- 2025-03-10
- Publication Date
- 2025-09-18
AI Technical Summary
Existing content-based publish/subscribe systems in IoT devices face inefficiencies due to high memory requirements and imprecision in content routing, necessitating a need for expressive, multi-modal semantic content routing with efficient memory management.
A neural routing apparatus utilizing a large language model to match and batch communications, implement caching strategies, and manage resources dynamically to optimize memory usage and processing efficiency.
Improves information dissemination efficiency and accuracy, reduces memory requirements, and enhances scalability and resource awareness in content-based publish/subscribe systems.
Smart Images

Figure FI2025050115_18092025_PF_FP_ABST
Abstract
Description
[0001] VIRTUAL MEMORY MANAGEMENT IN A NEURAL ROUTING APPARATUS
[0002] DESCRIPTION OF BACKGROUND
[0003] The following disclosure relates to data communication technology . More particularly the disclosure relates to machine learning assisted contentbased publish / subscribe systems and memory management in those .
[0004] Internet-of-Things ( loT ) device is a battery- constrained node with limited computational and storing capabilities . Nevertheless , loT devices are used in a wide range of applications , such as , smart city scenarios , autonomous transportation, industrial control and automati zation systems , and other 4G / 5G / 6G and similar network applications . loT devices enable sensing and monitoring the environment and, due to their limited functions , further transmit data to a processing node using wireless technology, such as Bluetooth Low Energy (BLE ) , LoRaWAN, Wi-Fi , or cellular technologies 4G / 5G / 6G . Devices establish a connection between each other using an appropriate network protocol . loT data protocols comprise several categories . One example of a category is so called publish / subscribe category .
[0005] The publish / subscribe stands out for an asynchronous communication between different components or systems , improved scalability, low latency and simplicity compared to traditional request / response ( sender / receiver ) model . The components of publish / subscribe category includes a publisher, that provide information about the event ; a subscriber, that express the interest about the event ; and a message broker, that is responsible for acquiring, filtering and routing messages between the publisher and the subscriber . The publish / subscribe category supports several types of publishing strategies , such as a content-based strategy . Content-based publish / subscribe systems typically filter messages based on subscriber' s interest on the specific attributes in the publi shed data . They disseminate information to a large number of subscribers based on their interests using keyword-based filters . Usually, the subscriber expresses interest in a form of a special query, which needs to have the predetermined format and technical details for accurate match with the service from publisher .
[0006] This beforementioned way of subscribing to the particular content is inefficient and imprecise and often has high memory requirements that are difficult to fulfill . There is a need for expressive , i . e . natural , multi-modal semantic content routing from publishers of content to subscribers of content that has an efficient memory management . In addition, there is a need for scalable and resource aware system .
[0007] SUMMARY
[0008] In the following disclosure an arrangement for managing memory in a content-based publish / subscribe system is disclosed . The system receives communications from subscribers and publishers . The subscriptions and published services , such as events and advertisements , are matched together and packed in batches together so that the context memory of a large language model is used in an efficient manner . The data that is repeating itsel f in plurality of communications can be stored in a shared prompt memory and the si ze of the batches is reduced .
[0009] In the first aspect a router is disclosed . The router comprises of at least one processor configured to execute computer programs , and at least one memory configured to store data and a computer program code comprising instructions . The router is configured to receive a set compri sing at least one communication indicating the interest from a subscriber, receive a set comprising at least one communication indicating a service from a publisher, match at least one communication in the set received from the subscriber with a communication in the set received from the publisher . When matching, the router is configured to divide the received sets into batches and store the divided batches into a context memory of a large language model stored in the at least one memory . The router is further configured to determine the batch si ze based on the avai lable context memory so that at least one batch comprising a communication from the subscriber and a communication from the publisher fit into the context memory . Here and further, any communication indicating a service may comprise any events , information on the events , advertisements , or other relevant data, that was sent by publisher . Communications indicating the interest is a data from subscriber, where subscribers express their interest on event . Large language model is a type of artificial intelligence that process and understand the natural human language . A disclosed apparatus benefits in improved information dissemination efficiency and accuracy by utili zing large language models . It is advantageous in management and processing of large-scale communications in content-based publish / subscribe systems .
[0010] In an embodiment of the first aspect the router is conf igured to determine a caching strategy based on cache si ze limits , at least one eviction policy and a cache location . It results in optimal cashing configuration for effective system' s performance .
[0011] In an embodiment of the first aspect the router is configured to identify at least one resource to be stored into a cache . Resources that could benef it from cashing are , for example , frequently accessed subscription information, precomputed results , or lookup tables . In an embodiment the router is configured to store at least one data item in a communication received from the subscriber in a cache memory before processing of data items is initiated . It is beneficial to store the data item to the cache memory before processing, such as matching or batching, so that the data item is rapidly available when needed .
[0012] In an embodiment of the first aspect the router is configured to initiali ze a cache according to the determined caching strategy and identified at least one resource .
[0013] In an embodiment of the first aspect the router is set to configure a monitoring policy, a prediction policy or a resource allocation policy based on the current system configuration and current workload . These provide a dynamic memory management configuration .
[0014] In an embodiment of the first aspect the router is configured to group the sets received from the subscribers and the sets received from the publishers into batches .
[0015] In an embodiment of the first aspect the router is configured to match each of the batches comprising sets received from a subscriber with at least one batch received from a publisher and pack the matches with an instruction to a large language model invocation . The process results in a set of large language model invocations that are then scheduled .
[0016] In an embodiment of the first aspect the router is conf igured to parallel proces sing of a set of large language model invocations by creating a plurality of instances to process each large language model invocation in a batch concurrently and configuring synchroni zation and resource management for the instances for process safety .
[0017] Alternatively, in an embodiment of the first aspect , the router is configured to sequential processing of a set of large language model invocations by processing large language model invocation sequentially in a predefined order .
[0018] In an embodiment of the first aspect the router is conf igured to validate and update cached data based on changes that occurred during processing of the large language model invocations . This ensures cache validation and update .
[0019] In an embodiment of the first aspect the router is configured to release resources after processing of the large language model invocations is complete . This ensures cache release resources .
[0020] In an embodiment of the first aspect the router is configured to remove unused or outdated entries based on at least one eviction pol icy and cache s i ze limits . This ensures cache cleanup .
[0021] In an embodiment of the first aspect the router is configured to identify at least one data item in at least two communications . It is beneficial to identify similar data items within communications so that the memory usage can be optimi zed by pointing from two or more communications to one memory location containing the data item . In this manner the data item does not need to be stored twice .
[0022] In an embodiment of the first aspect the router is configured to store the identified at least one data item in a shared memory . It is beneficial to use shared memory to store data items that have been used in a plurality communications so that the content can be retrieved from one location and the content of that location can be updated if needed so that all instances have updated data .
[0023] In an embodiment of the first aspect the router is configured to determine a destination for the at least one communication received from a subscriber and route the at least one communication received from a publisher to the determined destination . When determining the destination, the router is configured to input the received communications to a large language model stored in the at least one memory .
[0024] In a second aspect a method for managing memory is disclosed . The method comprises of receiving a set comprising at least one communication from a subscriber, receiving a set comprising at least one communication from a publisher, and matching at least one communication in the set received from the subscriber with a communication in the set received from the publisher . When matching, the method further comprises of dividing the received sets into batches ; and storing the divided batches into a context memory of a large language model stored in the at least one memory . The method further comprises determining the batch si ze based on the avai lable context memory so that at least one batch comprising a communication from the subscriber and a communication from the publisher fit into the context memory . The disclosed method has advantages in scalable and flexible communication management , efficient memory utili zation, reduced system ' s latency, improved system ' s performance , isolation for communications from direct interaction with large language model instances , which results in ensuring consistent behavior, reliability, and fault tolerance in processing .
[0025] In a third aspect a computer program is disclosed . A computer program comprising instructions which, when executed by a computing device , are further configured to cause the computer device to perform a method according to the second aspect .
[0026] BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings , which are included to provide a further understanding of the virtual memory management in neural routing apparatus and constitute a part of this specification, illustrate embodiments and together with the description help to explain the principles of the virtual memory management in neural routing apparatus . In the drawings :
[0028] Fig . 1 illustrates an example of a publish / subscribe system;
[0029] Fig . 2 presents an example of a method for large language model memory usage ; and
[0030] Fig . 3 is a flow chart of an example method for managing memory .
[0031] DETAILED DESCRIPTION
[0032] Reference will now be made in detail to the embodiments , examples of which are illustrated in the accompanying drawings .
[0033] In figure 1 an example of a of a system 100 comprising a neural routing apparatus 102 is disclosed . The apparatus 102 comprises at least one processor 104 and at least one memory 106 . The at least one processor 104 and at least one memory 104 are configured to work together and execute computer program code 110 that is stored to the at least one memory . The at least one memory is also configured to store data relevant to the computer programs , operating systems and similar . The apparatus further comprises a communication interface 112 , which may comprise several network technologies . The communication interface may comprise one or more transceivers configured to transmit and receive data over wired or wireless network technologies .
[0034] In the example of Figure 1 one the neural routing apparatus is configured to receive communications from a subscriber 114 and a publi sher 116 . In the example of figure 1 only one of each is shown, however, in typical operation the neural routing apparatus accepts and receives communications from a plurality of subscribers and publishers .
[0035] The neural routing apparatus 102 comprises an acces s to at least one large language model . The large language model may be stored in the memory of the apparatus or it may be used as a distributed resource so that the at least one memory 106 stores a pointer or access information to at least one large language model resource . I f the large language model i s stored in the at least one memory 106 , then a portion of the at least one memory is allocated to the large language model to be used as the context memory . Similarly, if the router 102 i s accessing a large language model in a dif ferent computing system, then for the large language model has been allocated a memory space for operations .
[0036] Publishers aim to publish information on the events , advertisements on the event , or other relevant data, i . e . , any service , that would be delivered to subscriber . Subscribers aim to express their interest on the event and further receive the relevant information from publisher based on their interest . In addition , there is a middleware entity in the publish- subscribe system, called router . Router receives communications indicating a service from publishers and communications indicating an interest from subscribers , process and route them .
[0037] The system of the example of figure 1 is typically processing very large numbers of communications . Thus , there is a need for high computing power and memory usage . The system of figure 1 is configured to use internally or externally stored large language models in a manner that the model processing communications stores necessary information in an efficient manner . Different large language models may be having different processing capabilities and memory amount . Furthermore , the geographical location of data centers where the large language model s are hosted may have an effect on efficiency, for example , because of communication latencies . The large language model instances may be available on geographically distributed servers and cloud platforms requiring scheduling of their use considering geographical requirements , network latency, processing latency . Some processing tasks may have stricter security and privacy requirements requiring local or very close-by large language model serving capability .
[0038] The subscription, advertisement , event and other service processing tasks may require the use of specific large language models that are expert for the specific type of content . This requires the capability of mapping processing tasks to the specific large language model types and instances . As mentioned, different large language models may have different capabilities . For example , a specific large language model may have reduced memory because it is focusing on a specific purpose instead of general coverage .
[0039] Because of this Memory requirements and memory usage of large language models may vary requiring the packing / compressing and batching of inputs to large language models in parallel and in sequence according to the avai lable resources . Some of the large language models , typically designed for specific tasks and / or resource constrained environments , may have significant memory constraints that must be taken into account .
[0040] Packing and compressing may also include other optimi zation mechanisms . For example , it is possible to process all received communications so that it is possible to determine if they have common data items within the communications . When the communications have common data items , it is possible to store the data item only once and refer to the stored data item from multiple communications . In this manner communications do not need to store the data item multiple times and the memory usage is reduced . Furthermore , in some implementations it is possible to use a shared memory space for storing these shared data items . In order to identify the data items that can shared it is possible to process a plurality of communications or limit processing to communications that comprise further communications . For example , it is possible to have a subscription comprising a plurality of subscriptions .
[0041] Figure 2 comprises an example of flow chart of a method for large language model memory usage . In the method the router , such as a neural router or a neural routing apparatus , receives at least one communication from a subscriber 202 and a publisher 204 . The communications from a publisher comprise events , messages , advertisements or similar . The communications from a subscriber comprise interests . Optionally the arrangement may be extended so that advertisements that indicate the availability of a publisher of content are transmitted . The advertisements are used setup state for processing events published under the advertisements . This makes the system more deterministic and helps in the management and scaling of the operations .
[0042] When the communications have been received from both subscribers and publishers , then the neural routing apparatus where the method is implemented determines a destination to at least one publisher communication . The communications may have been received directly or indirectly . Thus , a communication received from a subscriber or publisher maybe received through another router . The determination of the destination may be done using a neural router, such as the neural router of the example of figure 1 .
[0043] For processing purposes the example method of figure 2 stores the received communications to a memory . The memory may be a local memory or a memory located in a data center where the large language model used for processing is hosted . In order to achieve this the method matches 206 at least one communication in the set received from the subscriber with a communication in the set received from the publisher . When matching the received sets are divided 208 into batches and stored the divided batches into context memory of a large language model . This can be done , for example , by dividing sets received from subscribers and publishers into batched . Divided sets are distributed into smal ler batches for efficient processing . Batch si ze is determined by available large language model memory . Thus , when using different large language models for different sets , the batch si ze may be different . The batch si zes for subscriptions , advertisements and events are set based on static or dynamic configuration policy . The batch si ze is set so that the instruction, subscription batch and advertisement batch, or subscription batch and event batch fit into the available memory . Furthermore , the batch si ze can be optimi zed at runtime .
[0044] After receiving the set comprising at least one communication from a subscriber the communications may be processed . The communications comprise subscriptions that often comprise data that is shared among several subscriptions . For this data, it is possible to initiali ze a shared prompt memory . The shared prompt memory may be store in the memory of the neural routing apparatus or in a data center . When the shared prompt memory is in use , i . e . appropriately initiali zed at the start of the system or whenever needed, the neural routing apparatus preprocesses the received sets and the communications in the sets .
[0045] The neural routing apparatus tunes and optimi zes the instruction set based on large language model capabilities and characteristics , taking into account the shared prompt memory for storing common data used across multiple subscriptions . Shared prompt memory use may be identified in the subscriptions . Shared prompt memory constraint indicates the values of matching events that are stored in the shared prompt memory . In thi s manner each of the subscriptions us ing shared prompt memory can indicate that latest relevant data values can be fetched form the shared prompt memory and need not to be stored in each of the subscriptions . The large language model can be instructed to provide shared prompt memory return data from the matching process that is stored in the shared prompt memory .
[0046] Figure 3 discloses a f low chart of an example method for managing memory . The method is using a shared prompt memory described in the above , however , this is only an example and the method can be implemented also without shared prompt memory . The method is initiated by setting caching configurations 302 . The router is configured to determine a caching strategy based on cache si ze limits , at least one eviction policy and a cache location . Furthermore , the configuration step may include determination if the shared prompt memory is used and which are the parameters of the shared prompt memory .
[0047] The next step is resource identification for caching 304 . At this step, resources that could benefit from caching, for example , frequently accessed subscription information, precomputed results , or lookup tables , are identified . It is possible to identify resources for caching before they are actually used so that they are readily available . This provides better performance as cache memory is typically highspeed memory that facilitates fast acces s to the data . Similarly, it is possible to have an optional step to identify resources that could benefit from shared prompt memory . The benefit of the shared prompt memory is that the data used multiple times in different subscriptions or other process parts can be stored only once . This step may involve determining the data that is repeating itself .
[0048] After the resource determination cache is initiali zed 306 . At this point , the router is configured to initiali ze a cache according to the determined caching strategy on step 302 and identi fied at least one resource on step 304 . When the method is using shared prompt memory, the possible initiali zation of the shared prompt memory may be done at this step .
[0049] After initiali zation dynamic memory management is configured 308 . To provide a dynamic memory management , the apparatus is set to configure a monitoring policy, a prediction policy or a resource allocation policy based on the current system configuration and current workload .
[0050] In the fifth step is to group the sets received from the subscribers and the sets received from the publishers into batches 310 . At this step , if not done in a separate step earlier, the data items to be stored in the shared prompt memory are determined from the sets received from the subscribers . The batches are formed so that if the batch comprises subscriptions having data in shared prompt memory the batch comprises an indication about the use of shared prompt memory .
[0051] After producing the batches the method comprises matching the batches and packing to large language model 312 . The router matches each of the batches comprising sets received from a subscriber with at least one batch received from a publisher, and packs the matches with an instruction to a large language model invocation .
[0052] On the seventh step, a set of large language model invocations are processed . The invocations are processed in parallel 314 or in sequence 316 based on the schedule . In case of parallel processing 314 , the process initiates by creating a plurality of instances to process each large language model invocations in a batch concurrently, and configuring synchroni zation and resource management for the instances for process safety . In case of sequential proces sing 316 , large language model invocations are processed sequentially in a predefined order .
[0053] Scheduling of the large language model invocations is important element of the neural router system . The scheduler addresses the requirements for task heterogeneity, memory use, and the processing and operational considerations of the LLM invocation . A scheduling subsystem is responsible for selecting the most suitable large language model types given the input subscriptions , advertisements , events , and other services respectively . The specific chosen type for a given subscription / advertisement / event / service determines the large language model specific properties and optimi zations . The runtime characteristics that are monitored online determine which specific available instance of the chosen large language model type is used for invocations .
[0054] Large language models are stateless resources allowing flexible invocations where the state is at the requesting side . The actual invocations are prepared in the packing and batching proces s that takes the memory limitations into account . This results in a set of specific invocations that are then executed in parallel or in sequence based on the operational requirements given in online system configuration . The online configuration can be optimi zed at runtime . As end-result the system provides support for multi-scale distributed large language model usage , which scales very well for parallel processing .
[0055] After the batches are processed and the large language model invocations are complete , cache validation and update 318 are performed . The router is set to validate and update cached data based on changes that occurred during processing of the large language model invocations .
[0056] The router then performs a cleanup and releases resources 320 . It releases resources after processing of the large language model invocations is complete and these resources are not needed anymore . By releasing the resources they are not left unused but can be used again in the future processing . After the clean up and releasing resources also the cache is cleaned up 322 . The router is configured to remove unused or outdated entries based on at least one eviction policy and cache si ze limits .
[0057] Finally, a destination is determined for the at least one communication received from a subscriber and route the at least one communication received from a publisher to the determined destination 324 . When determining the destination, the router is configured to input the received communications to a large language model stored in the at least one memory .
[0058] The above discussed steps may be performed sequentially or in parallel . Furthermore , some of the steps may be associated with supporting tasks , such as monitoring cache performance , memory allocation or correct number of clusters formed for the batches . For example , cache hit / miss rates , latency, and other relevant metrics can be monitored continuously or at regular intervals to evaluate its effectiveness and identify potential bottlenecks or optimi zation opportunities . Likewise the number of clusters that are used can be increased or decreased according to the need .
[0059] The above mentioned method may be implemented as computer software which is executed in a computing device able to communicate with a mobile device . When the software is executed in a computing device it is configured to perform the above described inventive method . The software is embodied on a computer readable medium so that it can be provided to the computing device , such as the router 102 of figure 1 .
[0060] As stated above , the components of the exemplary embodiments can include computer readable medium or memories for holding instructions programmed according to the teachings of the present inventions and for holding data structures , tables , records , and / or other data described herein . Computer readable medium can include any suitable medium that participates in providing instructions to a processor for execution . Common forms of computer-readable media can include , for example , a floppy disk, a flexible disk, hard disk, magnetic tape , any other suitable magnetic medium, a CD- ROM, CD±R, CD1RW, DVD, DVD-RAM, DVD1RW, DVD1R, HD DVD, HD DVD-R, HD DVD-RW, HD DVD-RAM, Blu-ray Disc, any other suitable optical medium, a RAM, a PROM, an EPROM, a FLASH-EPROM, any other suitable memory chip or cartridge , a carrier wave or any other suitable medium from which a computer can read .
[0061] It is obvious to a person skil led in the art that with the advancement of technology, the basic idea of the virtual memory management in neural routing apparatus may be implemented in various ways . The virtual memory management in neural routing apparatus and its embodiments are thus not limited to the examples described above ; instead they may vary within the scope of the claims .
Claims
CLAIMS1 . A router comprising : at least one processor configured to execute computer programs ; and at least memory configured to store data and a computer program code comprising instructions which, when executed by the at least one processor, is configured to cause the router to : receive a set comprising at least one communication from a subscriber ; receive a set comprising at least one communication from a publisher ; match at least one communication in the set received from the subscriber with a communication in the set received from the publisher ; wherein when matching, the router is further configured to : divide the received sets into batches ; and store the divided batches into a context memory of a large language model stored in the at least one memory; wherein the router is configured to determine the batch si ze based on the available context memory so that at least one batch comprising a communication from the subscriber and a communication from the publisher fit into the context memory .2 . The router according to claim 1 , wherein the instructions which, when executed by the at least one processor, are further configured to cause the router to : determine a caching strategy comprising at least one of the following : cache si ze limits , at least one eviction policy and a cache location .3 . The router according to claim 2 , wherein the instructions which, when executed by the at least one processor, are further configured to cause the router to :identify at least one resource to be stored into a cache .4 . The router according to claim 3 , wherein the instructions which, when executed by the at least one processor, are further configured to cause the router to :Initiali ze a cache according to the determined caching strategy and identified at least one resource .5 . The router according to any of preceding claims 1 - 4 , wherein the instructions which, when executed by the at least one processor, are further configured to cause the router to : store at least one data item in a communication received from the subscriber in the cache memory before processing of data items is initiated .6 . The router according to any of preceding claims 1 - 4 , wherein the instructions which, when executed by the at least one processor, are further configured to cause the router to : configure at least one of the following : a monitoring policy, a prediction policy or a resource allocation policy based on the current system configuration and current workload .7 . The router according to any of preceding claims 1 - 5 , wherein the instructions which, when executed by the at least one processor, are further configured to cause the router to : group the sets received from the subscribers and the sets received from the publishers into batches .8 . The router according to claim 6 , wherein the instructions which, when executed by the at least one processor, are further configured to cause the router to : match each of the batches comprising sets received from a subscriber with at least one batch received from a publisher ;pack the matches with an instruction to an large language model invocation .9 . The router according to claim 7 , wherein the instructions which, when executed by the at least one processor, are further configured to cause the router to : create a plurality of instances to process each large language model invocation in a batch concurrently; and configure synchroni zation and resource management for the instances .10 . The router according to claims 8 or 9 , wherein the instructions which, when executed by the at least one processor, are further configured to cause the router to : release resources after processing of the large language model invocations is complete .11 . The router according to any of preceding claims 1 - 10 , wherein the instructions which, when executed by the at least one processor, are further configured to cause the router to : identifying at least one data item in at least two communications .12 . The router according to claim 11 , wherein the instructions which, when executed by the at least one processor, are further configured to cause the router to : store the identified at least one data item in a shared memory13 . The router according to any of preceding claims 1 - 12 , wherein the instructions which, when executed by the at least one processor, are further configured to cause the router to : determine a destination for the at least one communication received from a subscriber, wherein when determining the destination, the router is further configured to input the received communications to alarge language model stored in the at least one memory; and route the at least one communication received from a published to the determined destination .14 . A method for managing memory comprising : receiving a set comprising at least one communication from a subscriber ; receiving a set comprising at least one communication from a publisher ; matching at least one communication in the set received from the subscriber with a communication in the set received from the publisher ; wherein when matching, the method further comprises : dividing the received sets into batches ; and storing the divided batches into a context memory of a large language model stored in the at least one memory; wherein the method further comprises : determining the batch si ze based on the available context memory so that at least one batch comprising a communication from the subscriber and a communication from the publisher fit into the context memory .15 . A computer program comprising instructions which, when executed by a computing device , are further configured to cause the computer device to perform a method according to claim 14 .