Node, method, and multi-agent intelligent network to provide signalling protocol for providing distributed attention
The multi-agent intelligent network with layered nodes using multi-head linear attention models facilitates efficient, scalable, and secure data processing by sharing processed information, addressing bandwidth and privacy constraints in conventional systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-10-18
- Publication Date
- 2026-04-23
AI Technical Summary
Conventional communication systems face challenges in exchanging data between distributed agents due to bandwidth limitations and privacy constraints, making it difficult to achieve outputs identical to centralized models without raw data transmission.
A multi-agent intelligent network with K layers of nodes, each processing input information through a multi-head linear attention model, shares processed output information across layers to maintain privacy and reduce communication overhead, enabling distributed attention.
This approach allows efficient, scalable, and privacy-preserving data processing for complex tasks, reducing network bandwidth requirements and enhancing computational efficiency while maintaining data security.
Smart Images

Figure EP2024079447_23042026_PF_FP_ABST
Abstract
Description
[0001] NODE, METHOD, AND MULTI-AGENT INTELLIGENT NETWORK TO PROVIDE SIGNALLING PROTOCOL FOR PROVIDING DISTRIBUTED ATTENTION
[0002] TECHNICAL FIELD
[0003] The present disclosure relates generally to the field of wireless communication network management and, more specifically, to a node and a method for a multi-agent intelligent network to provide a signalling protocol for providing distributed linear attention, such as by providing a signalling for inference using a distributed neural network model via linear attention.
[0004] BACKGROUND
[0005] In rapidly advancing domain of communication systems, upcoming future communication systems are anticipated to heavily rely on cooperation of distributed agents that are used to combine processing compatibility of each of the intelligent agents and produce results. Moreover, each of the distributed agents are used to jointly process multimodal data (e.g., images, videos or sensing data collected by each of the distributed agent) to adapt changing environments, solve various complex tasks, and share obtained data with other distributed agents, such as by distributing all relevant data (i.e., tokens) across the distributed agents.
[0006] Conventional communication systems, such as conventional self-attention mechanism cannot be computed in a distributed manner and require all the relevant data (i.e., the tokens) to be transmitted to a central server for joint computation. On the other hand, the transmission of the relevant data from each of the distributed agents to the central server is not feasible due to limitations of communication bandwidth and privacy constraints. Moreover, certain attempts have been made to overcome the communication bandwidth and privacy constraints, such as by using self-attention of a transformer, ring attention, retentive networks, message passing neural networks, and the like. However, such attempts fail due to many reasons, such as unavailability of the relevant data at a single place, complex computational cost, restrictive data flow, and the like. Therefore, there exist a technical problem of how to exchange data between the distributed agents in order to obtain an output that is identical to the output obtained by a centralized model with all input data at one place.
[0007] Therefore, in light of the foregoing discussion, there exists a need to overcome the aforementioned drawbacks associated with the conventional nodes and conventional methods to provide a signalling protocol for providing distributed linear attention.
[0008] SUMMARY
[0009] The present disclosure provides a method and a node for providing a signalling protocol for a multi-agent intelligent network providing distributed attention. The present disclosure provides a solution to the existing problem of to exchange data between the distributed agents in order to obtain an output that is identical to the output obtained by a centralized model with all input data at one place. An objective of the present disclosure is to provide a solution that overcomes at least partially the problems encountered in the prior art and provides the node of the multi-agent intelligent network and the method for the node of the multi-agent intelligent network to provide a signalling protocol for multi-agent intelligent network providing distributed attention.
[0010] One or more objectives of the present disclosure are achieved by the solutions provided in the enclosed independent claims. Advantageous implementations of the present disclosure are further defined in the dependent claims.
[0011] In one aspect, the present disclosure provides a method for providing a signalling protocol for a multi-agent intelligent network providing distributed attention, the multi-agent intelligent network comprising K layers of nodes, each layer comprising a plurality of nodes and there is a first input layer of nodes, one or more intermediate layers of nodes and a final layer of nodes. Moreover, each node is arranged to process input information through a multi-head linear attention model to provide output information. The nodes of the first layer are arranged to receive input tokens as input information, the nodes of the one or more intermediate layers are arranged to receive input information from one or more nodes of a previous layer and send output information to one or more nodes of a subsequent layer, and the nodes of the final layer are arranged to provide output information as output labels. Moreover, the method includes providing the nodes of the first layer with an acquired input sequence as the input tokens and then, for each layer, process the input in each node in the current layer to generate the output information of the node, share output information with nodes of the same current layer, and then share output information as input information to the one or more nodes of the subsequent layer, whereby the nodes of the final layer provides processed output labels.
[0012] Advantageously, the method is used for a signalling protocol in a multi-agent intelligent network that utilizes distributed attention with K layers of nodes, where each layer processes input information using a multi-head linear attention model to generate output labels. Moreover, the method involves processing input tokens in the first layer, sharing output information within the same layer, and then passing the output information to the subsequent layer thereby, ensuring that the input tokens are processed effectively, and the output labels are shared efficiently. In addition, the method allows for distributed processing of input data across multiple nodes and layers, enabling the multi-agent intelligent network to handle large-scale tasks that would be impractical or impossible for a single centralized system. Furthermore, by distributing the computation across multiple nodes the utilization of the computational resources are improved leading to an enhanced data processing times and improved scalability. The protocol is used to minimize the data transfer between nodes by sharing only processed output information rather than raw input data, which can significantly reduce network bandwidth requirements. As the input data is processed locally at each node and only output information is shared, the method is used to maintain data privacy and security, which is crucial in many applications. Moreover, the multi-layer, multi-node structure allows for easy scaling of the network by adding more nodes or layers as needed, without fundamentally changing the protocol and allows for decentralized operation, reducing the need for a central coordinator and potentially improving fault tolerance. As a result, the method can be used in a wide range of applications, such as from video processing and autonomous driving to distributed intelligent computing and smart manufacturing. Additionally, the method is used to provide a signalling protocol for providing distributed attention in multi-agent intelligent networks, enabling an efficient, scalable, and privacy-preserving processing of complex tasks.
[0013] In another aspect, the present disclosure provides a multi-agent intelligent network configured to provide a signalling protocol for providing distributed attention in the multi-agent intelligent network, the multi-agent intelligent network comprising K layers of nodes, each layer comprising a plurality of nodes and there is a first input layer of nodes, one or more intermediate layers of nodes and a final layer of nodes. Moreover, each node is arranged to process input information through a multi-head linear attention model to provide output information, the nodes of the first layer are arranged to receive input tokens as input information, the nodes of the one or more intermediate layers are arranged to receive input information from one or more nodes of a previous layer and send output information to one or more nodes of a subsequent layer, and the nodes of the final layer are arranged to provide output information as output labels. Additionally, the multi-agent intelligent network is configured to provide the nodes of the first layer with an acquired input sequence as the input tokens and then, for each layer, process the input in each node in the current layer to generate the output information of the node, share output information with nodes of the same current layer, and then share output information as input information to the one or more nodes of the subsequent layer, whereby the nodes of the final layer provides processed output labels.
[0014] Advantageously, the multi-agent intelligent network is configured to provide a signalling protocol for distributed attention offers several technical advantages. By structuring the network into K layers of nodes, including a first input layer, intermediate layers, and a final layer, it enables efficient distributed processing of complex tasks. Each node's ability to process input through a multi-head linear attention model allows for sophisticated data analysis at every stage of the network. The multi-agent intelligent network is configured to facilitate seamless information flow, with input tokens entering at the first layer, being processed and shared within each layer, and then passed on to subsequent layers that supports parallel processing within layers while maintaining a logical progression of information through the multi-agent intelligent network. The sharing of output information between nodes in the same layer allows for collaborative processing and potentially more robust results. Moreover, the multi-agent intelligent network supports data privacy and security in order to provide the multi-agent intelligent network for a wide range of real- world problems, from image and speech recognition to complex decision-making tasks in fields, such as autonomous driving or financial analysis. As a result, the multi-agent intelligent network is configured to provide a scalable and flexible solution for distributed attention processing while maintaining data privacy and efficient resource utilization.
[0015] In yet another aspect, the present disclosure provides a node configured to be used in a multi-agent intelligent network configured to provide a signalling protocol for providing distributed attention in the multi-agent intelligent network, the multiagent intelligent network comprising K layers of nodes, each layer comprising a plurality of nodes and there is a first input layer of nodes, one or more intermediate layers of nodes and a final layer of nodes. Moreover, the node is arranged to process input information through a multi-head linear attention model to provide output information, and if the node is in the first layer, the node is configured to receive input tokens as input information, and if the node is in the one or more intermediate layers, the node is configured to receive input information from one or more nodes of a previous layer and send output information to one or more nodes of a subsequent layer, and if the node is in the final layer, the node is configured to provide output information as output labels. Furthermore, if the node is in the first layer, the node is configured to receive an acquired input sequence as the input tokens and, for each layer, the node is further configured to process its input to generate the output information of the node, share output information with nodes of the same current layer, and then share output information as input information to the one or more nodes of the subsequent layer, whereby if the node is in the final layer, the node is configured to provide processed output labels.
[0016] The node achieves all the advantages and technical effects of the method of the present disclosure.
[0017] It is to be appreciated that all the aforementioned implementation forms can be combined.
[0018] It has to be noted that all devices, elements, circuitry, units, and means described in the present application could be implemented in the software or hardware elements or any kind of combination thereof. All steps which are performed by the various entities described in the present application, as well as the functionalities described to be performed by the various entities are intended to mean that the respective entity is adapted to or configured to perform the respective steps and functionalities. Even if, in the following description of specific embodiments, a specific functionality or step to be performed by external entities is not reflected in the description of a specific detailed element of that entity that performs that specific step or functionality, it should be clear for a skilled person that these methods and functionalities can be implemented in respective software or hardware elements or any kind of combination thereof. It will be appreciated that features of the present disclosure are susceptible to being combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims.
[0019] Additional aspects, advantages, features, and objects of the present disclosure would be made apparent from the drawings and the detailed description of the illustrative implementations construed in conjunction with the appended claims that follow.
[0020] BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The summary above, as well as the following detailed description of illustrative embodiments, is better understood when read in conjunction with the appended drawings. For the purpose of illustrating the present disclosure, exemplary constructions of the disclosure are shown in the drawings. However, the present disclosure is not limited to specific methods and instrumentalities disclosed herein. Moreover, those in the art will understand that the drawings are not to scale. Wherever possible, like elements have been indicated by identical numbers. Embodiments of the present disclosure will now be described, by way of example only, with reference to the following diagrams wherein:
[0022] FIG. 1 is a block diagram that illustrates a multi-agent intelligent network providing distributed attention, in accordance with an embodiment of the present disclosure;
[0023] FIG. 2 is a diagram that illustrates a method for providing a signalling protocol for a multi-agent network providing distributed attention, in accordance with an embodiment of the present disclosure;
[0024] FIG. 3 is a diagram that depicts a node configured to be used in a multi-agent intelligent network configured to provide a signaling protocol for providing distributed attention in the multi-agent intelligent network, in accordance with an embodiment of the present disclosure;
[0025] FIG. 4A and 4B is a diagram that illustrates an exemplary scenario of signalling protocol for distributed inference over a network of devices with pieces of a large model with K layers of multi-head linear attention, in accordance with an embodiment of the present disclosure; and
[0026] FIG. 5 is a diagram that illustrates an exemplary scenario of a vehicle monitored by a network of cameras, in accordance with an embodiment of the present disclosure.
[0027] In the accompanying drawings, an underlined number is employed to represent an item over which the underlined number is positioned or an item to which the underlined number is adjacent. A non-underlined number relates to an item identified by a line linking the non-underlined number to the item. When a number is non-underlined and accompanied by an associated arrow, the non-underlined number is used to identify a general item at which the arrow is pointing.
[0028] DETAILED DESCRIPTION OF EMBODIMENTS
[0029] The following detailed description illustrates embodiments of the present disclosure and ways in which they can be implemented. Although some modes of carrying out the present disclosure have been disclosed, those skilled in the art would recognize that other embodiments for carrying out or practicing the present disclosure are also possible.
[0030] FIG. 1 is a block diagram that illustrates a multi-agent intelligent network providing distributed attention, in accordance with an embodiment of the present disclosure. With reference to FIG. 1, there is shown a multi-agent intelligent network 100 providing distributed attention. The multi-agent intelligent network 100 includes a multi-head network model 104 and a server 108. The multi-head network model 104 includes a plurality of model layers, such as a first layer 114A, a second layer 114B, kth layer 114K, nth layer 114N, and n+1 layer 114N+1. Furthermore, the multi-agent intelligent network 100 includes system layers, such as a first system layer 116A and a second system layer 116B that further includes one or more nodes, such as a first node 118A, a second node 118B, a third node 118C, and a fourth node 118D.
[0031] There is provided the multi-agent intelligent network 100 providing distributed attention. The multi-agent intelligent network 100 provides distributed attention by allowing multiple nodes to independently process and focus on different parts of the data while communicating and collaborating with each other to provide an enhanced efficiency, robustness, and data privacy while handling complex and large-scale operations in a decentralized and efficient manner with reduced communication overhead. The multi-agent intelligent network 100 includes the multi-head network model 104 that is configured to receive a model input data 102 (i.e., a sequence of input tokens), such as a query translated into tokens and provides a model output data 106, which is a sequence of output tokens translated into a response that is further utilized for computing the distributed attention. Moreover, the server 108 includes a scheduler 110, which is configured to control, schedule, or train the multi-head network model 104 based on the information (e.g. , the information received at operationl 12), such as information on system capabilities, number or nodes, and the like. The multi-head network model 104 is deployed to compute distributed attention within the multi-agent intelligent network 100 and includes multiple model layers, such as the first layer 114A, the second layer 114B, the kth layer 114k, the nth layer 114N, and the nth+1 layer 114N+1. Moreover, each layer of the model layers includes multiple head and is configured to transmit the set of local inputs along with the memory states (i.e., partially processed tokens) to the other layer of the multi-head network model 104, such as at operation 120. For example, at operation 120A, the nth layer 114N is configured to receive the local input embeddings from the previous node and, further, at operation 120B, transmits the local input embeddings (or the partially processed tokens) to the nth+1 layer 114N+1. Moreover, the detailed transmission of the local input embeddings is given and explained in detail in FIG. 2.
[0032] Furthermore, each of the model layers of the multi-head network model 104 includes at least one system layer, such as the first system layer 116A and the second system layer 116B, that are assigned to one or more nodes. In other words, each of the node executes one or more model layers and one or more heads, such as a first head 124A, a second head 124B, a third head 124C, and a fourth head 124D, are assigned to the one or more nodes. Moreover, each of the node executes all heads of the executed model layers. For example, the first head 124A, the second head 124B, the third head 124C, and the fourth head 124D are assigned to the second node 118B. As a result, the multi-agent intelligent network 100 is configured to provide distributed attention by allowing each node within the multi-agent intelligent network 100 to determine the combined node memory state and further transmit the same to the other nodes along with the set of local memory states.
[0033] In operation, the multi-agent intelligent network 100 is configured to receive a set of local input embeddings for the node. In an example, the multi-agent intelligent network 100 is configured to receive the set of local input embeddings (e.g., Xn, XI- X3) for the first node 118A. In another example, the multi-agent intelligent network 100 is configured to receive the set of local input embeddings for the second node 118B. Similarly, the multi-agent intelligent network 100 is configured to receive the set of local input embeddings for the third node 118C and the fourth node 118D. The node (e.g., the first node 118A, the second node 118B, the third node 118C, and the fourth node 118D) is configured to receive the set of local input embeddings to initialize the data processing at the local node level. Moreover, the set of local input embeddings refers to a set of tokens that represents the information about the local data input of the one or more nodes within the multi-agent intelligent network 100. As a result, the node is configured to receive the set of local input embeddings to process and interpret the local data efficiently for further computations and decision-making processes within the multi-agent intelligent network 100.
[0034] Furthermore, the multi-agent intelligent network 100 is configured to apply positional embedding to the set of local input embeddings for the node. The positional embedding, which is applied to the set of local input embeddings for the node is configured to provide the information about the position of each input element in order to identify the sequence order to provide accurate and context-aware predictions. The multi-agent intelligent network 100 calculates positional embeddings for each position in the input sequence. In an implementation, the positional embeddings can be pre-defined (e.g., sinusoidal functions) or learned during training. Thereafter, the calculated positional embeddings are added to the corresponding local input embeddings. As a result, the positional embeddings are used to capture the order of input elements to allow the multi-agent intelligent network 100 to handle a wide range of sequence lengths and structures. Additionally, the integration of positional information allows the multi-agent intelligent network 100 to handle complex patterns and relationships, providing an accurate and flexible approach to processing sequential data.
[0035] Furthermore, the multi-agent intelligent network 100 is configured to determine a set of keys, Kn, as Kn= WK Xn, WK being a learned key matrix WK. In an implementation, the set of local input embeddings (i.e., token Xn) are multiplied with the set of learned key matrix (i.e., WK) in order to determine the set of keys, for example, as shown in the equation (1) given below: -
[0036] The determination of the set of keys involves transforming the set of local input embeddings by using a learned key matrix to provide the set of keys that are further utilized to compute the memory states. Moreover, by using a learned matrix enables the multi-agent intelligent network 100 to provide an efficient and effective representation of the data that can be used for computing the linear attention.
[0037] Furthermore, the multi-agent intelligent network 100 is configured to determine a set of values, Vn, as Vn= Wv Xn, Wv being a learned value matrix, Wv. The set of values is used to encode the information that is used to compute the combined memory states that are further utilized to compute the linear computation. In an implementation, the set of local input embeddings (i.e., token Xn) are multiplied with the set of learned value matrix (i.e., Wv) in order to determine the set of keys, for example, as shown in the equation (2) given below: -
[0038] Moreover, the learned value matrix allows the encoding of the set of input embeddings into the set of learned values that are used to capture relevant information to handle varying input sizes and complexities without a significant increase in computational cost, such as by further computing the combined memory states. As a result, the determination of the set of values is used to improve the efficiency, scalability, and overall performance of the node while minimizing the communication overhead within the multi-agent intelligent network 100.
[0039] Furthermore, the multi-agent intelligent network 100 is configured to determine a set of queries, Qn, as Qn =WQ „, WQ being a learned query matrix, WQ, for the node. The set of values is used to encode the information that is used to compute the combined memory states that are further utilized to compute the linear computation. In an implementation, the set of local input embeddings (i.e., token Xn) are multiplied with the set of learned queries matrix (i.e., Wq) in order to determine the set of keys, for example, as shown in the equation (3) given below: -
[0040] Qn = WQXn(3)
[0041] Moreover, the learned value matrix allows the encoding of the set of input embeddings into the set of learned queries that are used to capture relevant information to handle varying input sizes and complexities without a significant increase in computational cost, such as by further computing the combined memory states. As a result, the determination of the set of queries is used to improve the efficiency, scalability, and overall performance of the node while minimizing the communication overhead within the multi-agent intelligent network 100.
[0042] Moreover, the multi-agent intelligent network 100 is characterized in that the multi-agent intelligent network 100 is further configured to determine a set of local memory states S’n, as S’n = XTn(WTKWv)Xn, for the node. In an implementation, the set of local memory states refers to input partial tokens that are associated with each node that is derived through a series of matrix multiplications involving the input token and a set of learned matrices, such as the set of values, set of keys, and the set of queries. Moreover, the local memory states are configured to encapsulate the information contained in the input tokens in a manner that is suitable for further processing within the node and for sharing the same with other nodes of the multi-agent intelligent network 100. Firstly, the set of input embeddings (or tokens) are received by the node locally. Thereafter, the set of keys, the set of values, and the set of queries are determined in order to further determine the set of local memory states for each node, such as the first node 118A, the second node 118B, the third node 118C, and the fourth node 118D of the multiagent intelligent network 100. The set of local memory states allows the node to perform linear distribution of attention in order to allow information sharing within the multi-agent intelligent network 100 locally, thereby maintaining data privacy within the multi-agent intelligent network 100. As a result, the set of local memory states is determined to provide an efficient local processing, distributed attention, data privacy, and reduced communication overhead in the multi-agent intelligent network 100. In accordance with an embodiment, the multi-agent intelligent network 100 is further configured to determine the set of local memory states S’n, as S’n = KTnVnfor the node. Each node computes the key vector (Kn) and the value vector (Vn) from the set of input embeddings by using the learned matrices. Thereafter, the key vector is multiplied by the value vector in order to determine the local memory state. Moreover, by calculating local memory states as S'n= KTnVn, the node of the multi-agent intelligent network 100 is configured to efficiently encapsulate the information from the input tokens of each node. As a result, the multi-agent intelligent network 100 is configured to ensure effective distributed data processing with an improved overall network performance.
[0043] Furthermore, the multi-agent intelligent network 100 is configured to receive a set of node memory states (Sn) from the one or more other nodes. In an example, the multi-agent intelligent network 100 is configured to receive the set of node memory states (Sn) from the first node 118A. In another example, the multi-agent intelligent network 100 is configured to receive the set of node memory states (Sn) from the first node 118A and the second node 118B. In yet another example, the multi-agent intelligent network 100 is configured to receive the set of node memory states (Sn) from the first node 118A, the second node 118B, and the third node 118C. Moreover, by receiving the set of node memory states from the one or more nodes, the node is configured to integrate external information with the local memory state of the node in order to facilitate a distributed linear attention mechanism without sharing raw data.
[0044] Furthermore, the multi-agent intelligent network 100 is configured to determine a set of combined node memory states (Sn) as Sj = Sj-i + S’n, and So is determined as the sum SS' for the node. In other words, each node in the multi-agent intelligent network 100 determines the local memory states from the set of input embeddings by using learned matrices and transmits these memory states to one or more other nodes. Finally, the received memory states are integrated with the local memory states of the node to further determine the combined memory states that are used to compute the linear attention. The set of combined node memory states is used to facilitate distributed attention in order to ensure that each node has access to a comprehensive set of node combined memory states with an efficient and effective data integration and processing with reduced communication overhead, such as by transmitting aggregated states rather than raw data. Additionally, the determination of the set of node combined memory states is used to maintain data privacy, such as by sharing the combined node memory state with an improved overall performance of the multi-agent intelligent network 100.
[0045] In accordance with an embodiment, the multi-agent intelligent network 100 is further configured to determine the set of combined node memory states (Sn) as Sj = A(H)Sj-i + S’n, and H for the node and H is a Hadamard product and wherein A is an adaptation factor. The adaptation factor enables the adjustment of the received node memory states thereby allowing the multi-agent intelligent network 100 to scale the contributions of previous memory states based on current conditions or data characteristics dynamically. Moreover, the Hadamard product refers to an element- wise multiplication operation between two matrices or vectors of the same dimensions, where each element in the resulting matrix or vector is the product of the corresponding elements from the input matrices or vectors. As a result, the determination of the set of combined node memory states by using the Hadamard product and the adaptation factor is used to allow the multi-agent intelligent network 100 to improve the accuracy and relevance of the combined memory states in order to enhance the decision-making and data processing capabilities of the multi-agent intelligent network 100. In addition, the utilization of the Hadamard product and adaptation factor ensures an efficient and scalable data integration with reduced communication overhead and maintained data privacy.
[0046] In such an implementation, the adaptation factor A is ge*' and the g is a decay factor and j is a rotating factor. The decay factor attenuates the contributions of previous memory states, effectively reducing the influence over time or distance, which helps to prioritize more recent or relevant information (or the node memory states received from the one or more nodes). Moreover, the rotating factor refers to a phase shift that enables the temporal modulation in order to align or misalign with the phases of memory states, which is further used in applications involving periodic or cyclical data patterns. As a result, the adaptation factor allows the multi-agent intelligent network 100 to control the attenuation and phase shift of the combined memory states in order to prioritize the received memory states and improve the overall decision-making and processing efficiently and effectively.
[0047] In accordance with an embodiment, the multi-agent intelligent network 100 is further configured to receive a counter (CTR=dj) for each received node memory state (S’), and So is determined as the sum SAd’S' for the node, and A is an adaptation factor. The counter for each received memory state allows the multi-agent intelligent network 100 to track the relevance or freshness of the data from the one or more nodes. Moreover, each node is configured to determine the combined memory state by summing the products of the adaptation factor, the counter, and the memory states that may include parameters, such as the decay factor and the rotation, to modulate the influence of each state. In such an implementation, the adaptation factor A is ge" and the g is a decay factor and j is a rotating factor. The decay factor reduces the contribution of older memory states, while the rotating factor represents the phase shift. As a result, the multi-agent intelligent network 100 is used to manage and prioritize the memory states, thereby enhancing the overall accuracy and relevance of the combined memory states. Additionally, such determination also reduces the risk of outdated information affecting current decisions and maintains flexibility and scalability within the multi-agent intelligent network 100.
[0048] Furthermore, the multi-agent intelligent network 100 is configured to provide at least a portion of the set of combined node memory states (Sn) to at least one next node for the node. In other words, the multi-agent intelligent network 100 is configured to determine the combined node memory state and further transmit the same to the next node within the multi-agent intelligent network 100. For example, the second node 118B is configured to receive at least the portion of the set of combined node memory states by the first node 118 A. As a result, by providing the combined memory states, the multi-agent intelligent network 100 is configured to adapt to changing environments and solve various challenges by leveraging the collective knowledge and insights of the nodes.
[0049] Furthermore, the multi-agent intelligent network 100 is configured to determine an output, On, as On = QnSnfor the node. By determining the output On, as On = QnSnfor the node, the multi-agent intelligent network 100 ensures that the output of each node includes both the local data (i.e., the local memory states) and the collective knowledge (i.e., the combined memory states) of the multi-agent intelligent network 100. As a result, the accuracy and relevance of the output are enhanced, leading to an improved overall network performance.
[0050] Advantageously, the multi-agent intelligent network 100 provides distributed attention with an efficient and reduced communication overhead. The multi-agent intelligent network 100 receives the set of input embeddings from the one or more nodes without directly sharing raw data in order to ensure that the sensitive information is not directly shared between nodes for maintaining the data privacy and data security. By sharing memory states instead of raw input data, the multi-agent intelligent network 100 is used to reduce the need for extensive data transmission, leading to minimized communication overhead and improved bandwidth utilization. The multi-agent intelligent network 100 provides a scalable and flexible accommodation of new nodes without requiring any changes to an existing network structure in order to provide seamless expansion and integration of additional nodes. In addition, each node processes data locally and in parallel, reduces the overall computation time and latency of the data and enhances the efficiency and responsiveness of the multi-agent intelligent network 100. The use of learned matrices (e.g., key, value, and query matrices) enables efficient parallel training on a server with large datasets, and once trained, the model parameters can be easily deployed across distributed nodes to ensure continuity of operations while handling complex computations with minimal disruption.
[0051] FIG. 2 is a diagram that illustrates a method for providing a signalling protocol for a multi-agent network providing distributed attention, in accordance with an embodiment of the present disclosure. With reference to FIG. 2, there is shown a diagram that illustrates a method 200 for providing a signalling protocol for a multi-agent intelligent network 100 providing distributed attention.
[0052] There is provided the method 200 for providing a signalling protocol for a multi-agent intelligent network 100 providing distributed attention. The signalling protocol is used to allow the multi-agent intelligent network 100 to provide an efficient, scalable, and privacy-preserving distributed processing. The signalling protocol is configured to minimize the communication overhead, such as by structuring data exchanges and ensures that sensitive local data remains private by preventing the data transmission to other nodes.
[0053] The multi-agent intelligent network 100 includes K layers of nodes and each layer includes a plurality of nodes and there is a first input layer of nodes, one or more intermediate layers of nodes and a final layer of nodes. In an implementation, the first input layer is used for receiving and processing the initial data or input tokens. Each node in the first input layer is used to process the local portion of the data and begins the exchange of information with other nodes in the multi-agent intelligent network 100. Moreover, the intermediate layers are configured to handle the sequential processing of tokens and states, where nodes in each layer communicate both within their own layer and with nodes in adjacent layers. Additionally, the final layer of nodes is configured to perform the computation by producing the output tokens, which represent the final processed result of the multi-agent intelligent network 100. As a result, such layered structure allows an efficient distributed processing, where each node focuses on a specific part of the computation and coordinates with others through a well-defined signalling protocol and the multiple layers ensures that the multi-agent intelligent network 100 can scale to handle large models and complex tasks while maintaining efficiency and accuracy in the distributed processing environment.
[0054] Moreover, each node is arranged to process input information through a multi-head linear attention model to provide output information. In an implementation, each node from the plurality of nodes is configured to capture complex relationships and dependencies between the tokens and such parallel processing is further used for handling large-scale data by ensuring that the node can attend to multiple aspects of the input at once thereby, improving both processing speed and accuracy. Moreover, after processing the input, each node generates the output information, which is then communicated to other nodes in the multiagent intelligent network 100, either within the same layer or in subsequent layers. Therefore, by utilizing the multi-head linear attention model, each node contributes to the overall computation in a distributed and coordinated manner, ensuring that the final output of the multi-agent intelligent network 100 that reflects the comprehensive analysis of all input data. As a result, an efficient distributed attention enables the multi-agent intelligent network 100 to handle complex tasks with high accuracy while maintaining scalability and minimal communication overhead.
[0055] The nodes of the first layer are arranged to receive input tokens as input information and the nodes of the one or more intermediate layers are arranged to receive input information from one or more nodes of a previous layer and send output information to one or more nodes of a subsequent layer. The nodes in the one or more intermediate layers are arranged to receive input information from one or more nodes in the previous layer. These nodes act as intermediaries in the distributed processing chain, taking the input information from the previous layer and using them as input to continue the computation. After processing the received input information, the intermediate nodes then send their output information to one or more nodes in the subsequent layer, ensuring the continuous flow of data through the network. As a result, the structured flow of the input information between layers allows the multi-agent intelligent network 100 to refine the data, eventually reaching the final layer where the complete output is produced, thereby ensuring seamless, scalable, and efficient processing throughout the multiagent intelligent network 100.
[0056] In accordance with an embodiment, the input information includes one or more tokens to be processed, and output information of a previous node, if any. In an implementation, the one or more tokens refers to the raw data or input that the node needs to analyse, such as by using the multi-head linear attention model. If the node is not in the first layer, then, in that case, the node is configured to receive output information from the previous node, which contains processed data that has been passed along the multi-agent intelligent network 100. As a result, the combination of input tokens and previous output information enables the node to perform the data processing tasks effectively, incorporating both new data and context from earlier computations to generate the output information.
[0057] In accordance with an embodiment, the output information of a node comprises memory states indicating information about processed input information and a counter which indicates the total number of processed tokens by the node. In an implementation, the memory states indicates detailed information about the processed input tokens, summarizing the data and the results of the computations. Moreover, the combined memory states are required for maintaining context and continuity across the multi-agent intelligent network 100. Additionally, the counter keeps track of the total number of processed tokens by the node, providing a numerical indicator of how much data has been handled, which is used to ensure that all the tokens are accounted for and processed correctly throughout the multi-agent intelligent network 100.
[0058] In accordance with an embodiment, the memory states include one memory state per head of the layer. Each memory state from the memory states corresponds to a specific attention head within the multi-head linear attention model that is used for capturing information related to the particular focus or aspect of the data that the head processes in order to allow the node to maintain detailed and differentiated information about the input tokens, as each head can attend to different parts or features of the data independently. As a result, the memory states with one memory head per head of the layer is used to ensure that the comprehensive representation of the data is accurately conveyed to subsequent nodes in the multi-agent intelligent network 100.
[0059] In accordance with an embodiment, each node is arranged to share to the nodes in the same layer when processing the output information of that node, an identifier of the node, a counter indicating a number of tokens processed so far at that layer, and memory states for each head of a multi-head layer of linear attention. In an implementation, the identifier of the node is configured to recognize the source of the information and manage communication effectively. In another implementation, the counter is used to provide a record of how many tokens the node has handled up to that point, ensuring that all tokens are accounted for and processed correctly. In yet another implementation, the memory states for each head of the multi-head layer of the linear attention is used to provide an information captured by each attention head, offering a detailed summary of the processed input data. Therefore, by sharing the identifier, the counter and the memory state for each head of the multi-head layer, the nodes within the same layer can collaboratively manage and synchronize the processing tasks thereby, ensuring a coherent and accurate flow of information throughout the multi-agent intelligent network 100 for effective distributed computation.
[0060] In accordance with an embodiment, at least one node is arranged to further share to the nodes in the same layer an identifier of a sub layer. In an implementation, the identifier is used to distinguish different sub-layers within the same layer, providing additional context and organization for the information being processed. By including the sub- layer identifier, the at least one node is configured to track and manage the data flow within the respective sub-layers effectively thereby, ensuring that the information is aligned and processed accurately, thereby enhancing the coordination and efficiency of data processing within the layer, facilitating precise and effective distributed computation.
[0061] In accordance with an embodiment, each node is arranged to share to the one or more nodes in the subsequent layer, an identifier of the sending device, and a counter of the output information. In an implementation, the identifier of the sending device is used to identify the origin of the output information, allowing nodes in the subsequent layer to properly recognize and integrate the incoming data from the specific node and the counter of the output information is used to manage and track the data movement to the next layer. As a result, the arrangement of each node is used to ensure that the nodes in the subsequent layer receive well-organized and contextually accurate information, facilitating effective data processing and maintaining the integrity of the distributed computation across the multi-agent intelligent network 100.
[0062] In accordance with an embodiment, at least one node is further arranged to share an identifier of the layer. The identifier of the layer is used to provide the context for the processing and interpretating the shared information. By including the layer identifier, the at least one node in the subsequent layers can handle the incoming data according to the originating layer thereby, ensuring proper coordination and alignment within the multi-agent intelligent network 100 in order to support effective communication and data management.
[0063] Furthermore, the nodes of the final layer are arranged to provide output information as output labels. After receiving and processing data through the preceding layers, the final layer nodes completes the distributed computation by generating the ultimate results or labels based on the processed input. Moreover, the output labels represent the final, comprehensive interpretation of the data as determined by the multi-agent intelligent network 100. As a result, the output labels are used to provide end-to-end computations along with the final conclusions or predictions of the multi-agent intelligent network 100.
[0064] At step 202, the method 200 includes providing the nodes of the first layer with an acquired input sequence as the input tokens and then process the input in each node in the current layer to generate the output information of the node for each layer, such as at step 204. The input token is distributed to all the nodes of the first layer and each node processes the one or more tokens that are allocated to the corresponding nodes thereby, capturing relevant features and creating memory states that summarize the information processed. Once processed, the output is passed on to the nodes in the next layer. Moreover, by processing the input tokens at each node in the first layer, the multi-agent intelligent network 100 is configured to perform the computational task efficiently, preparing the data for more complex operations in subsequent layers. As a result, the at least one or more nodes are configured to work in parallel rather than relying on centralized computation in order to reduce the overall data processing time along with the reduced communication overhead by limiting the amount of data that needs to be transmitted between layers.
[0065] At step 206, the method 200 includes sharing output information with nodes of the same current layer, and then share output information as input information to the one or more nodes of the subsequent layer, such as at step 208, whereby the node of the final layer provides processed output labels. Once each node in the current layer processes the input tokens, each of the node shares its output that includes the memory states and the counters with the one or more nodes within the same layer in order to ensure that all the nodes have consistent and synchronized information. The output information is then forwarded as input to the nodes in the next layer until the final layer receives the accumulated and processed data. The final layer then combines the received data to generate output labels, representing the final results of the distributed computation. As a result, the method 200 is used to ensure an efficient data flow along with the parallel processing of the data within and across layers thereby reducing the overall data computation time and resource usage.
[0066] In accordance with an embodiment, the method 200 further includes receiving a set of local input embeddings (Xn,Xl-X3) as the input information by each node. The node is configured to receive the set of local input embeddings as the input information to initialize the data processing at the local node level. Moreover, the set of local input embeddings, for example, a first local input embedding, a second local input embedding, and a third local input embedding, refers to a set of tokens that represents the input information about the local data input of the one or more nodes within the multi-agent intelligent network 100. As a result, the node is configured to receive the set of local input embeddings as the input information to process and interpret the local data efficiently for further computations and decision-making processes within the multi-agent intelligent network 100.
[0067] Furthermore, the method 200 includes applying positional embedding to the set of local input embeddings (Xn, X1-X3). In other words, the application of the positional embedding to the set of local input embeddings is used to encode the position of each input element in the sequence, allowing the node to distinguish between different positions. In an implementation, the node is configured to generate the positional embeddings, such as by using predefined or learned functions. Thereafter, the generated positional embeddings are added to the local input embeddings. As a result, by incorporating positional embedding, the node is configured to distinguish between different positions of the one or more nodes, thereby allowing the node to handle various types of sequential data effectively, making the multi-agent intelligent network 100 adaptable to different applications, such as text, audio, and video processing.
[0068] Furthermore, the method 200 includes determining a set of keys (Kn) by multiplying the set of local input embeddings (Xn,Xl- X3) with a learned key matrix (WK). In an implementation, the set of local input embeddings, such as the first local input embedding, the second local input embedding, and the third local input embedding, are multiplied with learned key matrix in order to obtain the set of keys, such as a first key, a second key, and a third key. Moreover, the determination of the set of keys involves transforming the set of local input embeddings by using the learned key matrix to provide the set of keys that are further utilized to compute the memory states. Moreover, by using the learned key matrix, the node is configured to provide an efficient and effective representation of the data that can be used for computing the linear attention.
[0069] Furthermore, the method 200 includes determining a set of values (Vn) by multiplying the set of local input embeddings (Xn, X1-X3) with a learned value matrix (Wv) and a set of queries (Qn) by multiplying the set of local input embeddings (Xn, XI- X3) with a learned query matrix (WQ). In an example, the set of local input embeddings, such as the first local input embedding, the second local input embedding, and the third local input embedding are multiplied with learned value matrix in order to obtain the set of values, such as a first value, a second value, and a third value. Similarly, the set of local input embeddings, such as the first local input embedding, the second local input embedding, and the third local input embedding, are multiplied with learned query matrix to determine the set of queries that includes a first query, a second query, and a third query. The determination of the set of queries and the set of values are used to improve the efficiency, scalability, and overall performance of the node while minimizing the communication overhead within the multi-agent intelligent network 100. Moreover, the node is characterized in that the node is further configured to determine a set of local memory states based on the set of local input embeddings, a learned key matrix, and a learned value matrix. In an implementation, the set of local memory states refers to input partial tokens that are associated with each node that is derived through a series of matrix multiplications involving the input token and a set of learned matrices, such as the set of values, set of keys, and the set of queries. Moreover, the local memory states, for example, a first local memory state, a second local memory state, and a third local memory state, are configured to encapsulate the information contained in the input tokens in a manner that is suitable for further processing within the node and for sharing the same with other nodes of the multi-agent intelligent network 100. Firstly, the set of input embeddings (or tokens) are received by the node locally. Thereafter, the set of keys, the set of values, and the set of queries are determined in order to further determine the set of local memory states for each node of the multi-agent intelligent network 100, such as based on the set of local embeddings, the set of learned key matrix, the learned value matrix. The set of local memory states allows the node to perform attention distribution in order to allow information sharing within the multi-agent intelligent network 100 locally, thereby maintaining the data privacy within the multi-agent intelligent network 100. As a result, the set of local memory states is determined to provide an efficient local processing, distributed attention, data privacy, and reduced communication overhead in the multi-agent intelligent network 100.
[0070] Furthermore, the method 200 includes determining a set of local memory states (S’n) based on the set of local input embeddings (Xn, X1-X3), a learned key matrix (WK) and a learned value matrix (Wv). By using learned matrices, the node is configured to transform the raw input embeddings into a format that highlights the key features and values in order to integrate and share such information with other nodes in the multi-agent intelligent network 100. By determining the local memory states by the multiplication of the set of input embeddings with the learned key and the learned value matrices, the multi-agent intelligent network 100 is configured to ensure that each node accurately encodes and processes the local data for efficient transformation of the raw data into meaningful representations, facilitating effective distributed processing and information sharing. Furthermore, the method 200 includes receiving one or more node memory states (S1) from at least one of the one or more other nodes, the received one or more node memory states forming a set of node memory states (Sn). By receiving the set of node memory states from the at least one or more other nodes, the node is configured to integrate external information with the local memory state of the node in order to facilitate a distributed linear attention mechanism without sharing raw data.
[0071] Furthermore, the method 200 includes determining a set of combined node memory states (Sn) based on the received set of node memory states (Sn) and the set of local memory states (S’n). In other words, each node of the multi-agent intelligent network 100 determines the local memory states from the set of input embeddings by using learned matrices and transmits the memory states to one or more other nodes. Finally, the received memory states are integrated with the local memory states of the node to further determine the combined memory states that are used to compute the linear attention. The combined memory state is scaled and rotated, which is added to the first memory state, for example, at operation in order to obtain a combined memory state. Furthermore, the combined memory state is scaled and rotated and is further add to the second memory state to obtain a combined memory state, for example, at operation. Thereafter, the combined memory state is scaled and rotated to further add with the third memory state, for example, to obtain a combined memory state. Hence, the combined memory state includes information from the combined memory state. As a result, the set of combined node memory states is used to facilitate distributed attention in order to ensure that each node has access to a comprehensive set of node combined memory states with an efficient and effective data integration and processing with reduced communication overhead, such as by transmitting aggregated states rather than raw data. Additionally, the determination of the set of node combined memory states is used to maintain data privacy, such as by sharing the combined node memory state with an improved overall performance of the multiagent intelligent network 100.
[0072] Furthermore, the method 200 includes providing the set of combined node memory states (Sn) to at least one next node. The set of combined memory states, which aggregates the information from all nodes within the current layer, is provided to the at least one next node to ensure that the subsequent nodes receive a comprehensive and cohesive representation of the processed data. Therefore, by forwarding the set of combined memory states, the method 200 is used to facilitate continuous and coherence in the distributed computation, allowing the next node to effectively integrate and utilize the accumulated information from the previous layer.
[0073] Furthermore, the method 200 includes determining an output (On) based on the determined set of queries (Qn) and the set of combined node memory states (Sn). Moreover, the determination of the output is calculated by utilizing the queries, which represent the specific aspects or features of the data that need to be focused on, and the combined memory states, which provide a comprehensive summary of the processed information from the current layer. By integrating these elements, the method 200 is used to generate the final output, ensuring that the generated final output reflects the nuanced and detailed processing performed across the distributed multi-agent intelligent network 100.
[0074] Advantageously, the method 200 is used for a signalling protocol in a multi-agent intelligent network 100 that utilizes distributed attention with K layers of nodes, where each layer processes input information using a multi-head linear attention model to generate output labels. Moreover, the method 200 involves processing input tokens in the first layer, sharing output information within the same layer, and then passing the output information to the subsequent layer thereby, ensuring that the input tokens are processed effectively, and the output labels are shared efficiently. In addition, the method 200 allows for distributed processing of input data across multiple nodes and layers, enabling the multi-agent intelligent network 100 to handle large-scale tasks that would be impractical or impossible for a single centralized communication system. Furthermore, by distributing the computation across multiple nodes the utilization of the computational resources are improved leading to an enhanced data processing times and improved scalability. The protocol is used to minimize the data transfer between nodes by sharing only processed output information rather than raw input data, which can significantly reduce network bandwidth requirements. As the input data is processed locally at each node and only output information is shared, the method 200 is used to maintain data privacy and security, which is crucial in many applications. Moreover, the multi-layer, multi-node structure allows for easy scaling of the network by adding more nodes or layers as needed, without fundamentally changing the protocol and allows for decentralized operation, reducing the need for a central coordinator and potentially improving fault tolerance. As a result, the method 200 can be used in a wide range of applications, from video processing and autonomous driving to distributed intelligent computing and smart manufacturing. Additionally, the method 200 is used to provide a signalling protocol for providing distributed attention in multi-agent intelligent networks, enabling an efficient, scalable, and privacypreserving processing of complex tasks.
[0075] The steps 202 to 208 are only illustrative, and other alternatives can also be provided where one or more steps are added, one or more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claims herein.
[0076] There is further provided a computer program product comprising program instructions for performing the method 200 when executed by one or more processors in the multi-agent intelligent network 100. The computer program product is implemented as an algorithm, embedded in a software stored in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage means may include, but are not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. Examples of implementation of computer-readable storage medium, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read Only Memory (ROM), Elard Disk Drive (1TDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), a computer-readable storage medium, and / or CPU cache memory.
[0077] FIG. 3 is a diagram that depicts a node configured to be used in a multi-agent intelligent network configured to provide a signaling protocol for providing distributed attention in the multi-agent intelligent network, in accordance with an embodiment of the present disclosure. With reference to FIG. 3, there is shown a node 302 that includes a controller 304, a memory 306, and a network interface 308.
[0078] The node 302 refers to an individual device within the multi-agent intelligent network 100 that processes local data and communicates with the one or more nodes, such as user equipment (UEs) and the like of the multi-agent intelligent network 100.
[0079] The controller 304 may include suitable logic, circuitry, interfaces, hardware, software, or code that is configured to provide a signaling protocol for providing distributed attention in the multi-agent intelligent network 100. Examples of implementation of the controller 304 may include but are not limited to a central data processing device, a microprocessor, a microcontroller, a complex instruction set computing (CISC) processor, an application-specific integrated circuit (ASIC) processor, a reduced instruction set (RISC) processor, a very long instruction word (VLIW) processor, a state machine, and other processors or control circuitry.
[0080] The memory 306 may include suitable logic, circuitry, interfaces, or code that is configured to store data for generating a relevant result. In an implementation, the memory 306 corresponds to a local memory, such as an Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read-Only Memory (ROM), a central processing unit (CPU) cache memory, and the like. In another implementation, the memory 306 corresponds to disc storage memory, such as a Hard Disk Drive (HDD), Flash memory, Solid-State Drive (SSD), and the like.
[0081] The network interface 308 includes hardware or software that is configured to establish communication among the controller 304 and the memory 306. Examples of the network interface 308 may include but are not limited to, a computer port, a network socket, a network interface controller (NIC), and any other network interface device. There is provided the node 302 to provide a signalling protocol for a multi-agent intelligent network 100 providing distributed attention in the multi-agent intelligent network. The signalling protocol is used to allow the multi-agent intelligent network 100 to provide an efficient, scalable, and privacy-preserving distributed processing. The signalling protocol is configured to minimize the communication overhead, such as by structuring data exchanges and ensures that sensitive local data remains private by preventing the data transmission to other nodes.
[0082] The multi-agent intelligent network 100 includes K layers of nodes and each layer includes a plurality of nodes and there is a first input layer of nodes, one or more intermediate layers of nodes and a final layer of nodes. In an implementation, the first input layer is used for receiving and processing the initial data or input tokens. The node 302 arranged in the first input layer is used to process the local portion of the data and begins the exchange of information with other nodes in the multi-agent intelligent network 100. Moreover, the intermediate layers are configured to handle the sequential processing of tokens and states, where nodes in each layer communicate both within their own layer and with nodes in adjacent layers. Additionally, the final layer of nodes is configured to perform the computation by producing the output tokens, which represent the final processed result of the multi-agent intelligent network 100. As a result, such layered structure allows an efficient distributed processing, where each node focuses on a specific part of the computation and coordinates with others through a well-defined signalling protocol and the multiple layers ensures that the multi-agent intelligent network 100 can scale to handle large models and complex tasks while maintaining efficiency and accuracy in the distributed processing environment.
[0083] Moreover, the node 302 is arranged to process input information through a multi-head linear attention model to provide output information. In an implementation, the node 302 is configured to capture complex relationships and dependencies between the tokens and such parallel processing is further used for handling large-scale data by ensuring that the node 302 can attend to multiple aspects of the input at once thereby, improving both processing speed and accuracy. Moreover, after processing the input, the node 302 generates the output information, which is then communicated to other nodes in the multi-agent intelligent network 100, either within the same layer or in subsequent layers. Therefore, by utilizing the multi-head linear attention model, the node 302 contributes to the overall computation in a distributed and coordinated manner, ensuring that the final output of the multi-agent intelligent network 100 that reflects the comprehensive analysis of all input data. As a result, an efficient distributed attention enables the multi-agent intelligent network 100 to handle complex tasks with high accuracy while maintaining scalability and minimal communication overhead.
[0084] Furthermore, if the node 302 is arranged in the first layer, the node 302 is configured to receive input tokens as input information and if the node 302 is arranged in the one or more intermediate layers, then, the node 302 is configured to receive input information from one or more nodes of a previous layer and send output information to one or more nodes of a subsequent layer. The node 302, which is arranged in the one or more intermediate layers are arranged to receive input information from one or more nodes in the previous layer acts as intermediaries in the distributed processing chain, taking the input information from the previous layer and using them as input to continue the computation. After processing the received input information, the node 302 sends the generated output information to one or more nodes in the subsequent layer, ensuring the continuous flow of data through the network. As a result, the structured flow of the input information between layers allows the multiagent intelligent network 100 to refine the data, eventually reaching the final layer where the complete output is produced, thereby ensuring seamless, scalable, and efficient processing throughout the multi-agent intelligent network 100.
[0085] Furthermore, is the node 302 is arranged in the final layer, the node 302 is configured to provide output information as output labels. After receiving and processing data through the preceding layers, the final layer nodes completes the distributed computation by generating the ultimate results or labels based on the processed input. Moreover, the output labels represent the final, comprehensive interpretation of the data as determined by the multi-agent intelligent network 100. As a result, the output labels are used to provide end-to-end computations along with the final conclusions or predictions of the multi-agent intelligent network 100. Furthermore, if the node 302 is in the first layer, the node 302 is configured to receive an acquired input sequence as the input tokens and, for each layer. Moreover, the node 302 is further configured to process its input to generate the output information of the node 302, share output information with nodes of the same current layer, and then share output information as input information to the one or more nodes of the subsequent layer, whereby if the node 302 is in the final layer, then, the node 302 is configured to provide processed output labels. The input token is distributed to all the nodes of the first layer and each node processes the one or more tokens that are allocated to the corresponding nodes thereby, capturing relevant features and creating memory states that summarize the information processed. Once processed, the output is passed on to the nodes in the next layer. Moreover, by processing the input tokens at each node in the first layer, the multi-agent intelligent network 100 is configured to perform the computational task efficiently, preparing the data for more complex operations in subsequent layers. As a result, the at least one or more nodes are configured to work in parallel rather than relying on centralized computation in order to reduce the overall data processing time along with the reduced communication overhead by limiting the amount of data that needs to be transmitted between layers. Once the node 302 in the current layer processes the input tokens, each of the node 302 shares its output that includes the memory states and the counters with the one or more nodes within the same layer in order to ensure that all the nodes have consistent and synchronized information. The output information is then forwarded as input to the nodes in the next layer until the final layer receives the accumulated and processed data. The final layer then combines the received data to generate output labels, representing the final results of the distributed computation. As a result, the node 302 is arranged to ensure an efficient data flow along with the parallel processing of the data within and across layers thereby reducing the overall data computation time and resource usage.
[0086] In accordance with an embodiment, the node 302 is configured to receive a set of local input embeddings (Xn, X1-X3) as the input information. Moreover, the node 302 is configured to receive the set of local input embeddings as the input information to process and interpret the local data efficiently for further computations and decision-making processes within the multiagent intelligent network 100.
[0087] Furthermore, the node 302 is configured to apply positional embedding to the set of local input embeddings (Xn, XI -X3). By incorporating positional embedding, the node 302 is configured to distinguish between different positions of the one or more nodes, thereby allowing the node to handle various types of sequential data effectively, making the multi-agent intelligent network 100 adaptable to different applications, such as text, audio, and video processing.
[0088] Furthermore, the node 302 is configured to determine a set of keys (Kn) by multiplying the set of local input embeddings (Xn, X1-X3) with a learned key matrix (WK). The determination of the set of keys involves transforming the set of local input embeddings by using the learned key matrix to provide the set of keys that are further utilized to compute the memory states. Moreover, by using the learned key matrix, the node 302 is configured to provide an efficient and effective representation of the data that can be used for computing the linear attention.
[0089] Furthermore, the node 302is configured to determine a set of values (Vn) by multiplying the set of local input embeddings (Xn, X1-X3) with a learned value matrix (Wv) and a set of queries (Qn) by multiplying the set of local input embeddings (Xn, XI- X3) with a learned query matrix (WQ). AS a result, the set of local memory states is determined to provide an efficient local processing, distributed attention, data privacy, and reduced communication overhead in the multi-agent intelligent network 100.
[0090] Furthermore, the node 302 is configured to determine a set of local memory states (S’n) based on the set of local input embeddings (Xn, X1-X3), a learned key matrix (WK) and a learned value matrix (Wv). By using learned matrices, the node 302 is configured to transform the raw input embeddings into a format that highlights the key features and values in order to integrate and share such information with other nodes in the multi-agent intelligent network 100. By determining the local memory states by the multiplication of the set of input embeddings with the learned key and the learned value matrices, the multi-agent intelligent network 100 is configured to ensure that each node accurately encodes and processes the local data for efficient transformation of the raw data into meaningful representations, facilitating effective distributed processing and information sharing.
[0091] Furthermore, the node 302 is configured to receive one or more node memory states (S1) from at least one of the one or more other nodes, the received one or more node memory states forming a set of node memory states (Sn). By receiving the set of node memory states from the at least one or more other nodes, the node 302 is configured to integrate external information with the local memory state of the node 302 in order to facilitate a distributed linear attention mechanism without sharing raw data.
[0092] Furthermore, the node 302 is configured to determine a set of combined node memory states (Sn) based on the received set of node memory states (Sn) and the set of local memory states (S’n). As a result, the set of combined node memory states is used to facilitate distributed attention in order to ensure that each node 302 has access to a comprehensive set of node combined memory states with an efficient and effective data integration and processing with reduced communication overhead, such as by transmitting aggregated states rather than raw data. Additionally, the determination of the set of node combined memory states is used to maintain data privacy, such as by sharing the combined node memory state with an improved overall performance of the multi-agent intelligent network 100.
[0093] Furthermore, the node 302 is configured to provide the set of combined node memory states (Sn) to at least one next node. By forwarding the set of combined memory states, the method 200 is used to facilitate continuous and coherence in the distributed computation, allowing the next node to effectively integrate and utilize the accumulated information from the previous layer.
[0094] Furthermore, the node 302 is configured to determine an output (On) based on the determined set of queries (Qn) and the set of combined node memory states (Sn). By integrating these elements, the method 200 is used to generate the final output, ensuring that the generated final output reflects the nuanced and detailed processing performed across the distributed multi-agent intelligent network 100.
[0095] Advantageously, the node 302 is configured to provide a signalling protocol in a multi-agent intelligent network 100 that utilizes distributed attention with K layers of nodes, where each layer processes input information using a multi-head linear attention model to generate output labels. Moreover, the node 302 is configured to process input tokens in the first layer, sharing output information within the same layer, and then passing the output information to the subsequent layer thereby, ensuring that the input tokens are processed effectively, and the output labels are shared efficiently. In addition, the node 302 is configured to allow the distributed processing of input data across multiple nodes and layers, enabling the multi-agent intelligent network 100 to handle large-scale tasks that would be impractical or impossible for a single centralized communication system. Furthermore, by distributing the computation across multiple nodes the utilization of the computational resources are improved leading to an enhanced data processing times and improved scalability. The protocol is used to minimize the data transfer between nodes by sharing only processed output information rather than raw input data, which can significantly reduce network bandwidth requirements. As the input data is processed locally at each node and only output information is shared, the node 302 is configured to maintain data privacy and security, which is crucial in many applications. Moreover, the multi-layer, multi-node structure allows for easy scaling of the network by adding more nodes or layers as needed, without fundamentally changing the protocol and allows for decentralized operation, reducing the need for a central coordinator and potentially improving fault tolerance. As a result, the node 302 is configured to be used in a wide range of applications, from video processing and autonomous driving to distributed intelligent computing and smart manufacturing. Additionally, the node 302 is configured to provide a signalling protocol for providing distributed attention in multi-agent intelligent networks, enabling an efficient, scalable, and privacy-preserving processing of complex tasks. FIG. 4A and 4B is a diagram that illustrates an exemplary scenario of signalling protocol for distributed inference over a network of devices with pieces of a large model with K layers of multi-head linear attention, in accordance with an embodiment of the present disclosure. FIG. 4A and 4B are described in conjunction with elements from FIG. 1 to 3. With reference to FIG. 4A, there is shown a diagram 400A that depicts an exemplary scenario of signalling protocol for distributed inference over the multi-agent intelligent network 100 of the one or more nodes with pieces of a large model within the same layer of multi-head linear attention.
[0096] In an implementation scenario, the signalling protocol consists of K rounds (i.e., one round for each layer) and each round includes two phases. The one or more nodes of the first layer acquire parts of the input sequence and token indices. During the first phase, each node of layer k processes the input tokens and forwards a message to the next node of the same layer that includes the ID of the sending device, ID of the layer, the index of the last processed token, ID of heads of the multi-head layer of linear attention, and related information. The first phase continues until the last node of layer k processes its input tokens. In this phase, the nodes of the first layer 406 process an acquired input sequence. For example, the first node of the first layer 406 processes its input sequence, such as a first input token 402A, a second input 402B, and a third input 402C using its local model with multi-head linear attention. The result of the processing includes memory states which gather information about the processed input tokens (i.e., one state per head of the multi-head attention layer) and a counter which indicates the total number of processed tokens by the first node. Moreover, such obtained states and the counter are then forwarded to the next node in the first layer 406, which processes its own input tokens, such as a fourth input token 402D, a fifth input token 402E, and a sixth input token 402F while attending to the received states. The result of this processing includes new memory states which gather information about all processed tokens and an updated counter indicating the total number of tokens processed so far at the first layer 406. Moreover, such process continues, with results being forwarded to subsequent nodes for further processing, until the last node of the first layer 406 processes its input tokens (e.g., 402N-1, 402N).
[0097] With reference to FIG. 4B, there is shown a diagram 400B that depicts an exemplary scenario of signalling protocol for distributed inference over the multi-agent intelligent network 100 of the one or more nodes with pieces of a large model within different layers of multi-head linear attention.
[0098] In an implementation scenario, the diagram 400B illustrates the flow of the signalling protocol across multiple layers that is from the first layer 406 through intermediate layer 408 to a final layer 410. Moreover, each layer in FIG. 4B corresponds to a round in the K-round signalling protocol. After the completion of processing within a layer (as described in FIG. 4A), the protocol moves to its second phase, which involves communication between consecutive layers. In this phase, the nodes of one layer (e.g., the first layer 406) use their previously obtained memory states to produce their local output sequence and the output tokens are then split and shared with selected devices of the subsequent layer (e.g., intermediate layer 408). Moreover, the message for the inter-layer communication includes ID of the sending node, ID of the layer, counters (indices) of the output tokens, and the corresponding output tokens. Additionally, as the process moves to the next layer (e.g., from the first layer 406 to the intermediate layer 408), the nodes in a new layer treats the received tokens as their input and then process these tokens using their local models with multi-head linear attention, following a similar procedure as described for the first layer in FIG. 4A. Furthermroe, such pattern of intra-layer processing followed by inter-layer communication continues through all K layers. The final layer 410 processes the input and produces the final output tokens (e.g., a first output token 404A, a second output token 404B, a third output token 404C, and the like), which collectively form the final output sequence of the distributed inference process. As a result, by distributing the workload and enabling parallel processing within layers while maintaining sequential processing across layers, the system achieves improved scalability and performance for inference tasks involving extensive computational requirements.
[0099] FIG. 5 is a diagram that illustrates an exemplary scenario of a vehicle monitored by a network of cameras, in accordance with an embodiment of the present disclosure. FIG. 5 is described in conjunction with elements from FIG. 1 to 4. With reference to FIG. 5, there is shown a diagram 500 that depicts a vehicle 502, which is moving in direction 512 and is monitored by the multi-agent intelligent network 100 of the cameras, such as a first camera 504A, a second camera 504B, and a third camera 504C.
[0100] In an implementation, the cameras, such as the first camera 504A, the second camera 504B, and the third camera 504C, are configured to provide the data, such as the captured frames (e.g., a first frame 510A, a second frame 510B, and a third frame 510C) to the associated nodes. For example, the first camera 504A is configured to provide the data captured to the first node 514A, and the second camera 504B is configured to provide the data captured to the second node 514B. Similarly, the third camera 504C is configured to provide the data captured to the third node 514C. Furthermore, each of the nodes transfers a last frame along with the combined memory state and the counter. For example, at operation 506, the first node 514A is configured to transfer the combined memory state along with the counter and a last frame to the second node 514B, and at operation 508, the second node 514B is configured to transfer the last frame along with the combined memory state along with the counter. As a result, the multi-agent intelligent network 100 can be utilized to perform computations that would normally require large computations and moving all recorded data at a single place for distributed linear attention, such as autonomous driving, distributed intelligent computing, smart manufacturing, e-health care, and the like.
[0101] Modifications to embodiments of the present disclosure described in the foregoing are possible without departing from the scope of the present disclosure as defined by the accompanying claims. Expressions such as "including", "comprising", "incorporating", "have", "is" used to describe and claim the present disclosure are intended to be construed in a non-exclusive manner, namely allowing for items, components or elements not explicitly described also to be present. Reference to the singular is also to be construed to relate to the plural. The word "exemplary" is used herein to mean "serving as an example, instance or illustration". Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or to exclude the incorporation of features from other embodiments. The word "optionally" is used herein to mean "is provided in some embodiments and not provided in other embodiments". It is appreciated that certain features of the present disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable combination or as suitable in any other described embodiment of the disclosure.
Claims
CLAIMS1. A method (200) for providing a signaling protocol for a multi-agent intelligent network (100) providing distributed attention, the multi-agent intelligent network (100) comprising K layers of nodes, each layer comprising a plurality of nodes, wherein there is a first input layer of nodes, one or more intermediate layers of nodes and a final layer of nodes, wherein each node is arranged to process input information through a multi-head linear attention model to provide output information, the nodes of the first layer are arranged to receive input tokens as input information the nodes of the one or more intermediate layers are arranged to receive input information from one or more nodes of a previous layer and send output information to one or more nodes of a subsequent layer, and the nodes of the final layer are arranged to provide output information as output labels, wherein the method (200) comprises providing the nodes of the first layer (406) with an acquired input sequence as the input tokens and then, for each layer, process the input in each node in the current layer to generate the output information of the node, share output information with nodes of the same current layer, and then share output information as input information to the one or more nodes of the subsequent layer, whereby the nodes of the final layer provide processed output labels.
2. The method (200) according to claim 1 , wherein the input information comprises one or more tokens to be processed, and output information of a previous node, if any.
3. The method (200) according to any preceding claim, wherein the output information of a node comprises memory states indicating information about processed input information and a counter which indicates the total number of processed tokens by the node.
4. The method (200) according to claim 3, wherein the memory states comprise one memory state per head of the layer.
5. The method (200) according to any preceding claim, wherein each node is arranged to share to the nodes in the same layer when processing the output information of that node; an identifier of the node, a counter indicating a number of tokens processed so far at that layer, memory states for each head of a multi-head layer of linear attention.
6. The method (200) according to claim 5, wherein at least one node is arranged to further share to the nodes in the same layer an identifier of a sub layer.
7. The method (200) according to any preceding claim, wherein each node is arranged to share to the one or more nodes in the subsequent layer; an identifier of the sending device, a counter of the output information.
8. The method (200) according to claim 7, wherein at least one node is further arranged to share an identifier of the layer.
9. The method (200) according to any preceding claim, wherein the method (200) further comprises each node (302) receiving a set of local input embeddings (Xn, X1-X3) as the input information,applying positional embedding to the set of local input embeddings (Xn, X1-X3), determining a set of keys (Kn) by multiplying the set of local input embeddings (Xn, X1-X3) with a learned key matrix (WK), determining a set of values (Vn) by multiplying the set of local input embeddings (Xn, X1-X3) with a learned value matrix (Wv), determining a set of queries (Qn) by multiplying the set of local input embeddings (Xn, X1-X3) with a learned query matrix (WQ), wherein the method (200) further comprises the node determining a set of local memory states (S’n) based on the set of local input embeddings (Xn, X1-X3), a learned key matrix (WK) and a learned value matrix (Wv), receiving one or more node memory states (S1) from at least one of the one or more other nodes, the received one or more node memory states forming a set of node memory states (Sn), determining a set of combined node memory states (Sn) based on the received set of node memory states (Sn) and the set of local memory states (S’n), providing the set of combined node memory states (Sn) to at least one next node, and determine an output (On) based on the determined set of queries (Qn) and the set of combined node memory states (Sn).
10. A computer program product comprising program instructions for performing the method (200) according to claim 9, when executed by one or more processors in a multi-agent intelligent network (100).
11. A multi-agent intelligent network (100) configured to provide a signaling protocol for providing distributed attention in the multi-agent intelligent network (100), the multi-agent intelligent network (100) comprising K layers of nodes, each layer comprising a plurality of nodes, wherein there is a first input layer of nodes, one or more intermediate layers of nodes and a final layer of nodes, wherein each node is arranged to process input information through a multi-head linear attention model to provide output information, the nodes of the first layer are arranged to receive input tokens as input information, the nodes of the one or more intermediate layers are arranged to receive input information from one or more nodes of a previous layer and send output information to one or more nodes of a subsequent layer, and the nodes of the final layer are arranged to provide output information as output labels, wherein the system is configured to provide the nodes of the first layer with an acquired input sequence as the input tokens and then, for each layer, process the input in each node in the current layer to generate the output information of the node, share output information with nodes of the same current layer, and then share output information as input information to the one or more nodes of the subsequent layer, whereby the nodes of the final layer provide processed output labels.
12. The multi-agent intelligent network (100) according to claim 11, wherein each node in the multi-agent intelligent network (100) is further configured to configured to receive a set of local input embeddings (Xn, X1-X3) as the input information, apply positional embedding to the set of local input embeddings (Xn, X1-X3), determine a set of keys (Kn) by multiplying the set of local input embeddings (Xn, X1-X3) with a learned key matrix (WK), determine a set of values (Vn) by multiplying the set of local input embeddings (Xn, X1-X3) with a learned value matrix (Wv),determine a set of queries (Qn) by multiplying the set of local input embeddings (Xn, XI -X3) with a learned query matrix (WQ), wherein the method further comprises the node determine a set of local memory states (S’n) based on the set of local input embeddings (Xn, X1-X3), a learned key matrix (WK) and a learned value matrix (Wv), receive one or more node memory states (S1) from at least one of the one or more other nodes, the received one or more node memory states forming a set of node memory states (Sn), determine a set of combined node memory states (Sn) based on the received set of node memory states (Sn) and the set of local memory states (S’n), provide the set of combined node memory states (Sn) to at least one next node, and determine an output (On) based on the determined set of queries (Qn) and the set of combined node memory states (Sn).
13. A node (302) configured to be used in a multi-agent intelligent network (100) configured to provide a signaling protocol for providing distributed attention in the multi-agent intelligent network (100), the multi-agent intelligent network (100) comprising K layers of nodes, each layer comprising a plurality of nodes, wherein there is a first input layer of nodes, one or more intermediate layers of nodes and a final layer of nodes, wherein the node (302) is arranged to process input information through a multi-head linear attention model to provide output information, and if the node (302) is in the first layer, the node is configured to receive input tokens as input information and if the node (302) is in the one or more intermediate layers, the node is configured to receive input information from one or more nodes of a previous layer and send output information to one or more nodes of a subsequent layer, and if the node (302) is in the final layer, the node is configured to provide output information as output labels, wherein if the node (302) is in the first layer, the node is configured to receive an acquired input sequence as the input tokens and, for each layer, the node (302) is further configured to process its input to generate the output information of the node (302), share output information with nodes of the same current layer, and then share output information as input information to the one or more nodes of the subsequent layer, whereby if the node (302) is in the final layer, the node (302) is configured to provide processed output labels.
14. The node (302) according to claim 12, wherein the node (302) is further configured to configured to receive a set of local input embeddings (Xn, X1-X3) as the input information, apply positional embedding to the set of local input embeddings (Xn, X1-X3), determine a set of keys (Kn) by multiplying the set of local input embeddings (Xn, X1-X3) with a learned key matrix (WK), determine a set of values (Vn) by multiplying the set of local input embeddings (Xn, X1-X3) with a learned value matrix (Wv), determine a set of queries (Qn) by multiplying the set of local input embeddings (Xn, XI -X3) with a learned query matrix (WQ), wherein the method further comprises the node determine a set of local memory states (S’n) based on the set of local input embeddings (Xn, X1-X3), a learned key matrix (WK) and a learned value matrix (Wv),receive one or more node memory states (Si) from at least one of the one or more other nodes, the received one or more node memory states forming a set of node memory states (Sn), determine a set of combined node memory states (Sn) based on the received set of node memory states (Sn) and the set of local memory states (S’n), provide the set of combined node memory states (Sn) to at least one next node, and determine an output (On) based on the determined set of queries (Qn) and the set of combined node memory states (Sn).